用户恳请 Google 保留 Gemini 2.5 Flash 模型
阅读原文· discuss.ai.google.dev多位用户在 Google Gemini API 论坛发帖,恳请保留 Gemini 2.5 Flash 模型。用户反馈,Gemini 3 Flash 和 3.1 Flash Lite 在内部基准测试中性能均不及 2.5 Flash,且 3.5 Flash 延迟更高(600-700ms vs 2.5 Flash 的 300-400ms),在澳大利亚等地区未部署,实际延迟达 700-800ms,无法用于语音智能体。此外,3.5 Flash 价格约为 2.5 Flash 的 3 倍。部分用户表示若该模型停用,将转向开源模型。
Please don't discontinue Gemini 2.5 Flash
Gemini API
models
,
feedback
,
api
,
gemini
Nick_D
July 10, 2026, 12:13am
1
Hi Everyone -
Firstly, I want to give my thanks to the Gemini team for providing access to such great models. It has been so helpful.
We have some very specific workflows that rely on Gemini 2.5 Flash. Our internal benchmarks show that Gemini 3 flash does not perform as well (even after attempting to tweak prompting following the new prompting guidelines and other changes).
I am sure there are others with the same experience as well, and there is no easy switch.
It would be extremely appreciated if 2.5 flash was not discontinued.
Ruthvik
July 10, 2026, 8:16am
2
Yes, even our own benchmarks show that the closest model in latency + performance, which is 3.1 flash lite, doesn’t even come close to 2.5 Flash. Seeing issues with thoughts leaking out. I should say, 2.5 flash has been the best model we’ve seen for all round usage and works pretty well with most of tasks.
I am pretty sure, a huge chunk of their traffic and usage comes from this model. Would really appreciate it if the Google team can retain this model for much longer time.
Joshua_Simpson
July 10, 2026, 6:09pm
3
Preach. Retiring 2.5 Flash will be such a massive loss. It’s the only low latency model that is deployed in Australia. I’m able to get 300-400ms completions from 2.5 Flash, making it suitable for voice agents.
3.5 Flash offers 600-700ms completions and doesn’t even get an Australian deployment, so the actual latency is closer to 700-800ms, completely breaking the voice agent use case.
There is honestly no other model deployed in this part of the world that comes close to the quality that 2.5 Flash offers for low latency applications.
tylertreat
July 10, 2026, 9:42pm
4
I am more concerned about the cost step up from Gemini 2.5 Flash to 3.5 Flash, with the latter being roughly 3x more expensive. I thought the intention of the Flash models was to be relatively low-latency and more affordable compared to Pro, but the newer Flash models aren’t being priced as such. Then again, the era of cheap and plentiful AI might be coming to an end…
Bcoun
July 10, 2026, 10:11pm
5
Seconded, the alternative for us is not upgrading to flash-3 but rather finding an appropriate open weight model 
merc
July 10, 2026, 10:26pm
6
+1 I have some critical workflows that no other model is good at for the same price/intelligence! It would be a huge hit to have this model discontinued - we’d likely switch to an open source model if this happened but the latency of 2.5 flash is something hard to beat.