本周 Chatbot Arena 与 Artificial Analysis 排行榜出现重大变动。头部两家实验室地位稳固,但 Grok 和 Muse Spark 排名显著上升,对 Gemini 和 GLM 形成压力,不过与前两名仍有明显差距。推文指出,模型智能提升速度日益关键,实验室需快速迭代以保持竞争力。更频繁的模型发布正为持续学习奠定基础,若趋势延续,排行榜可能最终实现每小时更新,以反映模型独立进化而非依赖低频次大版本发布。发布周期已从季度缩短至月度,年底前双周发布或成新常态。
This week brought one of the biggest shake-ups we've seen on both the @arena and @ArtificialAnlys Index leaderboards.
The top two labs remain firmly in place, but Grok and Muse Spark made notable gains, putting pressure on Gemini and GLM. Even so, the performance gap to the top two is still significant.
What's becoming increasingly important is the pace of intelligence improvement. It's no longer just about reaching the top once. Labs need to iterate quickly to remain competitive.