DogeDesigner · @cb_doge · X·2026-07-15 20:20·41天前
AI 导读

BREAKING: Grok 4.5 现已在 FrontierSWE 上排名第二,在测试人类能力极限软件工程的基准测试中,表现优于 Claude Opus 4.8 和 GLM-5.2。 Grok 已能与最强者竞争。进步令人难以置信。🔥

DogeDesigner@cb_doge
58AI 编辑部评分,满分 100
2026-07-15 20:20· 41天前
AI 导读

BREAKING: Grok 4.5 现已在 FrontierSWE 上排名第二,在测试人类能力极限软件工程的基准测试中,表现优于 Claude Opus 4.8 和 GLM-5.2。 Grok 已能与最强者竞争。进步令人难以置信。🔥

BREAKING: Grok 4.5 is now ranked #2 on FrontierSWE, outperforming Claude Opus 4.8 and GLM-5.2 on a benchmark testing software engineering at the edge of human ability.

Grok is already competing with the very best. The progress is incredible. 🔥

来源:DogeDesigner· x.com