Grok 4.6 ranks #1 on RuntimeWire's Newsroom Reliability v0.2 benchmark with a score of 0.79, beating GPT-5.6 Sol, Claude Opus 4.8, Gemini and DeepSeek.
AI 导读
Grok 4.6 在 RuntimeWire 的 Newsroom Reliability v0.2 基准测试中排名第一,得分 0.79,超越了 GPT-5.6 Sol、Claude Opus 4.8、Gemini 和 DeepSeek。
Grok 4.6 ranks #1 on RuntimeWire's Newsroom Reliability v0.2 benchmark with a score of 0.79, beating GPT-5.6 Sol, Claude Opus 4.8, Gemini and DeepSeek.
来源:DogeDesigner· x.com