Elon Musk · @elonmusk · X·2026-08-18 13:58·6小时前
AI 导读

xAI 的 Grok 4.6 在 MedAgentBench(智能体临床 EHR 任务基准)上登顶,以约 95.9% 的 pass@1(3 次运行均值)超越前榜首 GPT-5.6 Sol(约 94.7%),并较 Grok 4.5 提升约 2.5 个百分点。

Elon Musk@elonmusk
40AI 编辑部评分,满分 100
2026-08-18 13:58· 6小时前
AI 导读

xAI 的 Grok 4.6 在 MedAgentBench(智能体临床 EHR 任务基准)上登顶,以约 95.9% 的 pass@1(3 次运行均值)超越前榜首 GPT-5.6 Sol(约 94.7%),并较 Grok 4.5 提升约 2.5 个百分点。

Grok 4.6 takes top spot on this benchmark

Medical Sphere🥇 We evaluated Grok 4.6 on MedAgentBench, a benchmark for agentic clinical EHR tasks, and it took the top spot. Grok 4.6 posts the highest pass@1 we've measure...