Grok 4.6 takes top spot on this benchmark
🥇 We evaluated Grok 4.6 on MedAgentBench, a benchmark for agentic clinical EHR tasks, and it took the top spot. Grok 4.6 posts the highest pass@1 we've measure...
xAI 的 Grok 4.6 在 MedAgentBench(智能体临床 EHR 任务基准)上登顶,以约 95.9% 的 pass@1(3 次运行均值)超越前榜首 GPT-5.6 Sol(约 94.7%),并较 Grok 4.5 提升约 2.5 个百分点。
xAI 的 Grok 4.6 在 MedAgentBench(智能体临床 EHR 任务基准)上登顶,以约 95.9% 的 pass@1(3 次运行均值)超越前榜首 GPT-5.6 Sol(约 94.7%),并较 Grok 4.5 提升约 2.5 个百分点。
Grok 4.6 takes top spot on this benchmark
来源:Elon Musk· x.com