# Grok 4.6 登顶 MedAgentBench 临床基准

- 来源：Elon Musk (@elonmusk)
- 发布时间：2026-08-18 13:58
- AIHOT 分数：40
- AIHOT 链接：https://aihot.virxact.com/items/cmsya1wze0pd0roz0qqpi657m
- 原文链接：https://x.com/elonmusk/status/2089592732780032132

## AI 摘要

xAI 的 Grok 4.6 在 MedAgentBench（智能体临床 EHR 任务基准）上登顶，以约 95.9% 的 pass@1（3 次运行均值）超越前榜首 GPT-5.6 Sol（约 94.7%），并较 Grok 4.5 提升约 2.5 个百分点。

## 正文

Grok 4.6 takes top spot on this benchmark

### 引用推文

> Medical Sphere：🥇 We evaluated Grok 4.6 on MedAgentBench, a benchmark for agentic clinical EHR tasks, and it took the top spot. Grok 4.6 posts the highest pass@1 we've measure...
