BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6.
The benchmark tests real-world AI agents across conversation, tool use, memory and safety.
BREAKING: Grok 4.5 在 LaurenBench 上以 56.9% 的得分排名第一,领先于 Claude Sonnet 5、GLM 5.2、Claude Opus 5、Kimi K3 和 GPT-5.6。 该基准测试评估真实世界 AI 智能体在对话、工具使用、记忆和安全方面的表现。
BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6.
The benchmark tests real-world AI agents across conversation, tool use, memory and safety.
BREAKING: Grok 4.5 在 LaurenBench 上以 56.9% 的得分排名第一,领先于 Claude Sonnet 5、GLM 5.2、Claude Opus 5、Kimi K3 和 GPT-5.6。 该基准测试评估真实世界 AI 智能体在对话、工具使用、记忆和安全方面的表现。
BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6.
The benchmark tests real-world AI agents across conversation, tool use, memory and safety.