Artificial Analysis@ArtificialAnlys
63AI 编辑部评分,满分 100
2026-08-05 02:05· 16分钟前
跳到正文
AI 摘要

Artificial Analysis 发布端点精度指数,以官方权重自托管部署为基准(100%),从工具调用、科学推理和长上下文召回三维度评测各 serverless API 端点。首批覆盖 GLM-5.2、gpt-oss-120b 和 DeepSeek V4 Pro,Kimi K3 即将加入。结果显示输出 token 限制和工具调用处理方式会显著拉低部分端点得分。

Announcing the Artificial Analysis Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves. We are initiating coverage with GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 coming soon

Providers trade off accuracy to optimize for speed and cost. They quantize weights, write custom kernels and tune their inference stacks, and sometimes they simply ship bugs. We are bringing the rigor of our Artificial Analysis Intelligence Index to measuring endpoints, so developers can pick providers on accuracy, not just price and speed

We benchmark each serverless endpoint against our own self-hosted reference deployment of the official weights, where 100% represents matching the reference. An endpoint is at reference parity when its result falls within the 95% confidence interval of the reference. Coverage is live for GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 accuracy coverage launching soon

Key elements of the Endpoint Accuracy Index: ➤ Three areas, equally weighted: tool calling (BFCL-500, 500 questions, 3 repeats), scientific reasoning (HLE-250, 250 questions, 10 repeats) and long context recall (AA-LCR-25, 25 questions, 10 repeats). Each subset separates endpoints on the serving choices that drive accuracy differences, with repeats sized for tight confidence intervals ➤ Reference deployment: we self-host the official weights at the lab's recommended precision, following the lab's serving recipe, and publish the complete commands for each reference ➤ Inference parameters: we run the model's highest supported reasoning mode and each endpoint's highest supported output length and context window ➤ Confidence intervals: the parity test accounts for uncertainty in both the endpoint's runs and the reference's runs ➤ Rotating coverage: models enter once sufficient number of providers serve them and exit when a newer version in the same family supersedes them. We benchmark new endpoints as providers launch them and refresh all listed endpoints periodically ➤ Point in time: each result carries the date it was measured, with multi-day benchmarks dated to their final day

Key results for GLM-5.2 ➤ Output token limits restrict accuracy. Restrictive limits cut responses off before the model finishes reasoning, and the most restrictive endpoints score half the reference or less on HLE-250

Key results for gpt-oss-120b ➤ Tool call handling separates endpoints. Providers parse and format tool calls differently, and some endpoints score 22% on BFCL-500 against 37% for the reference ➤ Serving configuration changes what the model does at the same requested settings. Some endpoints produce far fewer reasoning tokens at the same configured level, and restricted context windows truncate long context tasks

Key results for DeepSeek V4 Pro ➤ DeepSeek V4 Pro endpoints are more in line with the reference. Majority of the endpoints are at reference parity, and DeepSeek's own first-party endpoint scores slightly above the reference

Artificial Analysis · @ArtificialAnlys · X·2026-08-05 02:05·16分钟前
在 X 看原推· x.com(在新标签页打开)
AI 摘要

Artificial Analysis 发布端点精度指数,以官方权重自托管部署为基准(100%),从工具调用、科学推理和长上下文召回三维度评测各 serverless API 端点。首批覆盖 GLM-5.2、gpt-oss-120b 和 DeepSeek V4 Pro,Kimi K3 即将加入。结果显示输出 token 限制和工具调用处理方式会显著拉低部分端点得分。

Announcing the Artificial Analysis Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves. We are initiating coverage with GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 coming soon

Providers trade off accuracy to optimize for speed and cost. They quantize weights, write custom kernels and tune their inference stacks, and sometimes they simply ship bugs. We are bringing the rigor of our Artificial Analysis Intelligence Index to measuring endpoints, so developers can pick providers on accuracy, not just price and speed

We benchmark each serverless endpoint against our own self-hosted reference deployment of the official weights, where 100% represents matching the reference. An endpoint is at reference parity when its result falls within the 95% confidence interval of the reference. Coverage is live for GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 accuracy coverage launching soon

Key elements of the Endpoint Accuracy Index: ➤ Three areas, equally weighted: tool calling (BFCL-500, 500 questions, 3 repeats), scientific reasoning (HLE-250, 250 questions, 10 repeats) and long context recall (AA-LCR-25, 25 questions, 10 repeats). Each subset separates endpoints on the serving choices that drive accuracy differences, with repeats sized for tight confidence intervals ➤ Reference deployment: we self-host the official weights at the lab's recommended precision, following the lab's serving recipe, and publish the complete commands for each reference ➤ Inference parameters: we run the model's highest supported reasoning mode and each endpoint's highest supported output length and context window ➤ Confidence intervals: the parity test accounts for uncertainty in both the endpoint's runs and the reference's runs ➤ Rotating coverage: models enter once sufficient number of providers serve them and exit when a newer version in the same family supersedes them. We benchmark new endpoints as providers launch them and refresh all listed endpoints periodically ➤ Point in time: each result carries the date it was measured, with multi-day benchmarks dated to their final day

Key results for GLM-5.2 ➤ Output token limits restrict accuracy. Restrictive limits cut responses off before the model finishes reasoning, and the most restrictive endpoints score half the reference or less on HLE-250

Key results for gpt-oss-120b ➤ Tool call handling separates endpoints. Providers parse and format tool calls differently, and some endpoints score 22% on BFCL-500 against 37% for the reference ➤ Serving configuration changes what the model does at the same requested settings. Some endpoints produce far fewer reasoning tokens at the same configured level, and restricted context windows truncate long context tasks

Key results for DeepSeek V4 Pro ➤ DeepSeek V4 Pro endpoints are more in line with the reference. Majority of the endpoints are at reference parity, and DeepSeek's own first-party endpoint scores slightly above the reference

在 X 查看原推x.com(在新标签页打开)