# Opus 5 基准测试全面超越 Fable 5，但实际体验远不及后者，公开基准几乎失效

- 来源：Orange AI (@oran_ge)
- 发布时间：2026-07-27 06:05
- AIHOT 分数：39
- AIHOT 链接：https://aihot.virxact.com/items/cms2db9dh00b3royzno0jtslg
- 原文链接：https://x.com/oran_ge/status/2081501133168947412

## AI 摘要

Opus 5 在多项通用基准测试上超越 Fable 5，但实际使用体验远不及后者，表明现有公开基准几乎完全失效。Anthropic 在 5 系列中采用先训练 Mythos 再蒸馏出 Sonnet 和 Opus 的新路线，但 Sonnet 5 表现不佳、Opus 5 评价两极。

## 正文

非常吊诡的是，Opus 5 几乎在每一项指标上都超越了 Fable 5
也许这只能说明我们的所有指标都已经失效了
而且 Claude 系列的模型用起来的味道不太对劲，现在味道比较对的是 Kimi 和 Grok
对 RLVR 而言，重要的只有结果本身，而人类的偏好并没有那么重要
对齐计划，失败了吗？

### 引用推文

> Kun Chen：opus 5 is a VERY interesting release for a few reasons 1. it showed that the general benchmarks we use today are almost completely useless now opus 5 is nowhere...
