非常吊诡的是,Opus 5 几乎在每一项指标上都超越了 Fable 5 也许这只能说明我们的所有指标都已经失效了 而且 Claude 系列的模型用起来的味道不太对劲,现在味道比较对的是 Kimi 和 Grok 对 RLVR 而言,重要的只有结果本身,而人类的偏好并没有那么重要 对齐计划,失败了吗?
opus 5 is a VERY interesting release for a few reasons 1. it showed that the general benchmarks we use today are almost completely useless now opus 5 is nowhere...