Chubby♨️ · @kimmonismus · X·2026-08-21 16:30·4天前
AI 导读

Wtf!Ben 让这个神秘模型跑了 10 个 DeepSWE 任务,得分超过 80%,而 Fable 是 65%,GPT-5.6-sol 只有 52%! 这太疯狂了。可能是家中国公司。我猜要么是新版 GLM,要么是 Kimi 模型。

Chubby♨️@kimmonismus
35AI 编辑部评分,满分 100
2026-08-21 16:30· 4天前
AI 导读

Wtf!Ben 让这个神秘模型跑了 10 个 DeepSWE 任务,得分超过 80%,而 Fable 是 65%,GPT-5.6-sol 只有 52%! 这太疯狂了。可能是家中国公司。我猜要么是新版 GLM,要么是 Kimi 模型。

Wtf! Ben ran this mystery model through 10 DeepSWE tasks and it scored over 80%, versus 65% for Fable and 52% for GPT-5.6-sol!

this is insane. Probably a chinese company. Either a new GLM or Kimi model, I reckon.

Ben DavisI ran this thing through 10 tasks on DeepSWE (so there could be a ton of variance in it's real score, this is a subset), but uh... gpt-5.6-sol: 52% fable: 65% w...

来源:Chubby♨️· x.com