elvis · @omarsar0 · X·2026-09-03 06:22·55分钟前
elvis@omarsar0
41AI 编辑部评分,满分 100
2026-09-03 06:22· 55分钟前

Brilliant effort worth checking out.

Ultra-long horizon coding tasks are where frontier models like Fable 5.1 will shine.

But that's a crazy gap (over ~25 percentage points).

What I think could be interesting is seeing results for a mixture of agents like what Cursor did.

ProximalWe are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large perfo...