Brilliant effort worth checking out.
Ultra-long horizon coding tasks are where frontier models like Fable 5.1 will shine.
But that's a crazy gap (over ~25 percentage points).
What I think could be interesting is seeing results for a mixture of agents like what Cursor did.
We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large perfo...