Wtf! Ben ran this mystery model through 10 DeepSWE tasks and it scored over 80%, versus 65% for Fable and 52% for GPT-5.6-sol!
this is insane. Probably a chinese company. Either a new GLM or Kimi model, I reckon.
I ran this thing through 10 tasks on DeepSWE (so there could be a ton of variance in it's real score, this is a subset), but uh... gpt-5.6-sol: 52% fable: 65% w...