Another exciting reasoning benchmark, MazeBench, where GPT-Astra outperforms any other model, especially Fable 5.1.
At this point it’s clear to me that Astra is by far the best model. And since we already know that OpenAI is close to release another model, I couldn’t be more exited.
MazeBench vs GPT-6 Astra Astra spent 60+ hours in this 3D open world spatial reasoning eval. Final score: 14%