What!? It looks correct.
I am convinced we are not pushing these LLMs enough. Benchmarks are simply not enough. That's exciting.
Now can you please remove all the unnecessary guardrails from Fable so more of us can try solving extremely hard problems in our respective fields?