Apodex released Apodex 1.1, its proprietary model that reached 44 on the Artificial Analysis Intelligence Index and performs strongly on agentic tasks versus models in its tier.
on GDPval-AA v2, which measures real-world agentic work, it reaches 1,348 Elo, ahead of several models with much higher general intelligence scores.
So Apodex 1.1, its performance seems concentrated around professional and agentic tasks rather than being evenly distributed across the evaluation suite.
I increasingly think this distinction matters for model selection. A model that wins broad reasoning benchmarks is not automatically the model you want sitting inside an agentic execution loop.
For agents, the relevant question is: once you give the model a goal and tools, how often does it actually get the job done?
Apodex 1.1 from @Apodex_AI looks unusually concentrated in that direction.