Arena’s AI leaderboard has become a $100M annualized revenue business.
By turning public model comparisons into paid performance testing for AI labs and enterprises.
Arena began as a UC Berkeley research project that asked users to compare 2 anonymous model answers and vote for the better one.
That setup created a large human preference dataset, because every vote says something about what people value in AI responses.
Model labs care about those votes because benchmarks alone often miss the messy cases where users judge tone, reasoning, code quality, visual skill, or task completion.
Arena’s commercial move was to package that public testing engine into AI Evaluations, a service that gives customers deeper analytics from the same community feedback loop.
The business works because model companies badly need high-quality human preference signals after training, since small ranking gains can decide which model wins users, enterprise contracts, and investor attention.
---
techcrunch. com/2026/06/29/arena-the-ai-leaderboard-everyone-uses-is-now-a-100m-business/