Agent memory benchmarks have had a basic attribution problem.
Every AI memory startup claims better recall, and until now there was no way to check.
That is why this new shared benchmark caught my attention.
It puts every entrant through the same pipeline: 5,000 questions, one fixed answer model, and the same judging process across systems.
The benchmark also separates open-source and commercial systems. That feels like the right choice: community projects can compete for prizes without being directly compared against heavily funded commercial products.
A group of 20+ research institutions is running all of them through one identical pipeline and publishing the results.
Entries close on August 7, with the first public rankings expected in mid-August.
I'll be watching the first results closely.