This is really great.
AI is now doing the research step that makes AI systems better, and doing it well enough to beat the people whose job that was.
Intology's automated research system, Locus, has post-trained a model that beats the official human-tuned Qwen3-1.7B release, scoring 51.6% against 49.4%.
PostTrainBench is the benchmark used to score AI agents that post-train other models, and it gives each agent one H100 and 10 hours. Intology's different path was to remove that limit and let agents run for 100 hours across a cluster.
So Intology raised the ceiling and called it PostTrainBench+, taking the total budget from 70 H100 hours to 4,500.
The capability being measured here is sustained experimental search: allocating compute, running parallel jobs, reading evaluations, abandoning weak branches, and scaling the promising ones.
One run per setting is a serious limitation. Still, this makes a good case that automated research systems need to be evaluated at the timescale where research decisions compound.