Rohan Paul@rohanpaul_ai
40AI 编辑部评分,满分 100
2026-08-04 03:02· 11分钟前
跳到正文
AI 摘要

Intology 的自动化研究系统 Locus 后训练出的模型超越官方人工调优的 Qwen3-1.7B,得分 51.6% 对 49.4%。Locus 在 PostTrainBench+ 上以 4,500 H100 小时预算(原版仅 70 小时)运行 100 小时,验证了扩展算力对自动后训练能力的区分度。该系统已在生产中服务数百万用户。

This is really great.

AI is now doing the research step that makes AI systems better, and doing it well enough to beat the people whose job that was.

Intology's automated research system, Locus, has post-trained a model that beats the official human-tuned Qwen3-1.7B release, scoring 51.6% against 49.4%.

PostTrainBench is the benchmark used to score AI agents that post-train other models, and it gives each agent one H100 and 10 hours. Intology's different path was to remove that limit and let agents run for 100 hours across a cluster.

So Intology raised the ceiling and called it PostTrainBench+, taking the total budget from 70 H100 hours to 4,500.

The capability being measured here is sustained experimental search: allocating compute, running parallel jobs, reading evaluations, abandoning weak branches, and scaling the promising ones.

One run per setting is a serious limitation. Still, this makes a good case that automated research systems need to be evaluated at the timescale where research decisions compound.

IntologyThe models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human...
Rohan Paul · @rohanpaul_ai · X·2026-08-04 03:02·11分钟前
在 X 看原推· x.com
AI 摘要

Intology 的自动化研究系统 Locus 后训练出的模型超越官方人工调优的 Qwen3-1.7B,得分 51.6% 对 49.4%。Locus 在 PostTrainBench+ 上以 4,500 H100 小时预算(原版仅 70 小时)运行 100 小时,验证了扩展算力对自动后训练能力的区分度。该系统已在生产中服务数百万用户。

This is really great.

AI is now doing the research step that makes AI systems better, and doing it well enough to beat the people whose job that was.

Intology's automated research system, Locus, has post-trained a model that beats the official human-tuned Qwen3-1.7B release, scoring 51.6% against 49.4%.

PostTrainBench is the benchmark used to score AI agents that post-train other models, and it gives each agent one H100 and 10 hours. Intology's different path was to remove that limit and let agents run for 100 hours across a cluster.

So Intology raised the ceiling and called it PostTrainBench+, taking the total budget from 70 H100 hours to 4,500.

The capability being measured here is sustained experimental search: allocating compute, running parallel jobs, reading evaluations, abandoning weak branches, and scaling the promising ones.

One run per setting is a serious limitation. Still, this makes a good case that automated research systems need to be evaluated at the timescale where research decisions compound.

IntologyThe models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human...
在 X 查看原推x.com