elvis@omarsar0
36AI 编辑部评分,满分 100
2026-08-01 00:45· 26分钟前
跳到正文
AI 摘要

微软提出Echoverse,通过将规格编译为有状态应用并基于应用数据库评分任务,配合共进化循环(读取评分轨迹两次,分别用于修复环境和生成训练信号),规模化训练计算机操作智能体。

New research from Microsoft.

This one is on training computer-use agents at scale.

Recent pipelines generate synthetic environments in bulk, which moved the bottleneck from how many exist to what is inside each one. Echoverse compiles specifications into stateful applications whose tasks are graded against the application's own database, then runs a co-evolution loop that reads every graded rollout twice. Once as repairs to the environment, its tasks and its verifier. Once as training signal for the model.

On the same domains, shallow environments pushed live-site accuracy below the base model, from 80.0 down to 75.0. Deep ones raised it, 80.0 to 85.0 and 48.0 to 65.0.

Repairing a single environment lifted the model trained on it from 16.2% to 38.5%. Across twelve environments, a 9B model went from 36.5% to 67.1% on fourteen evaluation splits, within fourteen points of the much larger frontier model that taught it.

They release four environments as a benchmark with applications, seed data and grounded graders.

Paper: https://arxiv.org/abs/2607.28074

Track trending AI papers in our academy: https://academy.dair.ai/

elvis · @omarsar0 · X·2026-08-01 00:45·26分钟前
在 X 看原推· x.com
AI 摘要

微软提出Echoverse,通过将规格编译为有状态应用并基于应用数据库评分任务,配合共进化循环(读取评分轨迹两次,分别用于修复环境和生成训练信号),规模化训练计算机操作智能体。

New research from Microsoft.

This one is on training computer-use agents at scale.

Recent pipelines generate synthetic environments in bulk, which moved the bottleneck from how many exist to what is inside each one. Echoverse compiles specifications into stateful applications whose tasks are graded against the application's own database, then runs a co-evolution loop that reads every graded rollout twice. Once as repairs to the environment, its tasks and its verifier. Once as training signal for the model.

On the same domains, shallow environments pushed live-site accuracy below the base model, from 80.0 down to 75.0. Deep ones raised it, 80.0 to 85.0 and 48.0 to 65.0.

Repairing a single environment lifted the model trained on it from 16.2% to 38.5%. Across twelve environments, a 9B model went from 36.5% to 67.1% on fourteen evaluation splits, within fourteen points of the much larger frontier model that taught it.

They release four environments as a benchmark with applications, seed data and grounded graders.

Paper: https://arxiv.org/abs/2607.28074

Track trending AI papers in our academy: https://academy.dair.ai/

在 X 查看原推x.com