Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs (5.4%) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores 47.0/100. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks, yet only 26% reach 100%. Agent capability differs across migration categories: agents score 31.4 on build toolchain rewrites but only 5.6 on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
SWE Refactor Bench:编码智能体能否完成长周期、全仓库的栈迁移?
AI 导读
SWE Refactor Bench 新基准用 20 个全仓库迁移任务、三阶段评估协议(迁移审计、行为测试、智能体验证)衡量编码智能体的迁移能力。8 个前沿模型的 520 次运行中仅 28 次(5.4%)通过全部阶段,最佳模型 claude-opus-5 得分 47.0/100。迁移完整性与行为正确性是两种独立能力,智能体在构建工具链重写得分 31.4,语言重写仅 5.6。
HuggingFace Daily Papers(社区热门论文)
46
AI 编辑部评分,满分 100SWE Refactor Bench:编码智能体能否完成长周期、全仓库的栈迁移?
SWE Refactor Bench 新基准用 20 个全仓库迁移任务、三阶段评估协议(迁移审计、行为测试、智能体验证)衡量编码智能体的迁移能力。8 个前沿模型的 520 次运行中仅 28 次(5.4%)通过全部阶段,最佳模型 claude-opus-5 得分 47.0/100。迁移完整性与行为正确性是两种独立能力,智能体在构建工具链重写得分 31.4,语言重写仅 5.6。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org