Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.
LongWoF-Bench:用 EvoMap 基因评估可验证的长流程任务
AI 导读
研究团队提出 LongWoF-Bench,含 778 个可机器验证任务,覆盖代码生成、智能体环境合成、数学推理与规则遵循。在 252 个经验证的 Opus 轨迹上,进化得到的 EvoMap 基因在全部 7 个模型中比 Skill 平均高出 8.7-15.5 个百分点,且优势扩展至消费级模型;对 Claude Opus,基因复用还多完成 39 个任务,并减少 9.9% 的推理 token 消耗。
HuggingFace Daily Papers(社区热门论文)
45
AI 编辑部评分,满分 100LongWoF-Bench:用 EvoMap 基因评估可验证的长流程任务
研究团队提出 LongWoF-Bench,含 778 个可机器验证任务,覆盖代码生成、智能体环境合成、数学推理与规则遵循。在 252 个经验证的 Opus 轨迹上,进化得到的 EvoMap 基因在全部 7 个模型中比 Skill 平均高出 8.7-15.5 个百分点,且优势扩展至消费级模型;对 Claude Opus,基因复用还多完成 39 个任务,并减少 9.9% 的推理 token 消耗。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org