MT-EditFlow:基于流匹配的多轮图像编辑强化学习

Apple Machine Learning Research(RSS)·2026-07-07 08:00·48天前
AI 导读

针对单轮图像编辑模型在多轮交互中因单轮失败导致序列中断、误差累积的问题,提出 MT-EditFlow 流匹配强化学习框架。该框架整合多轮视角与多奖励公式,统一支持 GRPO 和 NFT 类强化学习方法,并系统分析轮次聚合评分策略、VLM 推理模式(权衡奖励偏差与方差)以及优势融合层级(防止奖励黑客行为)。研究发现将聚合优势广播至整个编辑轨迹能弥合局部规划与全局多轮任务成功之间的差距。实验表明,MT-EditFlow 在多种基模型上显著提升性能,其中 FLUX.1-Kontext-dev 的 turn-3 整体性能提升 6.85 分,超越 Qwen-Image-Edit 等开源模型。

Apple Machine Learning Research(RSS)
40AI 编辑部评分,满分 100

MT-EditFlow:基于流匹配的多轮图像编辑强化学习

2026-07-07 08:00· 48天前
AI 导读

针对单轮图像编辑模型在多轮交互中因单轮失败导致序列中断、误差累积的问题,提出 MT-EditFlow 流匹配强化学习框架。该框架整合多轮视角与多奖励公式,统一支持 GRPO 和 NFT 类强化学习方法,并系统分析轮次聚合评分策略、VLM 推理模式(权衡奖励偏差与方差)以及优势融合层级(防止奖励黑客行为)。研究发现将聚合优势广播至整个编辑轨迹能弥合局部规划与全局多轮任务成功之间的差距。实验表明,MT-EditFlow 在多种基模型上显著提升性能,其中 FLUX.1-Kontext-dev 的 turn-3 整体性能提升 6.85 分,超越 Qwen-Image-Edit 等开源模型。

AuthorsJiahui Huang*, Yasi Zhang†*, Tianyu Chen‡, Shu Wang, Jianwen Xie§, Oscar Leong†, Mingyuan Zhou‡, Nanzhu Wang, Ying Nian Wu†

Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-turn editing—the natural interactive setting where a user iteratively refines an image based on the model’s own previous outputs. This failure stems from the all-or-nothing requirement, where a single failed turn compromises the entire sequence, and error propagation, where exposure bias leads to compounding editing errors. To address these challenges, we introduce MT-EditFlow, a flow-matching reinforcement learning framework designed to optimize reward signals for sequential image editing. MT-EditFlow integrates a multi-turn perspective with a multi-reward formulation to provide a unified structure applicable to both GRPO and NFT-based reinforcement learning methods. We systematically analyze and optimize the reward signal by investigating effective scoring strategies for turn-level aggregation, VLM reasoning modes to trade off reward bias and variance, and advantage fusion levels to prevent reward hacking. Our findings reveal that broadcasting the aggregated advantage across the entire editing trajectory effectively bridges the gap between local planning and global multi-turn task success. Extensive experiments demonstrate that MT-EditFlow significantly improves performance across diverse base models. Notably, it boosts FLUX.1-Kontext-dev by 6.85 points in turn-3 overall performance, surpassing state-of-the-art open-source models such as Qwen-Image-Edit. By maintaining high marginal success rates and reducing exposure bias, MT-EditFlow provides a foundation for more reliable and natural human-AI collaboration in visual content creation.

  • † University of California, Los Angeles
  • ‡ University of Texas at Austin
  • § Lambda, Inc
  • * Equal contribution

Related readings and updates.

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

December 16, 2025research area Computer Visionconference CVPR

We present UniGen-1.5, a unified multimodal large language model (MLLM) for advanced image understanding, generation and editing. Building upon UniGen, we comprehensively enhance the model architecture and training pipeline to strengthen the image understanding and generation capabilities while unlocking strong image editing ability. Especially, we propose a unified Reinforcement Learning (RL) strategy that improves both image generation and…

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

October 27, 2025research area Computer Visionconference CVPR

Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community’s progress remains constrained by the absence of large-scale, high-quality, and openly accessible datasets built from real images. We introduce Pico-Banana-400K, a comprehensive 400K-image dataset for instruction-based image editing. Our dataset…

Bottom banner

Discover opportunities in Machine Learning.

来源:Apple Machine Learning Research(RSS)· machinelearning.apple.com