E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scenarios. It maps image, video, audio, and text inputs into a unified representation space. For long-form live streaming analysis, we introduce Per-vGrid, a timestamped token organization that groups each video grid with its temporally corresponding audio within explicit boundary tokens to facilitate temporal alignment. We design a three-stage supervised training recipe that progressively develops live-commerce understanding, from omni-modal perception to instruction-following responses. We then propose Faithful-RFT, a reinforcement fine-tuning stage that further improves answer faithfulness and expression quality while meeting real-time demands, scoring final responses directly with task-verifiable feedback rather than optimizing for reasoning-style exploration during rollout. Moreover, TLive-Omni is supported by a scenario-oriented atomic capability taxonomy and a compact data production engine that converts live-commerce audio, image, and video streams into training signals for speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc. For scalable training, a synchronized length-grouped sampler reduces padding while preserving comparable workloads across workers, while a lightweight dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful relative advantages for GRPO. Experiments on e-commerce live streaming benchmarks demonstrate strong performance across live-commerce domain tasks, together with excellent generalization on general benchmarks.
TLive-Omni:面向电商直播的全模态理解模型
AI 导读
TLive-Omni 将图像、视频、音频和文本输入映射到统一表示空间,专为电商直播场景设计。其引入 Per-vGrid 时间戳 token 组织以对齐视频与音频,并通过三阶段监督训练和 Faithful-RFT 强化微调提升回答忠实度与表达质量,同时满足实时性要求。实验表明,该模型在电商直播基准任务上表现强劲,并在通用基准上展现出良好泛化能力。
HuggingFace Daily Papers(社区热门论文)
39
AI 编辑部评分,满分 100TLive-Omni:面向电商直播的全模态理解模型
TLive-Omni 将图像、视频、音频和文本输入映射到统一表示空间,专为电商直播场景设计。其引入 Per-vGrid 时间戳 token 组织以对齐视频与音频,并通过三阶段监督训练和 Faithful-RFT 强化微调提升回答忠实度与表达质量,同时满足实时性要求。实验表明,该模型在电商直播基准任务上表现强劲,并在通用基准上展现出良好泛化能力。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org