# DataOrchestra：逐样本编排预训练数据清洗

- 来源：Dongxi 东锡 NLP (@dongxi_nlp)
- 发布时间：2026-08-05 22:51
- AIHOT 分数：32
- AIHOT 链接：https://aihot.virxact.com/items/cmsg7n5e700odrorep7ye9t2j
- 原文链接：https://x.com/dongxi_nlp/status/2085015920544534682

## AI 摘要

DataOrchestra 是 GAIR NLP 推出的开源框架，将预训练数据清洗转为逐样本决策：1.7B 轻量编排器对每个数据块决定 Drop、Untouch 或 Clean，并只为需清洗块选择必要操作。基于执行反馈生成 300K 训练对，在 11 项基准上从 0.5B 到 7B 规模一致优于固定与单独处理策略，优势随模型增大而扩大。约 35% 数据块跳过清洗，节省约 37% 清洗算力。

## 正文

DataOrchestra:

Learning to Orchestrate Per-Example Curation of Pretraining Data

### 引用推文

> Yikun Wang ✈️ ICML 2026：Happy to introduce DataOrchestra by @Z_Huang_02 and GAIR NLP, and feel thrilled to participate!! Most pretraining pipelines ask: Which processing recipe should ...
