面壁智能开源 Ultra-FineWeb-L1 数据集

OpenBMB · @OpenBMB · X·2026-08-20 21:01·14天前
AI 导读

面壁智能开源 Ultra-FineWeb-L1,一个经精细处理的英文网页语料库,包含 1T+ tokens 和约 1.14B 文档,数据覆盖至 CC-MAIN-2025-51。该数据集作为 UltraData 生态的 L1 过滤层,结合文本提取、语言过滤、启发式过滤及去重等流程提升数据质量,已上线 HuggingFace。

OpenBMB@OpenBMB
62AI 编辑部评分,满分 100

面壁智能开源 Ultra-FineWeb-L1 数据集

2026-08-20 21:01· 14天前
AI 导读

面壁智能开源 Ultra-FineWeb-L1,一个经精细处理的英文网页语料库,包含 1T+ tokens 和约 1.14B 文档,数据覆盖至 CC-MAIN-2025-51。该数据集作为 UltraData 生态的 L1 过滤层,结合文本提取、语言过滤、启发式过滤及去重等流程提升数据质量,已上线 HuggingFace。

🔥 Ultra-FineWeb-L1 is here —— 1T+ tokens of high-quality web data are now available!

High-quality data is the foundation of powerful LLMs. Ultra-FineWeb-L1 is now open-sourced as part of the UltraData ecosystem, providing a carefully processed English web corpus from Common Crawl.

Built for Better Data Quality 1️⃣ Fresh Web Data: Collected from recent Common Crawl snapshots, covering data up to CC-MAIN-2025-51, with 1T+ tokens and ~1.14B documents. 2️⃣ Refined Processing Pipeline: Combines text extraction, language filtering, heuristic filtering, sensitive-field replacement, deduplication, and customized cleaning to improve data quality.

⚡ Highlights: • 1T+ tokens of filtered English web data • Enhanced cleaning for noisy web content, encoding issues, and abnormal documents • Designed as the L1 filtered layer in the UltraData management framework • Aligned with the latest Ultra-FineWeb updates http://huggingface.co/datasets/openbmb/Ultra-FineWeb

🔗 Resources 🤗 Dataset: http://huggingface.co/datasets/openbmb/Ultra-FineWeb-L1 📄 Paper: http://arxiv.org/abs/2505.05427 🧩 Classifier: http://ultradata.openbmb.cn/