🔥 Ultra-FineWeb-L1 is here —— 1T+ tokens of high-quality web data are now available!
High-quality data is the foundation of powerful LLMs. Ultra-FineWeb-L1 is now open-sourced as part of the UltraData ecosystem, providing a carefully processed English web corpus from Common Crawl.
Built for Better Data Quality 1️⃣ Fresh Web Data: Collected from recent Common Crawl snapshots, covering data up to CC-MAIN-2025-51, with 1T+ tokens and ~1.14B documents. 2️⃣ Refined Processing Pipeline: Combines text extraction, language filtering, heuristic filtering, sensitive-field replacement, deduplication, and customized cleaning to improve data quality.
⚡ Highlights: • 1T+ tokens of filtered English web data • Enhanced cleaning for noisy web content, encoding issues, and abnormal documents • Designed as the L1 filtered layer in the UltraData management framework • Aligned with the latest Ultra-FineWeb updates http://huggingface.co/datasets/openbmb/Ultra-FineWeb
🔗 Resources 🤗 Dataset: http://huggingface.co/datasets/openbmb/Ultra-FineWeb-L1 📄 Paper: http://arxiv.org/abs/2505.05427 🧩 Classifier: http://ultradata.openbmb.cn/