Pre-training Data Curation — Common Crawl, The Pile, Proprietary Mix
As of 2026-10-09
What is Pre-training Data Curation?
The Chinchilla scaling law's 20-tokens-per-parameter rule means a 70B model needs 1.4T+ training tokens, and frontier models now require 10-15T — making the five-stage curation pipeline the real product behind any frontier model.
What it is
Pre-training data curation is the foundation phase of building any large language model, and arguably the most consequential yet least visible one. Before a single training step runs, teams spend months assembling, filtering, deduplicating, and rebalancing the trillions of tokens that will permanently shape a model's vocabulary, reasoning style, factual grounding, and blind spots. Everything downstream — instruction tuning, RLHF, deployment behaviour — inherits the corpus's biases and gaps. For anyone advising a client on which foundation model to adopt inside Joule or a BDC-grounded agent, understanding what went into the corpus is not academic trivia; it is the difference between an informed model-selection decision and a marketing-driven one.
Why it matters
- Common Crawl provides 250B+ raw pages but is unusable raw — CCNet, C4 and RefinedWeb pipelines strip 90-95% of raw bytes before it's trainable.
- Deduplication via MinHash + LSH typically removes 30-50% of the raw corpus at document and paragraph level, and decontamination explicitly strips benchmark test sets (MMLU, HumanEval, GSM8K) to prevent eval leakage.
- Domain rebalancing deliberately under-weights raw web text and over-weights books/code/papers — the Chinchilla mix itself was 67% web plus 33% high-quality sources — to lift reasoning quality.
Key points
- Chinchilla scaling — compute-optimal ≈ 20 tokens/parameter; 70B model needs 1.4T+ tokens, frontier now 10T-15T.
- Three sources — Common Crawl (open web, 90-95% filtered out) + The Pile / RedPajama / FineWeb (curated research mix) + proprietary (licensed books, code, synthetic).
- Five-stage pipeline — quality filter, deduplication (MinHash+LSH, removes 30-50%), toxicity/PII, domain rebalance, benchmark decontamination.
- Domain mix decides downstream capability — code-heavy mixes lift reasoning; SAP-light mixes guarantee SAP-domain hallucination.
- Anti-pattern — assuming an LLM 'knows' proprietary or recent content; if it wasn't in the corpus before the training cutoff, it isn't there.
- SAP Domain Models (SAP-ABAP-2, S/4HANA/Ariba variants) solve the SAP-vocabulary gap via continued training on SAP's own corpus, not a from-scratch pre-training run — Early Adopter Care today, GA targeted Q3 2026 (past due: no GA announcement found as of 9 October 2026), not yet a committed delivery item.
- SAP-RPT and TabPFN sidestep this entire discussion — they predict via in-context learning from a supplied table at inference time, with no pre-training-corpus assembly step at all; don't apply this card's framework to that product family.
Terms used on this page
- Chinchilla scaling law
- Empirical result from DeepMind (Hoffmann et al. 2022) showing that compute-optimal LLM training requires roughly 20 training tokens per model parameter; reshaped the field away from over-parameterised under-trained models.
- Common Crawl
- Non-profit open web scrape published monthly since 2008; 250B+ pages of raw HTML; the largest single source in almost every LLM pre-training corpus, but requires 90-95% filtering before use.
- The Pile
- 825GB open research corpus released by EleutherAI in 2020 mixing Common Crawl, books, PubMed, ArXiv, GitHub, StackExchange, Wikipedia in declared proportions; superseded by SlimPajama, RedPajama-v2, FineWeb in 2024-2025.
- Decontamination
- The pipeline stage that explicitly removes benchmark evaluation test sets (MMLU, HumanEval, GSM8K) from the training corpus to prevent eval leakage and inflated reported scores.
- RedPajama-V2
- Together Computer's October 2023 successor corpus: 100B+ documents from 84 Common Crawl snapshots, a deduplicated head_middle partition of roughly 20.8B documents (~30.4T tokens across five languages), with published quality-signal annotations (word count, repetitiveness, toxicity) enabling customisable filtering (huggingface.co, checked 2026-09-27).
- SAP Domain Models corpus
- SAP's SAP-specific continued-training corpus (ABAP code, S/4HANA documentation and process content, Ariba data) layered onto an existing foundation model to produce SAP-ABAP-2 and S/4HANA/Ariba-specialised variants running under Joule/Joule Studio; Early Adopter Care today, GA targeted Q3 2026 — a domain-dense mix, not a from-scratch Common-Crawl-scale pre-training run.
Sources
- Hoffmann et al. — Training Compute-Optimal Large Language Models (Chinchilla, DeepMind 2022)
- Gao et al. — The Pile: An 800GB Dataset of Diverse Text for Language Modeling (EleutherAI 2020)
- Penedo et al. — The RefinedWeb Dataset for Falcon LLM (2023)
- HuggingFace — FineWeb dataset card and curation methodology
- SAP Business AI — official product page
- SAP Generative AI — official product page
- Hugging Face — RedPajama-Data-V2 dataset card, size and filtering methodology (checked 2026-09-27)
- SAP News Center — Joule Studio, enterprise-scale agentic development (SAP Domain Models context, 2026-05-13)
- SAP Learning — reference architecture, SAP Domain Models under Joule/Joule Studio
- SAP — SAP-RPT tabular foundation model product page (contrast with pre-training-corpus discussion)
- Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research — Soldaini et al., arXiv
- DataComp-LM: In search of the next generation of training sets for language models — Li et al., arXiv
- Deduplicating Training Data Makes Language Models Better — Lee et al., arXiv
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (C4) — Raffel et al., arXiv
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale — Penedo et al., arXiv
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.