AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

Pre-training Data Curation — Common Crawl, The Pile, Proprietary Mix

Pre-training Data Curation — Common Crawl, The Pile, Proprietary Mix — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-10-09

What is Pre-training Data Curation?

The Chinchilla scaling law's 20-tokens-per-parameter rule means a 70B model needs 1.4T+ training tokens, and frontier models now require 10-15T — making the five-stage curation pipeline the real product behind any frontier model.

What it is

Pre-training data curation is the foundation phase of building any large language model, and arguably the most consequential yet least visible one. Before a single training step runs, teams spend months assembling, filtering, deduplicating, and rebalancing the trillions of tokens that will permanently shape a model's vocabulary, reasoning style, factual grounding, and blind spots. Everything downstream — instruction tuning, RLHF, deployment behaviour — inherits the corpus's biases and gaps. For anyone advising a client on which foundation model to adopt inside Joule or a BDC-grounded agent, understanding what went into the corpus is not academic trivia; it is the difference between an informed model-selection decision and a marketing-driven one.

Why it matters

  • Common Crawl provides 250B+ raw pages but is unusable raw — CCNet, C4 and RefinedWeb pipelines strip 90-95% of raw bytes before it's trainable.
  • Deduplication via MinHash + LSH typically removes 30-50% of the raw corpus at document and paragraph level, and decontamination explicitly strips benchmark test sets (MMLU, HumanEval, GSM8K) to prevent eval leakage.
  • Domain rebalancing deliberately under-weights raw web text and over-weights books/code/papers — the Chinchilla mix itself was 67% web plus 33% high-quality sources — to lift reasoning quality.

Key points

  • Chinchilla scaling — compute-optimal ≈ 20 tokens/parameter; 70B model needs 1.4T+ tokens, frontier now 10T-15T.
  • Three sources — Common Crawl (open web, 90-95% filtered out) + The Pile / RedPajama / FineWeb (curated research mix) + proprietary (licensed books, code, synthetic).
  • Five-stage pipeline — quality filter, deduplication (MinHash+LSH, removes 30-50%), toxicity/PII, domain rebalance, benchmark decontamination.
  • Domain mix decides downstream capability — code-heavy mixes lift reasoning; SAP-light mixes guarantee SAP-domain hallucination.
  • Anti-pattern — assuming an LLM 'knows' proprietary or recent content; if it wasn't in the corpus before the training cutoff, it isn't there.
  • SAP Domain Models (SAP-ABAP-2, S/4HANA/Ariba variants) solve the SAP-vocabulary gap via continued training on SAP's own corpus, not a from-scratch pre-training run — Early Adopter Care today, GA targeted Q3 2026 (past due: no GA announcement found as of 9 October 2026), not yet a committed delivery item.
  • SAP-RPT and TabPFN sidestep this entire discussion — they predict via in-context learning from a supplied table at inference time, with no pre-training-corpus assembly step at all; don't apply this card's framework to that product family.

Terms used on this page

Chinchilla scaling law
Empirical result from DeepMind (Hoffmann et al. 2022) showing that compute-optimal LLM training requires roughly 20 training tokens per model parameter; reshaped the field away from over-parameterised under-trained models.
Common Crawl
Non-profit open web scrape published monthly since 2008; 250B+ pages of raw HTML; the largest single source in almost every LLM pre-training corpus, but requires 90-95% filtering before use.
The Pile
825GB open research corpus released by EleutherAI in 2020 mixing Common Crawl, books, PubMed, ArXiv, GitHub, StackExchange, Wikipedia in declared proportions; superseded by SlimPajama, RedPajama-v2, FineWeb in 2024-2025.
Decontamination
The pipeline stage that explicitly removes benchmark evaluation test sets (MMLU, HumanEval, GSM8K) from the training corpus to prevent eval leakage and inflated reported scores.
RedPajama-V2
Together Computer's October 2023 successor corpus: 100B+ documents from 84 Common Crawl snapshots, a deduplicated head_middle partition of roughly 20.8B documents (~30.4T tokens across five languages), with published quality-signal annotations (word count, repetitiveness, toxicity) enabling customisable filtering (huggingface.co, checked 2026-09-27).
SAP Domain Models corpus
SAP's SAP-specific continued-training corpus (ABAP code, S/4HANA documentation and process content, Ariba data) layered onto an existing foundation model to produce SAP-ABAP-2 and S/4HANA/Ariba-specialised variants running under Joule/Joule Studio; Early Adopter Care today, GA targeted Q3 2026 — a domain-dense mix, not a from-scratch Common-Crawl-scale pre-training run.

Sources

  1. Hoffmann et al. — Training Compute-Optimal Large Language Models (Chinchilla, DeepMind 2022)
  2. Gao et al. — The Pile: An 800GB Dataset of Diverse Text for Language Modeling (EleutherAI 2020)
  3. Penedo et al. — The RefinedWeb Dataset for Falcon LLM (2023)
  4. HuggingFace — FineWeb dataset card and curation methodology
  5. SAP Business AI — official product page
  6. SAP Generative AI — official product page
  7. Hugging Face — RedPajama-Data-V2 dataset card, size and filtering methodology (checked 2026-09-27)
  8. SAP News Center — Joule Studio, enterprise-scale agentic development (SAP Domain Models context, 2026-05-13)
  9. SAP Learning — reference architecture, SAP Domain Models under Joule/Joule Studio
  10. SAP — SAP-RPT tabular foundation model product page (contrast with pre-training-corpus discussion)
  11. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research — Soldaini et al., arXiv
  12. DataComp-LM: In search of the next generation of training sets for language models — Li et al., arXiv
  13. Deduplicating Training Data Makes Language Models Better — Lee et al., arXiv
  14. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (C4) — Raffel et al., arXiv
  15. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale — Penedo et al., arXiv

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →