Continual Pre-training — Domain Adaptation Without Forgetting
As of 2026-10-09
What is Continual Pre-training?
Continual pre-training injects billions of domain tokens into a foundation model to internalize its dialect before task fine-tuning — but risks catastrophic forgetting unless replay buffers, low learning rates, and loss masking anchor it.
What it is
Continual pre-training (CPT) is the procedure of taking a fully-pretrained foundation model — Llama-3, Mistral, Qwen, or a frontier closed model exposed through fine-tuning APIs — and feeding it billions more tokens drawn from a target domain, so that the model internalises the vocabulary, idioms, and reasoning patterns of that domain before any task-specific fine-tuning happens. For SAP analytics, the target corpus is a curated mix of SAP Help Portal, OSS notes, ABAP and HANA SQL code, SAP Community threads, and customer-implementation documentation.
The central risk is catastrophic forgetting — the phenomenon where new training shifts the weights so far that the model loses general capabilities it had at the start (math, multi-lingual reasoning, instruction following). Three guardrails dominate the literature. Replay buffers mix 5-15% of the original pretraining distribution back into each CPT batch so the general distribution stays anchored. Low learning rates (1e-5 to 5e-5 versus the 1e-4 of from-scratch pretraining) keep weight updates small. Loss masking on system tokens and reserved instruction markers prevents the model from over-fitting to formatting artefacts.
The data pipeline is where most CPT projects fail. Deduplication using MinHash or SemDeDup is non-negotiable — repeated content drives memorisation rather than generalisation. Quality filtering with perplexity scoring against a known-good model removes machine-translated junk. PII scrubbing and EU-AI-Act-Art-10-aligned provenance logging are mandatory for any model that will be exposed to end users.
Why it matters
- Three concrete guardrails prevent forgetting: 5-15% replay of original pretraining data per batch, learning rates 1e-5 to 5e-5 (vs 1e-4 from-scratch), and loss masking on system/instruction tokens.
- Most CPT projects fail at the data pipeline, not the training loop — deduplication (MinHash/SemDeDup) and perplexity-based quality filtering are non-negotiable.
- The decision rule is explicit: reach for CPT with 10B+ tokens of domain dialect (SAP ABAP, legal, biomedical); reach for LoRA (C258) instead with only 10K-10M task examples.
Key points
- Procedure — feed 10B-100B tokens of domain text into a pretrained model at a low learning rate (1e-5 to 5e-5) before any task-level fine-tuning.
- Catastrophic forgetting — the principal risk; mitigated by replay buffers (5-15% original distribution), low LR, and loss masking on system tokens.
- Data pipeline — MinHash/SemDeDup deduplication, perplexity-based quality filtering, PII scrubbing, EU AI Act Art. 10 provenance logging are mandatory.
- When to use — domain has its own dialect (ABAP, legal, biomed) AND you have 10B+ tokens; otherwise use LoRA fine-tuning (C258) or RAG (C203).
- Evaluation — hold out a domain benchmark AND a general benchmark (MMLU, GSM8K); track both every checkpoint to detect forgetting in real time.
- SAP-specific caution — SAP Domain Models (SAP-ABAP-2 and successors) exist and run under Joule/Joule Studio, but SAP has not published that they are built via continual pre-training specifically; treat the training method as undocumented, not as CPT by default.
- SAP's generative AI hub defaults to orchestration (grounding) and LoRA-style fine-tuning before anything resembling CPT — reach for CPT only once token count and dialect breadth cross the thresholds in the decision table.
- Measure post-deduplication token count before scoping the technique — a corpus that looks large in raw bytes can fall well under the 10B-token threshold once boilerplate and near-duplicates are removed.
Terms used on this page
- Catastrophic forgetting
- The phenomenon where new training shifts model weights enough that previously-learned general capabilities (math, multi-lingual, instruction following) degrade; the central failure mode of continual pre-training.
- Replay buffer
- A pool of original pretraining-distribution samples mixed into every new training batch (typically 5-15%) to anchor the general distribution and prevent forgetting.
- Perplexity filtering
- Removing low-quality text from a training corpus by scoring each document with a known-good model and dropping the top-decile-perplexity rows (machine-translated junk, OCR noise, broken HTML).
- SemDeDup
- Semantic deduplication (Abbas et al., 2023) using sentence-embedding cosine similarity to remove near-duplicates that simple hash-based deduplication misses; reduces memorisation and improves generalisation.
- Elastic Weight Consolidation (EWC)
- Kirkpatrick et al.'s 2017 technique that slows learning on weights identified as important to previously-learned tasks, an early and influential approach to mitigating catastrophic forgetting that predates the replay-buffer approach dominant in 2026 LLM practice.
- Domain-adaptive pretraining (DAPT)
- Gururangan et al.'s term for a lighter-weight alternative to full CPT: a single additional phase of in-domain pretraining, shown to produce measurable gains even on modest domain corpora, often the right choice when the target corpus falls short of CPT's multi-billion-token threshold.
- Learning-rate re-warming / re-decaying
- The practice of temporarily raising the learning rate back toward its original schedule when starting a continual pretraining run on new data, then decaying it again — shown to let a continually pretrained model match from-scratch retraining quality at a fraction of the compute.
Sources
- arXiv: Continual Pre-Training of Large Language Models — How to (re)warm your model? (Ibrahim et al. 2024)
- arXiv: Simple and Scalable Strategies to Continually Pre-train Large Language Models (Gupta et al. 2024)
- Anthropic Responsible Scaling Policy — model training governance
- NIST AI Risk Management Framework 1.0
- SAP Generative AI — official product page
- arXiv — Abbas et al., "SemDeDup: Data-efficient learning at web-scale through semantic deduplication" (2023, fetched 2026-09-27)
- arXiv — Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021, fetched 2026-09-27)
- arXiv — Gururangan et al., "Don't Stop Pretraining: Adapt Language Models to Domains and Tasks" (ACL 2020, fetched 2026-09-27)
- arXiv — Kirkpatrick et al., "Overcoming catastrophic forgetting in neural networks" — the Elastic Weight Consolidation paper (2017, fetched 2026-09-27)
- SAP Help Portal — Orchestration service in the generative AI hub, SAP AI Core (fetched 2026-09-27)
- SAP Help Portal — Metering and pricing for generative AI, SAP AI Core service guide (fetched 2026-09-27)
- EUR-Lex — Regulation (EU) 2024/1689 (EU AI Act), consolidated text incl. Art. 10 data governance (fetched 2026-09-27)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.