Long-Context Training — From 8K to 1M+ Tokens
As of 2026-09-27
What is Long-Context Training?
Extending context from 8K to 1M+ tokens is three coupled engineering problems, not one — position-encoding extrapolation (YaRN), O(N²) attention cost (FlashAttention-3, Ring Attention), and training-data construction that teaches reasoning, not just memorization.
Long-context training is the discrete engineering effort that takes a language model built and trained at a modest context length — typically 8,000 or 32,000 tokens — and extends its usable window to the 200,000-to-2-million-token ranges that frontier models now advertise. It is not a single trick but three coupled problems that must be solved together: how position information survives far beyond its training range, how attention computation stays tractable at quadratic cost, and how training data is constructed so the model actually learns to reason over long spans rather than merely memorising them.
Position encoding: the first bottleneck
Rotary Position Embeddings, the dominant scheme in modern decoder-only models, encode position by rotating the query and key vectors by an angle proportional to token position before the attention dot product. The rotation frequency sets the range over which position information stays coherent, and naively running a RoPE-based model past its trained length degrades output quality sharply because the model is extrapolating into rotation angles it never saw during training. Three techniques address this. Position interpolation compresses the position index so that a longer sequence maps back into the frequency range the model was trained on. NTK-aware scaling adjusts the rotation base to preserve the high-frequency components that carry fine-grained positional detail. YaRN combines NTK-style scaling with a temperature correction to attention itself, and it has become the standard 2026 recipe for stretching an 8,000-token base model to 128,000 tokens and beyond with only a few hundred fine-tuning steps rather than a full retrain.
Why it matters
- Naive RoPE extrapolation degrades sharply past training length — YaRN (NTK scaling + attention temperature correction) is the 2026 standard for stretching 8K base to 128K+ in a few hundred fine-tuning steps.
- Ring Attention shards the sequence across GPUs in a ring topology, which is specifically how 1M+ context training becomes tractable on a 64-GPU cluster instead of requiring petabytes of attention memory.
- Random document concatenation for training data produces a model that memorizes long sequences but can't reason over them — curriculum data (book-length texts, multi-document QA chains) is what actually builds usable long-context capability.
Key points
- Three coupled components — position-encoding extrapolation (YaRN, NTK, PI), attention compute (FlashAttention-3, Ring Attention), training-data curriculum.
- RoPE extension — YaRN is 2026 standard; takes a 8K-base model to 128K+ with a few hundred fine-tuning steps at minimal quality loss.
- Attention scaling — full self-attention is O(N²); Ring Attention shards sequence across GPUs and rotates KV blocks to enable 1M+ context training.
- Data curriculum — synthesise long-context training data from book-length texts, code repos, multi-doc QA; ramp 8K → 32K → 128K during fine-tune phase.
- Evaluation — needle-in-haystack (NIAH) and RULER benchmarks; long-context accuracy degrades non-monotonically — test at every context length you plan to use.
- SAP's orchestration service implements 'grounding' as a managed RAG pipeline (document/DB grounding), not a long-context prompt-stuffing pattern — precisely so SAP Data Access Controls can be enforced per retrieved fragment, which a giant single prompt cannot support.
- Current Claude models, reachable from Joule via the SAP-Anthropic MCP integration, ship up to 1M-token context windows (Sonnet 5, Opus 5.5/5.5; Haiku 4.5 at 200K) — real headroom for holistic single-document review, not a reason to abandon retrieval for governed enterprise corpora.
- A bigger advertised window does not remove the need to test recall at the length actually used — needle-in-the-middle degradation is a property of the attention mechanism, not of any one vendor's marketing number.
Terms used on this page
- RoPE (Rotary Position Embedding)
- Position encoding scheme that rotates query and key vectors by an angle proportional to their position; dominant in 2026 LLMs because it extrapolates better than absolute or sinusoidal encodings.
- YaRN
- Yet another RoPE extensioN; a 2023-introduced technique combining NTK-aware frequency scaling with attention temperature correction; the 2026 standard for extending pretrained context windows by 16-32x.
- Ring Attention
- Distributed attention algorithm (Liu et al. 2023) that shards the sequence across N GPUs and rotates key-value blocks in a ring topology so each GPU sees the full context across N communication rounds; enables 1M+ token training.
- Needle-in-Haystack (NIAH)
- Long-context evaluation that inserts a target fact at a random position in a long distractor document and measures whether the model retrieves it; the basic sanity check for long-context capability.
- Position interpolation
- A RoPE-extension technique that compresses the position index of a longer sequence so it maps back into the rotation-frequency range the model was originally trained on; simpler than YaRN but degrades fine-grained positional detail more.
- NTK-aware scaling
- A RoPE-extension technique that adjusts the rotation base to preserve high-frequency (fine-grained) positional components while still extending the usable range; one of the two ingredients YaRN combines with an attention-temperature correction.
- Attention thinning / lost-in-the-middle
- The empirically observed drop in retrieval accuracy for facts placed in the middle of a very long prompt, even when the same facts near the start or end are retrieved reliably (Liu et al., 2023); the reason a large advertised context window does not guarantee uniform recall across it.
Sources
- arXiv: YaRN — Efficient Context Window Extension of Large Language Models (Peng et al. 2023)
- arXiv: Ring Attention with Blockwise Transformers for Near-Infinite Context (Liu et al. 2023)
- arXiv: RoFormer — Enhanced Transformer with Rotary Position Embedding (Su et al. 2021)
- Anthropic Engineering — long-context internals
- arXiv — Shah et al., "FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision" (2024, fetched 2026-09-27)
- arXiv — Hsieh et al., "RULER: What's the Real Context Size of Your Long-Context Language Models?" (2024, fetched 2026-09-27)
- arXiv — Liu et al., "Lost in the Middle: How Language Models Use Long Contexts" (TACL 2023, fetched 2026-09-27)
- arXiv — Gemini Team, Google, "Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context" (2024, fetched 2026-09-27)
- Claude Docs — Context windows (current model context-window limits), fetched 2026-09-27
- SAP Help Portal — Orchestration service (document/DB grounding) in the generative AI hub, SAP AI Core (fetched 2026-09-27)
- SAP Community — Why SAP needs a Knowledge Graph: giving enterprise AI a map of the business (2026, fetched 2026-09-27)
- SAP News Center — SAP and Anthropic to bring Claude to the SAP Business AI Platform (2026-05-12)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.