Sliding Window Attention — Mistral's Long-Context Trick
As of 2026-10-04
What is Sliding Window Attention?
Mistral 7B claims 32K effective context by restricting each token's attention to a 4K window and letting the effective receptive field compound across 32 stacked layers instead of paying the O(n^2) cost of full attention.
What it is
Sliding Window Attention (SWA) restricts each token's attention to a fixed-size window of nearby tokens — typically 4K — instead of attending to the entire sequence. By stacking many such layers, information still propagates across long distances (each layer expands the effective receptive field), but the per-layer cost drops from O(n²) to O(n·w) where w is the window size. The architecture was popularised in research by Longformer (2020) and BigBird (2020) and brought to production scale by Mistral 7B (2023), which combined SWA with a rolling KV cache to claim 32K effective context on hardware that previously could not afford it.
Why it exists. Full attention's quadratic cost dominates GPU memory and compute past ~8K tokens; FlashAttention (C214) helps with the constant factor but not the asymptotic. For applications where most relevant information sits within a few thousand tokens of any given position (long documents, codebases, multi-turn chat), full attention is wasteful. SWA exploits this locality directly.
Why it matters
- Per-layer cost drops from O(n^2) to O(n*w), and a rolling KV cache lets memory per request stay O(w) instead of O(n) by evicting tokens outside the window.
- A 32-layer model with w=4096 reaches a roughly 128K theoretical receptive field purely from stacking sliding windows.
- Pure SWA can lose information that needed more hops than the model's depth allows, which is why Mistral Large and Mixtral 8x22B mix SWA with periodic full-attention layers.
Key points
- Each token attends only to the w previous tokens (causal sliding window) — O(n·w) cost per layer vs O(n²) for full attention.
- Effective receptive field grows to L · w across L stacked layers; a 32-layer Mistral 7B (w=4096) reaches ~128K theoretical reach.
- Mistral 7B uses pure SWA-4096 + rolling KV cache — capped per-request KV memory and 32K advertised context.
- Mistral Large, Mixtral, Gemini 1.5 and Claude use hybrid SWA + periodic full-attention layers to recover long-range fidelity.
- Predecessors — Longformer (2020) and BigBird (2020) — established the pattern in research; Mistral 7B (2023) was the production breakthrough.
- Nothing in Joule, Joule Studio or the generative AI hub exposes a sliding-window size to configure — SWA is entirely a property of whichever open-weight checkpoint a team self-hosts on SAP AI Core.
- The locality assumption behind SWA breaks specifically on densely cross-referenced enterprise documents — a purchase order pointing to a contract clause pointing to a master-data record — the exact structure of many SAP business documents.
Terms used on this page
- Sliding Window Attention (SWA)
- An attention pattern in which each token attends to only the w nearest previous tokens, reducing per-layer cost from O(n²) to O(n·w). Information propagates across longer distances by stacking layers.
- Rolling KV cache
- An inference-time optimisation that evicts Key/Value entries for tokens outside the current sliding window, capping per-request KV memory at O(w) regardless of total sequence length. Mistral 7B's signature trick.
- Effective receptive field
- The maximum distance over which information can propagate from input to output through stacked SWA layers, equal to L · w for L layers with window w. Distinct from advertised 'context window'.
- Hybrid attention
- An architecture interleaving SWA layers with periodic full-attention layers (e.g. every 4th or every 8th). Used by Mistral Large, Mixtral, Gemini 1.5 and Claude to recover the long-range fidelity that pure SWA can lose.
- Locality assumption
- The working bet behind every sparse-attention pattern, SWA included, that tokens near each other in the sequence matter more to a given prediction than distant tokens. Holds well for natural prose; breaks down on documents built around dense cross-referencing, such as many SAP business documents.
Sources
- Beltagy et al. — Longformer: The Long-Document Transformer (2020)
- Jiang et al. — Mistral 7B (2023) — Sliding Window Attention and Rolling Buffer Cache
- Mistral's Model Lets You Vibe Long-Running Code in the Cloud — AI Business, 2026-05-12
- Mistral AI — Announcing Mistral 7B (2023): confirms sliding window attention w=4096 + GQA
- arXiv — Big Bird: Transformers for Longer Sequences (Zaheer et al., 2020)
- Mistral AI — Mixtral of Experts announcement (2023)
- SAP Help Portal — Generative AI hub in SAP AI Core (hosted model endpoints)
- SAP Help Portal — Metering and pricing for generative AI on SAP AI Core (AI Units)
- SAP Community — Why SAP needs a Knowledge Graph: giving enterprise AI a map of the business (2026)
- Databricks — Model Serving documentation (self-hosting open-weight checkpoints)
- arXiv — FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (Dao et al., 2022)
- Efficient Streaming Language Models with Attention Sinks — Xiao et al., arXiv
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — Dao, arXiv
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Ainslie et al., arXiv
- Lost in the Middle: How Language Models Use Long Contexts — Liu et al., arXiv
- Mistral model documentation — Hugging Face Transformers
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.