AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

KV Cache

KV Cache — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-09-27

What is KV Cache?

The KV cache, not model weights, is the binding VRAM constraint on LLM inference — an 8,000-token context on a 70B model already consumes ~42GB, meaning doubling context length roughly doubles hardware cost.

The KV cache — key-value cache — is the memory structure that makes fast, affordable large language model inference possible, and it sits directly behind cost and latency decisions for any SAP Joule or AI Foundation deployment that runs at real user volume. Understanding it is what separates an architect who can explain why a longer context window is expensive from one who can only point at an invoice.

What it is and why it matters

When a transformer model generates a response one token at a time, each new token has to "attend" to every token that came before it — the whole conversation so far, including the system prompt and any retrieved grounding content. Recomputing the key and value projections for that entire history at every single step would make generation slower with every token, an effect that compounds badly on long enterprise prompts stuffed with SAP context. The KV cache avoids that: it stores the key and value projections for every token the moment they are first computed, so generating token number one thousand only requires processing that one new token, not recomputing the previous nine hundred and ninety-nine. This is the single mechanism that makes chat-style, multi-turn generation practical at all.

Why it matters

  • Without caching, generating token N would require an O(N²) recomputation explosion — the cache is what makes autoregressive generation at enterprise scale feasible at all.
  • Prefix/prompt caching is directly relevant to Joule: sessions prefixed with the same enterprise context block can reuse cached KV projections instead of recomputing them every call.
  • The concrete memory formula (2 × L × H × d × bytes) lets an architect size hardware (single A100-80GB vs. multi-GPU pod) before committing to a context-length requirement.

Key points

  • KV cache memory = 2 × n_layers × n_kv_heads × d_head × seq_len × batch × dtype_bytes.
  • PagedAttention (vLLM) raises GPU utilisation from ~40% to ~90% by virtualising the KV cache.
  • Prefix caching (RadixAttention) reuses KV representations of shared prompt prefixes across calls — critical for SAP Joule system prompts.
  • INT8 KV quantisation halves cache memory with ~1% quality loss when per-token dynamic quantisation is applied.
  • H2O eviction reduces cache by 20× for long generation tasks with <5% quality loss by keeping only high-attention tokens.
  • On a managed generative AI hub endpoint the KV cache itself is invisible to the customer — the only exposed lever is prompt/prefix caching (implicit for OpenAI and Gemini, explicit cache_control breakpoints for Anthropic Claude and Amazon Nova via orchestration v2); a self-hosted BYOM deployment on SAP AI Core (SAP's own sample repo documents vLLM, llama.cpp, Ollama and LocalAI paths) exposes the full PagedAttention and quantisation stack instead.
  • Cache size is a hard capacity ceiling before it is a cost curve: concurrency is capped the moment the sum of active KV caches exceeds GPU memory, regardless of how much compute is still idle — this is why a Joule-style deployment can look 'slow' while GPU utilisation graphs show headroom.
  • GQA-based models (Llama 3, Mistral) cut KV cache 4-8× versus classic multi-head attention at near-zero quality loss — an architecture choice baked into the model, not a runtime optimisation, so two similarly-sized models are not automatically equally expensive to serve at long context.

Terms used on this page

KV cache
Memory buffer storing key and value matrices for all past tokens; enables O(seq_len) per-step generation cost instead of O(seq_len²) full recomputation.
PagedAttention
Virtual memory-inspired KV cache allocation (Kwon 2023, vLLM) that stores cache in non-contiguous pages, eliminating fragmentation and raising GPU utilisation.
Prefix caching
Mechanism (RadixAttention, Zheng 2023) that indexes KV representations of shared prompt prefixes so repeated calls reuse the cached computation.
H2O
Heavy Hitter Oracle (Zhang 2023): KV cache eviction policy that retains the tokens with highest cumulative attention mass, reducing cache size 20× with <5% quality loss.
INT8 KV quantisation
Compression of KV cache entries from BF16 (2 bytes) to INT8 (1 byte), halving memory with ~1% quality loss when per-token dynamic scales are maintained.

Sources

  1. Kwon et al. 2023 — Efficient Memory Management for LLM Serving with PagedAttention (vLLM)
  2. Zheng et al. 2023 — Efficiently Programming Large Language Models using SGLang (RadixAttention)
  3. Zhang et al. 2023 — H2O: Heavy-Hitter Oracle for Efficient Generative Inference
  4. Shazeer 2019 — Fast Transformer Decoding: One Write-Head is All You Need
  5. NVIDIA H100 SXM Datasheet
  6. OpenAI prompt caching documentation
  7. Anthropic prompt caching documentation
  8. vLLM documentation — PagedAttention and prefix caching
  9. GitHub SAP-samples — btp-generative-ai-hub-use-cases: bring-your-own OSS LLM on SAP AI Core (Ollama, LocalAI, llama.cpp, vLLM, custom HF Transformers server)
  10. SAP News Center — How SAP and NVIDIA Advance Enterprise AI Transformation (NIM inference optimisation, SAP-ABAP-1, 2026-03-17)
  11. PyTorch — torch.nn.functional.scaled_dot_product_attention documentation
  12. SAP Help Portal — Orchestration service in SAP AI Core generative AI hub (prompt/document grounding pipeline where prefix caching sits)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →