AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

KV Cache Optimization — PagedAttention, Quantized KV, Speculative Cache

KV Cache Optimization — PagedAttention, Quantized KV, Speculative Cache — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-10-05

What is KV Cache Optimization?

At 32K context, Llama 3 70B's KV cache runs roughly 10-11 GB per request in FP16 — a fraction of the INT8 weights for one user, but the term that scales with concurrency — making KV-cache optimisation the highest-leverage lever in LLM inference economics.

What it is

During autoregressive LLM inference, every previously generated token's key and value tensors must be kept in GPU memory so future tokens can attend to them — the KV cache. At long context lengths the KV cache dominates total memory usage: for Llama 3 70B (80 layers, 8 GQA key-value heads, head dimension 128) at 32K context the KV cache is roughly 10-11 GB per request in FP16 — a fraction of the ~70 GB INT8 weights for one request, but a few dozen concurrent long-context requests exceed the weights, which is why the KV cache, not the weights, sets the concurrency ceiling. KV cache optimisation is therefore the single highest-leverage area for LLM inference economics — every technique here directly translates into either more concurrent users on the same GPU or longer context at the same cost.

Why it matters

  • PagedAttention eliminates memory fragmentation and enables up to 24x higher throughput by safely overlapping multiple sequences — the single biggest inference-serving advance of 2023-2024.
  • KV quantization to INT8 halves cache memory for under 1% quality loss; INT4 quarters it for 2-3% loss.
  • Prefix caching reuses a shared prompt prefix's KV state across requests, often delivering a 5-10x throughput win for chatbot workloads with a common system prompt.

Key points

  • KV cache (per-token K + V tensors held in GPU memory for future attention) dominates long-context memory — Llama 3 70B at 32K context ≈ 10-11 GB KV cache per request in FP16 (80 layers × 8 GQA KV heads × 128 dims) — the term that scales with concurrent users.
  • PagedAttention (vLLM, 2023) — page-table-style KV allocation eliminates fragmentation, 24× throughput improvement; now standard in vLLM, TensorRT-LLM, SGLang.
  • KV quantisation — INT8 halves memory at <1% quality loss, INT4 quarters at ~2-3% loss; standard in modern serving stacks.
  • MQA (single K/V across all heads) and GQA (grouped heads share K/V, used in Llama 2/3, Mistral, Mixtral) shrink KV cache 4-8× at architecture level.
  • Prefix / speculative caching — cache shared prompt prefix KV once, reuse across requests; 5-10× throughput win for chatbot workloads with shared system prompts.
  • The raw concurrency ceiling for a self-hosted deployment is arithmetic, not a serving-stack feature: (GPU memory − model weights) ÷ per-request KV-cache size at the target context length; PagedAttention improves how close real utilisation gets to that ceiling, it does not move the ceiling itself.
  • Prefix caching and KV quantisation are not simply additive — a cached prefix's KV entries are locked to whatever precision they were cached at, and a precision mismatch on a later request forces a full recompute instead of a cache hit.
  • On a managed generative AI hub endpoint none of this is a customer-facing lever; it becomes the consultant's own sizing decision only on a self-hosted SAP AI Core BYOM deployment (vLLM, per SAP's own sample repository), where the GPU resource plan chosen from SAP's catalogue sets the hard ceiling above.

Terms used on this page

KV cache
The cached key and value tensors of all previously generated tokens, held in GPU memory so future tokens can attend to them during autoregressive inference; dominates memory consumption at long context lengths.
PagedAttention
An inference-serving technique introduced by vLLM (Kwon et al., 2023) that allocates the KV cache in fixed-size blocks like OS virtual-memory pages, eliminating fragmentation and enabling 24× higher throughput than naive contiguous allocation.
Grouped-Query Attention (GQA)
An attention architecture variant where groups of query heads share a single set of K/V tensors (e.g. 8 K/V groups for 32 query heads); used in Llama 2/3, Mistral, Mixtral and gives 4-8× KV cache reduction at near-zero quality loss.
Multi-Query Attention (MQA)
An extreme attention variant where all query heads share a single K and V tensor; cuts KV cache by H (number of heads, typically 32-128) but with measurable quality loss compared to GQA.
Prefix caching
An inference optimisation that stores the KV cache for a shared prompt prefix once and reuses it across all requests that share that prefix; canonical in chatbot workloads with shared system prompts; 5-10× throughput win typical.

Sources

  1. Efficient Memory Management for Large Language Model Serving with PagedAttention (Kwon et al., 2023)
  2. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints (Ainslie et al., 2023)
  3. Fast Transformer Decoding: One Write-Head is All You Need (Shazeer, 2019, MQA)
  4. vLLM documentation — PagedAttention and prefix caching
  5. NVIDIA TensorRT-LLM — KV cache management
  6. GitHub SAP-samples — btp-generative-ai-hub-use-cases: bring-your-own OSS LLM on SAP AI Core (vLLM serving backend)
  7. SAP Help Portal — Choose a Resource Plan for training/inference in SAP AI Core (GPU sizing catalogue, incl. flexible instance types per SAP Note 3660109)
  8. SAP News Center — How SAP and NVIDIA Advance Enterprise AI Transformation (NIM inference optimisation context)
  9. GitHub SGLang — sgl-project/sglang (RadixAttention prefix-cache implementation)
  10. GitHub vllm-project — vLLM (PagedAttention reference implementation and KV-cache quantisation support)
  11. SAP Help Portal — What is SAP AI Core (platform context for a self-hosted BYOM deployment)
  12. GitHub SAP-samples — btp-gen-ai-hub-sdk-samples: setting up the Extended AI Core service plan required for self-hosted serving
  13. Hugging Face — GQA (Grouped-Query Attention) explainer and model configuration reference
  14. vLLM docs — Quantized KV Cache (FP8 KV-cache dtypes fp8_e4m3/fp8_e5m2 supported in vLLM)
  15. vLLM docs — Automatic Prefix Caching (reuse of shared-prefix KV blocks across requests)
  16. vLLM blog — vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention (original launch, up to 24x vs Hugging Face Transformers; project/vendor source)
  17. Liu et al. — KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache (low-bit KV quantization research)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →