Multi-head Attention vs Grouped-Query Attention vs Multi-query Attention
As of 2026-09-27
Multi-head Attention vs Grouped-Query Attention vs Multi-query Attention: what is the difference?
Switching from MHA to GQA-8 at equal model size cuts KV-cache memory roughly 8x for under 1% MMLU drop, doubling or tripling how many concurrent users one H100 can serve — the real reason Mistral and Llama 3 serve cheaper than equivalent MHA models.
What it is
Multi-head Attention (MHA), Grouped-Query Attention (GQA) and Multi-Query Attention (MQA) are three points on the same trade-off curve: how many distinct Key/Value projections does each layer compute, given a fixed number of Query heads. The choice does not change what a model can express; it changes how much KV-cache memory inference burns per token, and therefore the per-request cost and the maximum concurrency a given GPU can serve.
Why it exists. The original Transformer (2017) used MHA — every Query head has its own Key and Value head. That doubles representational diversity but also doubles KV-cache memory at inference, which is the dominant memory cost at long context (companion C160). MQA (Shazeer 2019) collapses all heads to a single shared K/V — minimal KV memory, but a measurable quality drop at scale. GQA (Ainslie et al. 2023) is the practical compromise that every frontier model since Llama 2 has adopted.
How they work. With n q Query heads and n kv Key/Value heads: MHA uses n kv = n q (e.g. Llama 1 70B: 64 Q, 64 KV). MQA uses n kv = 1 (one shared K/V across all Q heads). GQA groups Q heads into g groups that share K/V (e.g. Llama 3 70B: 64 Q, 8 KV — 8 Q heads per shared KV pair). Llama 3/4, Mistral, Mixtral, Gemma 2, and Claude (per inferred architecture) all use GQA. GPT-4 family uses MHA per public reverse-engineering analyses; GPT-5 details remain undisclosed.
Why it matters
- MHA (Llama 1 70B: 64 Q, 64 KV) maximises representational diversity but doubles KV-cache memory versus alternatives; MQA collapses to one shared K/V head with a measurable quality drop at scale.
- GQA (Llama 3 70B: 64 Q, 8 KV) is the practical compromise every frontier model since Llama 2 has adopted.
- If you self-host or fine-tune via SAP AI Core, this single architectural choice controls sustainable concurrency per GPU — invisible to a hosted-API consumer, but essential to know why some models cost less to serve.
Key points
- Three points on one curve — MHA (n_kv = n_q), GQA (n_kv = n_q / g), MQA (n_kv = 1) — trading KV memory for quality.
- GQA-8 (8 Q heads per KV pair) is the default for Llama 3/4, Mistral 7B/Large, Mixtral 8x7B, Gemma 2 — ~8× KV memory reduction vs MHA, <1% MMLU loss.
- MQA (n_kv = 1) maximises memory savings but degrades multi-task quality; rarely used in frontier models post-2023.
- KV-cache memory at inference = 2 · n_layers · n_kv · d_head · seq_len · batch · bytes — n_kv is the dominant lever.
- Quality preservation requires GQA training from scratch or careful 'uptraining' from an MHA checkpoint (Ainslie 2023).
- GQA has become the default even for closed frontier models where the exact architecture is not disclosed — Claude's head configuration is not publicly confirmed, but its serving economics are consistent with a GQA-family design.
- Multi-Head Latent Attention (MLA), introduced in DeepSeek-V2 (2024), compresses K/V into a shared low-rank latent vector, pushing KV-cache reduction further than GQA-8 at comparable quality — the frontier of this same curve.
- On SAP AI Core, the MHA/GQA/MQA choice is invisible for every model served as a managed generative AI hub endpoint; it becomes a real consulting decision only when self-hosting or fine-tuning an open-weight checkpoint.
Terms used on this page
- n_q / n_kv
- The number of Query heads (n_q) and Key/Value heads (n_kv) in a Transformer layer. MHA has n_kv = n_q; GQA has n_kv < n_q; MQA has n_kv = 1.
- KV cache
- The per-request memory store of computed Key and Value tensors for all past tokens, reused on each generation step to avoid recomputation. Dominates GPU memory at long context.
- Uptraining
- The procedure (Ainslie et al. 2023) for converting an existing MHA checkpoint into a GQA model with a small amount of extra training, preserving most original quality at a fraction of from-scratch cost.
- Concurrency
- The number of simultaneous user sessions a fixed GPU pool can serve at target latency. Inversely proportional to KV-cache memory per session — the key business KPI GQA improves.
- Attention head ratio (Q:KV)
- The single number (n_q divided by n_kv) that predicts inference memory and concurrency more reliably than total parameter count; worth asking for explicitly in any open-weight model card or self-hosting RFP.
- AI Units (SAP)
- SAP's metered pricing unit for Premium AI consumption on the generative AI hub; abstracts the underlying model's n_kv configuration away from the customer's bill, which is why a self-hosted deployment's cost behaviour can surprise a team used to hosted-endpoint pricing.
- Multi-Head Latent Attention (MLA)
- A further compression beyond GQA, introduced in DeepSeek-V2 (2024): keys and values are projected into a shared low-rank latent vector before caching, cutting KV-cache memory even further than GQA-8 at comparable quality — the current frontier of the memory-vs-quality curve this card describes.
Sources
- Ainslie et al. — GQA: Training Generalized Multi-Query Transformer Models (2023)
- Meta AI — Llama 3 Model Card and Architecture Notes
- Mistral's Model Lets You Vibe Long-Running Code in the Cloud — AI Business, 2026-05-12
- SAP Generative AI — official product page
- arXiv — Fast Transformer Decoding: One Write-Head is All You Need (Shazeer, 2019 — origin of MQA)
- Mistral AI — Announcing Mistral 7B (2023): confirms GQA + sliding window attention
- SAP Help Portal — Generative AI hub in SAP AI Core (hosted model endpoints)
- SAP Help Portal — Metering and pricing for generative AI on SAP AI Core (AI Units)
- Databricks — Model Serving documentation (self-hosting open-weight checkpoints)
- arXiv — FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (Dao et al., 2022)
- SAP — AI Units pricing page (Premium AI consumption model)
- arXiv — DeepSeek-V2: introduces Multi-Head Latent Attention (MLA) (2024)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.