AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

Multi-head Attention vs Grouped-Query Attention vs Multi-query Attention

Multi-head Attention vs Grouped-Query Attention vs Multi-query Attention — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-09-27

Multi-head Attention vs Grouped-Query Attention vs Multi-query Attention: what is the difference?

Switching from MHA to GQA-8 at equal model size cuts KV-cache memory roughly 8x for under 1% MMLU drop, doubling or tripling how many concurrent users one H100 can serve — the real reason Mistral and Llama 3 serve cheaper than equivalent MHA models.

What it is

Multi-head Attention (MHA), Grouped-Query Attention (GQA) and Multi-Query Attention (MQA) are three points on the same trade-off curve: how many distinct Key/Value projections does each layer compute, given a fixed number of Query heads. The choice does not change what a model can express; it changes how much KV-cache memory inference burns per token, and therefore the per-request cost and the maximum concurrency a given GPU can serve.

Why it exists. The original Transformer (2017) used MHA — every Query head has its own Key and Value head. That doubles representational diversity but also doubles KV-cache memory at inference, which is the dominant memory cost at long context (companion C160). MQA (Shazeer 2019) collapses all heads to a single shared K/V — minimal KV memory, but a measurable quality drop at scale. GQA (Ainslie et al. 2023) is the practical compromise that every frontier model since Llama 2 has adopted.

How they work. With n q Query heads and n kv Key/Value heads: MHA uses n kv = n q (e.g. Llama 1 70B: 64 Q, 64 KV). MQA uses n kv = 1 (one shared K/V across all Q heads). GQA groups Q heads into g groups that share K/V (e.g. Llama 3 70B: 64 Q, 8 KV — 8 Q heads per shared KV pair). Llama 3/4, Mistral, Mixtral, Gemma 2, and Claude (per inferred architecture) all use GQA. GPT-4 family uses MHA per public reverse-engineering analyses; GPT-5 details remain undisclosed.

Why it matters

  • MHA (Llama 1 70B: 64 Q, 64 KV) maximises representational diversity but doubles KV-cache memory versus alternatives; MQA collapses to one shared K/V head with a measurable quality drop at scale.
  • GQA (Llama 3 70B: 64 Q, 8 KV) is the practical compromise every frontier model since Llama 2 has adopted.
  • If you self-host or fine-tune via SAP AI Core, this single architectural choice controls sustainable concurrency per GPU — invisible to a hosted-API consumer, but essential to know why some models cost less to serve.

Key points

  • Three points on one curve — MHA (n_kv = n_q), GQA (n_kv = n_q / g), MQA (n_kv = 1) — trading KV memory for quality.
  • GQA-8 (8 Q heads per KV pair) is the default for Llama 3/4, Mistral 7B/Large, Mixtral 8x7B, Gemma 2 — ~8× KV memory reduction vs MHA, <1% MMLU loss.
  • MQA (n_kv = 1) maximises memory savings but degrades multi-task quality; rarely used in frontier models post-2023.
  • KV-cache memory at inference = 2 · n_layers · n_kv · d_head · seq_len · batch · bytes — n_kv is the dominant lever.
  • Quality preservation requires GQA training from scratch or careful 'uptraining' from an MHA checkpoint (Ainslie 2023).
  • GQA has become the default even for closed frontier models where the exact architecture is not disclosed — Claude's head configuration is not publicly confirmed, but its serving economics are consistent with a GQA-family design.
  • Multi-Head Latent Attention (MLA), introduced in DeepSeek-V2 (2024), compresses K/V into a shared low-rank latent vector, pushing KV-cache reduction further than GQA-8 at comparable quality — the frontier of this same curve.
  • On SAP AI Core, the MHA/GQA/MQA choice is invisible for every model served as a managed generative AI hub endpoint; it becomes a real consulting decision only when self-hosting or fine-tuning an open-weight checkpoint.

Terms used on this page

n_q / n_kv
The number of Query heads (n_q) and Key/Value heads (n_kv) in a Transformer layer. MHA has n_kv = n_q; GQA has n_kv < n_q; MQA has n_kv = 1.
KV cache
The per-request memory store of computed Key and Value tensors for all past tokens, reused on each generation step to avoid recomputation. Dominates GPU memory at long context.
Uptraining
The procedure (Ainslie et al. 2023) for converting an existing MHA checkpoint into a GQA model with a small amount of extra training, preserving most original quality at a fraction of from-scratch cost.
Concurrency
The number of simultaneous user sessions a fixed GPU pool can serve at target latency. Inversely proportional to KV-cache memory per session — the key business KPI GQA improves.
Attention head ratio (Q:KV)
The single number (n_q divided by n_kv) that predicts inference memory and concurrency more reliably than total parameter count; worth asking for explicitly in any open-weight model card or self-hosting RFP.
AI Units (SAP)
SAP's metered pricing unit for Premium AI consumption on the generative AI hub; abstracts the underlying model's n_kv configuration away from the customer's bill, which is why a self-hosted deployment's cost behaviour can surprise a team used to hosted-endpoint pricing.
Multi-Head Latent Attention (MLA)
A further compression beyond GQA, introduced in DeepSeek-V2 (2024): keys and values are projected into a shared low-rank latent vector before caching, cutting KV-cache memory even further than GQA-8 at comparable quality — the current frontier of the memory-vs-quality curve this card describes.

Sources

  1. Ainslie et al. — GQA: Training Generalized Multi-Query Transformer Models (2023)
  2. Meta AI — Llama 3 Model Card and Architecture Notes
  3. Mistral's Model Lets You Vibe Long-Running Code in the Cloud — AI Business, 2026-05-12
  4. SAP Generative AI — official product page
  5. arXiv — Fast Transformer Decoding: One Write-Head is All You Need (Shazeer, 2019 — origin of MQA)
  6. Mistral AI — Announcing Mistral 7B (2023): confirms GQA + sliding window attention
  7. SAP Help Portal — Generative AI hub in SAP AI Core (hosted model endpoints)
  8. SAP Help Portal — Metering and pricing for generative AI on SAP AI Core (AI Units)
  9. Databricks — Model Serving documentation (self-hosting open-weight checkpoints)
  10. arXiv — FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (Dao et al., 2022)
  11. SAP — AI Units pricing page (Premium AI consumption model)
  12. arXiv — DeepSeek-V2: introduces Multi-Head Latent Attention (MLA) (2024)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →