Attention Mechanism — From Scaled Dot-Product to Production
As of 2026-09-27
What is Attention Mechanism?
Attention replaced recurrence and convolution by computing every token-pair interaction in one parallel step — but that same design is O(n^2) in sequence length, exactly why every modern variant (FlashAttention, GQA, sliding window) exists to relax it.
What it is
Attention is the mechanism that lets every token in a sequence decide, in one parallel computation, which other tokens matter to it and by how much. It replaced recurrence and convolution as the dominant sequence-modelling primitive after the 2017 "Attention Is All You Need" paper, and every consumer-relevant model a consultant will cite — GPT, Claude, Gemini, Llama, Mistral — is a stack of attention layers wrapped in feed-forward blocks. Understanding it is not academic trivia for an SAP analytics practitioner; it is the concept that explains why long context is expensive, why GPU sizing for a self-hosted deployment is hard, and why a per-token API bill behaves the way it does.
The reason attention exists is a limitation of what came before. A recurrent network processes tokens one at a time and carries information forward through a hidden state, which means information from early in a long document has to survive many sequential update steps to influence a late token — in practice it degrades over distance. A convolutional network has a fixed receptive field per layer, so distant relationships need many stacked layers to reach each other. Attention computes a pairwise interaction between every pair of tokens in a single step: any token can directly inform any other token regardless of how far apart they sit in the sequence, and because it is a matrix operation rather than a sequential loop, it parallelizes cleanly on GPU and TPU silicon. That parallelism is a large part of why transformer-based training scaled the way it did through the 2020s.
Why it matters
- The sqrt(d_k) scaling in softmax(QK^T/sqrt(d_k)) exists specifically to stop dot products saturating softmax at high embedding dimension.
- Causal masking in decoders zeroes attention to future tokens — the mechanical reason autoregressive generation works one token at a time.
- The O(n^2*d) cost is the direct driver of the per-token bill on a Joule deployment — tuning context length or debugging latency means tuning attention cost.
Key points
- Scaled dot-product formula — softmax(Q·Kᵀ / √dₖ)·V — three projections per token, one softmax-weighted sum.
- √dₖ scaling prevents softmax saturation at high embedding dimension; without it, gradients vanish on large models.
- Cost is O(n²·d) in both compute and memory — quadratic in sequence length n, linear in head dimension d.
- Causal mask in decoders zeros out future-token attention, enforcing left-to-right generation.
- Every LLM optimisation since 2017 (FlashAttention, GQA, MQA, sliding window, sparse) attacks the O(n²) wall.
- Multi-head attention runs several attention computations in parallel with different learned Q/K/V projections, then concatenates and re-projects the results — the actual production unit, not the single-head formula viewed in isolation.
- FlashAttention and FlashAttention-2 do not change what attention computes, only how the GPU computes it — IO-aware kernel fusion cuts wall-clock latency and memory traffic without touching the O(n²) asymptotic complexity.
- SAP's generative AI hub bills per token via AI Units, so attention's quadratic cost shows up directly as metered consumption on a customer's SAP invoice, not as a separate infrastructure line item.
Terms used on this page
- Query / Key / Value (Q, K, V)
- Three linear projections of each input token embedding. Q asks the question ('what am I looking for?'), K advertises the answer ('what do I contain?'), V carries the payload that gets mixed into the output.
- Softmax
- The normalisation function that turns raw dot-product scores into a probability distribution over tokens. Saturates (gradients vanish) when inputs are too large — the reason for the √dₖ scaling.
- Causal mask
- A triangular mask applied before softmax in decoder-only models so each position attends only to itself and earlier positions; enforces left-to-right autoregressive generation.
- Head dimension (dₖ)
- The size of each attention head's Q/K/V vectors. Typical values 64 or 128; total model embedding dim = num_heads × dₖ.
- Multi-head attention
- Running h independent attention computations in parallel, each with its own learned Q/K/V projections, then concatenating and linearly re-projecting the results — lets a layer attend to different relationships (syntax, coreference, position) at once.
- FlashAttention
- An IO-aware GPU kernel (Dao et al., 2022; v2 in 2023) that fuses the attention computation into blocks so the full n×n attention matrix is never materialised in GPU high-bandwidth memory, cutting wall-clock latency without changing the O(n²) compute complexity.
Sources
- Vaswani et al. — Attention Is All You Need (NeurIPS 2017)
- Stanford CS25 — Transformers United (lecture notes 2024-2026)
- SAP Business AI — official product page
- SAP Generative AI — official product page
- arXiv — FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (Dao et al., 2022)
- arXiv — FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning (Dao, 2023)
- SAP Help Portal — Orchestration service, generative AI hub (grounding, masking, filtering pipeline)
- SAP Help Portal — Generative AI hub in SAP AI Core (model access overview)
- Anthropic — Context windows documentation (2026 model context/output limits)
- SAP Help Portal — Metering and pricing for generative AI on SAP AI Core (AI Units)
- SAP Community — Why SAP needs a Knowledge Graph: giving enterprise AI a map of the business (2026)
- SAP — SAP-RPT tabular foundation model product page (contrast: in-context learning without sequential-token attention)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.