AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

Attention Mechanism — From Scaled Dot-Product to Production

Attention Mechanism — From Scaled Dot-Product to Production — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-09-27

What is Attention Mechanism?

Attention replaced recurrence and convolution by computing every token-pair interaction in one parallel step — but that same design is O(n^2) in sequence length, exactly why every modern variant (FlashAttention, GQA, sliding window) exists to relax it.

What it is

Attention is the mechanism that lets every token in a sequence decide, in one parallel computation, which other tokens matter to it and by how much. It replaced recurrence and convolution as the dominant sequence-modelling primitive after the 2017 "Attention Is All You Need" paper, and every consumer-relevant model a consultant will cite — GPT, Claude, Gemini, Llama, Mistral — is a stack of attention layers wrapped in feed-forward blocks. Understanding it is not academic trivia for an SAP analytics practitioner; it is the concept that explains why long context is expensive, why GPU sizing for a self-hosted deployment is hard, and why a per-token API bill behaves the way it does.

The reason attention exists is a limitation of what came before. A recurrent network processes tokens one at a time and carries information forward through a hidden state, which means information from early in a long document has to survive many sequential update steps to influence a late token — in practice it degrades over distance. A convolutional network has a fixed receptive field per layer, so distant relationships need many stacked layers to reach each other. Attention computes a pairwise interaction between every pair of tokens in a single step: any token can directly inform any other token regardless of how far apart they sit in the sequence, and because it is a matrix operation rather than a sequential loop, it parallelizes cleanly on GPU and TPU silicon. That parallelism is a large part of why transformer-based training scaled the way it did through the 2020s.

Why it matters

  • The sqrt(d_k) scaling in softmax(QK^T/sqrt(d_k)) exists specifically to stop dot products saturating softmax at high embedding dimension.
  • Causal masking in decoders zeroes attention to future tokens — the mechanical reason autoregressive generation works one token at a time.
  • The O(n^2*d) cost is the direct driver of the per-token bill on a Joule deployment — tuning context length or debugging latency means tuning attention cost.

Key points

  • Scaled dot-product formula — softmax(Q·Kᵀ / √dₖ)·V — three projections per token, one softmax-weighted sum.
  • √dₖ scaling prevents softmax saturation at high embedding dimension; without it, gradients vanish on large models.
  • Cost is O(n²·d) in both compute and memory — quadratic in sequence length n, linear in head dimension d.
  • Causal mask in decoders zeros out future-token attention, enforcing left-to-right generation.
  • Every LLM optimisation since 2017 (FlashAttention, GQA, MQA, sliding window, sparse) attacks the O(n²) wall.
  • Multi-head attention runs several attention computations in parallel with different learned Q/K/V projections, then concatenates and re-projects the results — the actual production unit, not the single-head formula viewed in isolation.
  • FlashAttention and FlashAttention-2 do not change what attention computes, only how the GPU computes it — IO-aware kernel fusion cuts wall-clock latency and memory traffic without touching the O(n²) asymptotic complexity.
  • SAP's generative AI hub bills per token via AI Units, so attention's quadratic cost shows up directly as metered consumption on a customer's SAP invoice, not as a separate infrastructure line item.

Terms used on this page

Query / Key / Value (Q, K, V)
Three linear projections of each input token embedding. Q asks the question ('what am I looking for?'), K advertises the answer ('what do I contain?'), V carries the payload that gets mixed into the output.
Softmax
The normalisation function that turns raw dot-product scores into a probability distribution over tokens. Saturates (gradients vanish) when inputs are too large — the reason for the √dₖ scaling.
Causal mask
A triangular mask applied before softmax in decoder-only models so each position attends only to itself and earlier positions; enforces left-to-right autoregressive generation.
Head dimension (dₖ)
The size of each attention head's Q/K/V vectors. Typical values 64 or 128; total model embedding dim = num_heads × dₖ.
Multi-head attention
Running h independent attention computations in parallel, each with its own learned Q/K/V projections, then concatenating and linearly re-projecting the results — lets a layer attend to different relationships (syntax, coreference, position) at once.
FlashAttention
An IO-aware GPU kernel (Dao et al., 2022; v2 in 2023) that fuses the attention computation into blocks so the full n×n attention matrix is never materialised in GPU high-bandwidth memory, cutting wall-clock latency without changing the O(n²) compute complexity.

Sources

  1. Vaswani et al. — Attention Is All You Need (NeurIPS 2017)
  2. Stanford CS25 — Transformers United (lecture notes 2024-2026)
  3. SAP Business AI — official product page
  4. SAP Generative AI — official product page
  5. arXiv — FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (Dao et al., 2022)
  6. arXiv — FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning (Dao, 2023)
  7. SAP Help Portal — Orchestration service, generative AI hub (grounding, masking, filtering pipeline)
  8. SAP Help Portal — Generative AI hub in SAP AI Core (model access overview)
  9. Anthropic — Context windows documentation (2026 model context/output limits)
  10. SAP Help Portal — Metering and pricing for generative AI on SAP AI Core (AI Units)
  11. SAP Community — Why SAP needs a Knowledge Graph: giving enterprise AI a map of the business (2026)
  12. SAP — SAP-RPT tabular foundation model product page (contrast: in-context learning without sequential-token attention)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →