LayerNorm and RMSNorm
As of 2026-09-27
What is LayerNorm and RMSNorm?
RMSNorm is 7-15% cheaper to compute than LayerNorm with negligible quality loss — why every major 2026 LLM (Llama, Mistral, Qwen) uses it, and why pre-norm, not post-norm, is what makes 80+ layer models trainable without warmup tricks.
What it is
LayerNorm and RMSNorm are normalisation layers that stabilise the distribution of activations across the feature dimension of each token at every layer of the transformer. They are not optional add-ons — without them, deep transformers (>12 layers) diverge in training within the first thousand steps.
The problem they solve is internal covariate shift: as activations flow through many matrix multiplications, their scale grows or shrinks exponentially, saturating nonlinearities and causing gradient explosion or vanishing. Normalisation anchors each token's feature vector to a stable scale before it enters the attention and FFN sublayers.
LayerNorm (Ba 2016) computes γ⊙((x−μ)/(σ+ε))+β, subtracting the mean μ and dividing by the standard deviation σ, then applying learned scale γ and bias β. RMSNorm (Zhang 2019) drops the mean subtraction: γ⊙(x/RMS(x)) where RMS(x) = √(∑xᵢ²/n). The bias term β is also removed. This makes RMSNorm 7–15% computationally cheaper than LayerNorm with negligible quality loss on language modelling, which is why Llama, Mistral, Qwen, and all major modern LLMs use RMSNorm. Pre-norm (x + Sublayer(Norm(x))) applies normalisation before the sublayer, enabling gradient flow to very early layers and making depth beyond 80 layers stable — all modern LLMs use pre-norm. Post-norm (Norm(x + Sublayer(x))) applies after, and was the original Vaswani 2017 design; it requires careful learning-rate warmup and fails without it at depth >24.
Choose RMSNorm pre-norm for any new transformer implementation. Use LayerNorm only for compatibility with BERT-era models or when importing pre-trained weights that require it.
Why it matters
- Deep transformers (>12 layers) diverge in training within the first thousand steps without normalisation — this isn't an optimisation, it's a hard requirement.
- Pre-norm's gradient-flow benefit is quantified in effect: it's what enables depth beyond 80 layers to be stable at all.
- Post-norm (the original 2017 design) fails without careful learning-rate warmup at depth beyond 24 layers — relevant only for BERT-era compatibility work.
Key points
- RMSNorm drops mean subtraction and bias vs LayerNorm — 7–15% fewer FLOPs, negligible quality loss.
- Pre-norm placement (Norm before sublayer) enables stable training beyond 80 layers; post-norm requires careful warmup.
- Llama 3.3 applies RMSNorm at 3 positions per layer: input_layernorm (before MHSA), post_attention_layernorm (before FFN), and model.norm (final output).
- The γ weight of RMSNorm is the scaling vector — large γ values indicate high-variance features that are sensitive to quantisation.
- INT8 KV quantisation interacts with normalisation: per-token dynamic quantisation is required when γ varies widely across tokens.
- bf16 training's coarser mantissa makes RMSNorm's simpler arithmetic (no running-mean subtraction) meaningfully less prone to cancellation error than LayerNorm at extreme sequence positions — one of the underappreciated reasons it displaced LayerNorm once bf16 became the default training format after 2022.
- At extreme depth (Wang et al.'s DeepNet reached 1,000 layers), plain pre-norm alone under-controls residual-stream variance growth; DeepNorm (a constant residual-branch scaling factor α applied at initialisation) and 'sandwich norm' (normalising both before and after each sublayer, as used in Gemma 2) are the two documented fixes beyond vanilla pre-norm.
Terms used on this page
- LayerNorm
- Normalisation layer (Ba 2016) that subtracts the mean and divides by standard deviation of each token's feature vector, then applies learned scale γ and bias β.
- RMSNorm
- Root Mean Square normalisation (Zhang 2019): normalises by RMS only, drops mean subtraction and bias term, 7–15% cheaper than LayerNorm.
- Pre-norm
- Placement of the normalisation layer at the INPUT of each sublayer: x + Sublayer(Norm(x)). Used by all modern LLMs for training stability at depth.
- Post-norm
- Placement of normalisation at the OUTPUT of each sublayer: Norm(x + Sublayer(x)). Original Vaswani 2017 design; requires careful LR warmup.
- Internal covariate shift
- Change in the distribution of layer inputs during training as parameters update; normalisation layers mitigate this to stabilise gradient flow.
Sources
- Ba et al. 2016 — Layer Normalisation
- Zhang & Sennrich 2019 — Root Mean Square Layer Normalisation
- Wang et al. 2020 — On Layer Normalisation in the Transformer Architecture (pre-norm analysis)
- Meta AI — Llama 3 Model Card
- AutoGPTQ documentation — quantisation and outlier handling
- SAP Generative AI — official product page
- Meta AI — Llama model research
- SAP Help Portal — Generative AI Hub in SAP AI Core, model library overview (2026)
- SAP Help Portal — Orchestration service, prompting vs fine-tuning guidance (2026)
- Wang et al. — DeepNet: Scaling Transformers to 1,000 Layers, arXiv:2203.00555 (2022)
- PyTorch documentation — torch.nn.LayerNorm reference
- Jiang et al. — Mistral 7B technical report, arXiv:2310.06825 (2023)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.