Quantization Strategies — INT8, INT4, AWQ, GPTQ, GGUF
As of 2026-10-06
What is Quantization Strategies?
INT4 quantization shrinks a 70B model's footprint from 140GB to 35GB — the difference between needing two H100s and just one A100 — for only 1-3% accuracy loss on MMLU/HumanEval.
What it is
Quantization shrinks a model's weights from 16-bit floats (FP16/BF16, the training format) down to 8-bit (INT8) or 4-bit (INT4) integers, cutting GPU memory 2-4x and inference latency 1.5-3x — at the cost of small, controlled accuracy loss. For a 70B-parameter model, INT4 turns a 140 GB FP16 footprint into 35 GB, the difference between needing two H100s and one A100 — a cost gap that decides whether a use case is economically viable for on-premises or edge deployment.
Four families dominate. AWQ (Activation-aware Weight Quantization, MIT 2023) measures which weight channels carry the most signal on a calibration dataset, leaves those in higher precision, and aggressively quantizes the rest; preserves accuracy within 0.5-1 perplexity point of FP16 at 4-bit. GPTQ (Frantar et al. 2022) is the older one-shot quantizer using second-order information from a calibration set; very fast to apply, slightly worse accuracy than AWQ on instruction-tuned models. INT8 SmoothQuant moves activation outliers into the weight matrix so both can be 8-bit cleanly; the conservative choice when accuracy must be near-FP16. GGUF (GPT-Generated Unified Format, llama.cpp ecosystem) is a container format with multiple quantization variants (Q2_K, Q4_K_M, Q5_K_M, Q8_0); the de-facto format for llama.cpp and edge deployment.
Why it matters
- AWQ preserves accuracy within 0.5-1 perplexity point of FP16 at 4-bit by leaving high-signal weight channels in higher precision, edging out the older GPTQ on instruction-tuned models.
- INT8 SmoothQuant is the conservative choice with under 1% degradation, the default for data-centre serving where accuracy can't be compromised.
- Below INT4 (3-bit, 2-bit) degradation becomes task-dependent — math and code reasoning collapse first, conversational chat degrades last.
Key points
- Quantization shrinks weights from FP16 to INT8 (2x memory cut) or INT4 (4x memory cut) — at small accuracy cost.
- AWQ (activation-aware, MIT 2023) — best 4-bit accuracy, 0.5-1 perplexity point of FP16; default for instruction-tuned models.
- GPTQ — older one-shot quantizer, fast to apply, slightly worse than AWQ on chat models.
- INT8 SmoothQuant — conservative 2x compression, <1% accuracy hit; default for production data-centre serving.
- GGUF — container format with Q2_K to Q8_0 variants; canonical for llama.cpp edge + laptop deployments.
- Below INT4 (3-bit, 2-bit) — math and code reasoning collapse first; chat degrades last.
Terms used on this page
- Quantization
- Reducing the numerical precision of a model's weights (and optionally activations) from FP16/BF16 down to INT8 or INT4 to shrink memory and accelerate inference.
- AWQ
- Activation-aware Weight Quantization (MIT 2023); identifies salient weight channels from a calibration dataset and preserves them in higher precision while aggressively quantizing the rest.
- GPTQ
- One-shot post-training quantizer (Frantar et al. 2022) using approximate second-order information from a small calibration set; fast but slightly less accurate than AWQ on instruction-tuned models.
- GGUF
- Container format for quantized models in the llama.cpp ecosystem; carries weight tensors at variants from Q2_K (2-bit) to Q8_0 (8-bit) with metadata.
- SmoothQuant (INT8)
- A quantization method that migrates activation outliers into the weight matrix mathematically before quantizing, so both weights and activations can run cleanly at 8-bit — the conservative, near-FP16-accuracy choice this card recommends for regulated or safety-relevant outputs.
- Calibration dataset
- The small, representative sample of real inputs used by AWQ, GPTQ and similar post-training quantizers to decide which weight channels matter most and preserve them in higher precision — the quality of this sample directly bounds the quantized model's accuracy.
- Distillation vs. quantization
- Two distinct compression techniques often confused: distillation trains a smaller model to mimic a larger one's behaviour (a new model, a new training run); quantization keeps the same architecture and weights but stores them at lower numerical precision (no retraining).
- BYOM (Bring Your Own Model)
- SAP AI Core's documented path for self-hosting an open-source or custom-quantized model — via Ollama, LocalAI, llama.cpp, vLLM or a custom Hugging Face Transformers server — instead of consuming a hosted foundation model through the generative AI hub.
Sources
- Lin et al. — AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- Frantar et al. — GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Xiao et al. — SmoothQuant: Accurate and Efficient Post-Training Quantization for LLMs
- llama.cpp GGUF format specification
- GitHub SAP-samples — btp-generative-ai-hub-use-cases: bring-your-own OSS LLM on SAP AI Core (Ollama, LocalAI, llama.cpp, vLLM, custom HF Transformers server)
- SAP Help Portal — Choose a Resource Plan for training/inference in SAP AI Core (GPU-tier consequence of a quantization choice)
- GitHub — casper-hansen/AutoAWQ (AWQ quantization library)
- GitHub SAP-docs — sap-artificial-intelligence: SAP Domain Models / foundation models overview
- SAP News Center — SAP Sapphire keynote: Business AI Platform to power the Autonomous Enterprise (SAP Domain Models, Early Adopter Care → GA target Q3 2026, 2026-05-12)
- NVIDIA Newsroom — SAP and NVIDIA to Accelerate Generative AI Adoption Across Enterprise Applications
- GitHub NVIDIA — TensorRT-LLM (FP8/INT8 quantized execution graphs)
- ggml-org — llama.cpp release v0.6.0 (NVFP4/MXFP4 W4A4 path, Vulkan sparse flash attention for quantised K/V, imatrix activation statistics; 5 Oct 2026)
- vLLM project — release v0.31.0 (NVFP4 compressed KV cache default, fp8_per_tensor, vllm preload; 5 Oct 2026)
- vLLM — Quantization (supported quantisation methods and hardware matrix)
- ggml-org — llama.cpp quantize tool README (GGUF quantisation types and imatrix usage)
- Hugging Face — Transformers quantization overview (bitsandbytes, AWQ, GPTQ and other backends compared)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.