AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

Quantization Strategies — INT8, INT4, AWQ, GPTQ, GGUF

Quantization Strategies — INT8, INT4, AWQ, GPTQ, GGUF — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-10-06

What is Quantization Strategies?

INT4 quantization shrinks a 70B model's footprint from 140GB to 35GB — the difference between needing two H100s and just one A100 — for only 1-3% accuracy loss on MMLU/HumanEval.

What it is

Quantization shrinks a model's weights from 16-bit floats (FP16/BF16, the training format) down to 8-bit (INT8) or 4-bit (INT4) integers, cutting GPU memory 2-4x and inference latency 1.5-3x — at the cost of small, controlled accuracy loss. For a 70B-parameter model, INT4 turns a 140 GB FP16 footprint into 35 GB, the difference between needing two H100s and one A100 — a cost gap that decides whether a use case is economically viable for on-premises or edge deployment.

Four families dominate. AWQ (Activation-aware Weight Quantization, MIT 2023) measures which weight channels carry the most signal on a calibration dataset, leaves those in higher precision, and aggressively quantizes the rest; preserves accuracy within 0.5-1 perplexity point of FP16 at 4-bit. GPTQ (Frantar et al. 2022) is the older one-shot quantizer using second-order information from a calibration set; very fast to apply, slightly worse accuracy than AWQ on instruction-tuned models. INT8 SmoothQuant moves activation outliers into the weight matrix so both can be 8-bit cleanly; the conservative choice when accuracy must be near-FP16. GGUF (GPT-Generated Unified Format, llama.cpp ecosystem) is a container format with multiple quantization variants (Q2_K, Q4_K_M, Q5_K_M, Q8_0); the de-facto format for llama.cpp and edge deployment.

Why it matters

  • AWQ preserves accuracy within 0.5-1 perplexity point of FP16 at 4-bit by leaving high-signal weight channels in higher precision, edging out the older GPTQ on instruction-tuned models.
  • INT8 SmoothQuant is the conservative choice with under 1% degradation, the default for data-centre serving where accuracy can't be compromised.
  • Below INT4 (3-bit, 2-bit) degradation becomes task-dependent — math and code reasoning collapse first, conversational chat degrades last.

Key points

  • Quantization shrinks weights from FP16 to INT8 (2x memory cut) or INT4 (4x memory cut) — at small accuracy cost.
  • AWQ (activation-aware, MIT 2023) — best 4-bit accuracy, 0.5-1 perplexity point of FP16; default for instruction-tuned models.
  • GPTQ — older one-shot quantizer, fast to apply, slightly worse than AWQ on chat models.
  • INT8 SmoothQuant — conservative 2x compression, <1% accuracy hit; default for production data-centre serving.
  • GGUF — container format with Q2_K to Q8_0 variants; canonical for llama.cpp edge + laptop deployments.
  • Below INT4 (3-bit, 2-bit) — math and code reasoning collapse first; chat degrades last.

Terms used on this page

Quantization
Reducing the numerical precision of a model's weights (and optionally activations) from FP16/BF16 down to INT8 or INT4 to shrink memory and accelerate inference.
AWQ
Activation-aware Weight Quantization (MIT 2023); identifies salient weight channels from a calibration dataset and preserves them in higher precision while aggressively quantizing the rest.
GPTQ
One-shot post-training quantizer (Frantar et al. 2022) using approximate second-order information from a small calibration set; fast but slightly less accurate than AWQ on instruction-tuned models.
GGUF
Container format for quantized models in the llama.cpp ecosystem; carries weight tensors at variants from Q2_K (2-bit) to Q8_0 (8-bit) with metadata.
SmoothQuant (INT8)
A quantization method that migrates activation outliers into the weight matrix mathematically before quantizing, so both weights and activations can run cleanly at 8-bit — the conservative, near-FP16-accuracy choice this card recommends for regulated or safety-relevant outputs.
Calibration dataset
The small, representative sample of real inputs used by AWQ, GPTQ and similar post-training quantizers to decide which weight channels matter most and preserve them in higher precision — the quality of this sample directly bounds the quantized model's accuracy.
Distillation vs. quantization
Two distinct compression techniques often confused: distillation trains a smaller model to mimic a larger one's behaviour (a new model, a new training run); quantization keeps the same architecture and weights but stores them at lower numerical precision (no retraining).
BYOM (Bring Your Own Model)
SAP AI Core's documented path for self-hosting an open-source or custom-quantized model — via Ollama, LocalAI, llama.cpp, vLLM or a custom Hugging Face Transformers server — instead of consuming a hosted foundation model through the generative AI hub.

Sources

  1. Lin et al. — AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
  2. Frantar et al. — GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
  3. Xiao et al. — SmoothQuant: Accurate and Efficient Post-Training Quantization for LLMs
  4. llama.cpp GGUF format specification
  5. GitHub SAP-samples — btp-generative-ai-hub-use-cases: bring-your-own OSS LLM on SAP AI Core (Ollama, LocalAI, llama.cpp, vLLM, custom HF Transformers server)
  6. SAP Help Portal — Choose a Resource Plan for training/inference in SAP AI Core (GPU-tier consequence of a quantization choice)
  7. GitHub — casper-hansen/AutoAWQ (AWQ quantization library)
  8. GitHub SAP-docs — sap-artificial-intelligence: SAP Domain Models / foundation models overview
  9. SAP News Center — SAP Sapphire keynote: Business AI Platform to power the Autonomous Enterprise (SAP Domain Models, Early Adopter Care → GA target Q3 2026, 2026-05-12)
  10. NVIDIA Newsroom — SAP and NVIDIA to Accelerate Generative AI Adoption Across Enterprise Applications
  11. GitHub NVIDIA — TensorRT-LLM (FP8/INT8 quantized execution graphs)
  12. ggml-org — llama.cpp release v0.6.0 (NVFP4/MXFP4 W4A4 path, Vulkan sparse flash attention for quantised K/V, imatrix activation statistics; 5 Oct 2026)
  13. vLLM project — release v0.31.0 (NVFP4 compressed KV cache default, fp8_per_tensor, vllm preload; 5 Oct 2026)
  14. vLLM — Quantization (supported quantisation methods and hardware matrix)
  15. ggml-org — llama.cpp quantize tool README (GGUF quantisation types and imatrix usage)
  16. Hugging Face — Transformers quantization overview (bitsandbytes, AWQ, GPTQ and other backends compared)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →