AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

LLM Inference Engines — vLLM vs TensorRT-LLM vs llama.cpp vs SGLang

LLM Inference Engines — vLLM vs TensorRT-LLM vs llama.cpp vs SGLang — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-09-27

What is LLM Inference Engines?

The inference engine, not the model, is the single biggest lever on LLM serving cost — the same Llama-3-70B can cost 3-10× more on a naive generate loop than on a tuned vLLM or TensorRT-LLM deployment.

An LLM inference engine is the runtime layer that turns a trained model's weights into served tokens: it schedules incoming requests, manages GPU memory, batches prompts together, and decodes outputs token by token. The engine choice is arguably the single biggest lever on serving cost in the entire generative AI stack — the same seventy-billion-parameter open-weight model can cost three to ten times more to run on a naive, unoptimised serving loop than on a properly tuned production engine, before anyone has touched model size or prompt design.

The Four Engines That Matter

vLLM, born out of UC Berkeley research, is the open-source reference implementation most teams reach for first. Its signature technique, PagedAttention, manages the key-value cache the way an operating system manages virtual memory, which lets it pack far more concurrent requests onto the same GPU than earlier serving stacks. Combined with continuous batching — folding new requests into a running batch instead of waiting for the whole batch to finish — vLLM covers a huge range of open-weight model families and is genuinely easy to operate, which is why it has become the default choice for teams self-hosting Llama, Mistral, Qwen or similar models on NVIDIA or AMD hardware.

Why it matters

  • vLLM is the open-source default for broad model coverage and Python ergonomics; TensorRT-LLM is 1.5-2× faster on H100/H200 but needs a 30-90 minute compilation step per model variant and locks you to NVIDIA.
  • llama.cpp is the edge/laptop engine (10-15 tok/s on a MacBook M-series), not a data-centre choice; SGLang's RadixAttention is 2-5× faster than vLLM specifically on agentic workloads with heavy prefix reuse.
  • The decision maps directly to workload: TensorRT-LLM for latency-sensitive single-user chat, vLLM for high-throughput batch RAG, SGLang for multi-turn Joule agents with shared system prompts.

Key points

  • Four dominant engines — vLLM (open-source default), TensorRT-LLM (max throughput on NVIDIA), llama.cpp (CPU/edge), SGLang (prefix reuse for agents).
  • vLLM combines PagedAttention (C246) + continuous batching (C245) — 10-24x throughput vs HuggingFace generate() baseline.
  • TensorRT-LLM compiles per-model engines (FP8/INT8 fused kernels) — 1.5-2x faster than vLLM on H100 but heavier build step + NVIDIA-only.
  • SGLang's RadixAttention shares KV-cache across requests with common prefixes — 2-5x speedup on agentic + few-shot workloads.
  • Engine choice is the single biggest cost lever — wrong engine can 3-10x the GPU bill before any model tuning matters.
  • Q4 quantisation on a 70B model: VRAM from 140 GB to 40 GB, 2-4% accuracy loss — often imperceptible for SAP extraction tasks.
  • OpenAI-compatible API (/v1/chat/completions): vLLM and TensorRT-LLM both implement it — swap model or engine without touching the SAP integration layer.
  • Production monitoring metrics: tokens/s per GPU, TTFT P50/P99, KV-cache hit rate, GPU memory utilisation (all exposed via Prometheus by vLLM).
  • llama.cpp: Llama-3-8B at 10-15 tokens/s on a MacBook M-series, 20-30 tokens/s on an RTX 4090 — a laptop/edge engine, not a datacenter one.
  • SAP deployment pattern: Docker container on SAP AI Core + OpenAI-compatible endpoint = BTP/Datasphere/Joule integration unchanged when the model is swapped.

Terms used on this page

Inference engine
The runtime that loads model weights into GPU/CPU memory, schedules incoming requests, manages the KV-cache, and decodes output tokens; distinct from the model itself.
vLLM
Open-source LLM inference engine from UC Berkeley (2023) built around PagedAttention and continuous batching; the de-facto reference for self-hosting open-weight models — and one of the serving backends SAP's own sample repository documents for the Bring-Your-Own-Model path on SAP AI Core.
TensorRT-LLM
NVIDIA's compiled inference engine that fuses CUDA kernels and applies FP8/INT8 quantization per model; highest absolute throughput on NVIDIA hardware, and architecturally close to what NVIDIA NIM microservices build on for SAP's own domain-model serving.
RadixAttention
SGLang's technique of organising KV-cache as a prefix tree so multiple requests sharing system prompt / few-shot exemplars reuse the cached compute.
llama.cpp
A C/C++ inference engine for quantised models (GGUF format) that runs on ordinary CPUs, Apple Silicon or a single consumer GPU; the natural choice for edge, laptop or air-gapped deployment, not for data-centre-scale serving.
Continuous batching
A scheduling technique that folds newly arrived requests into an already-running batch rather than waiting for the whole batch to finish, keeping GPU utilisation high under variable, real-world traffic.
OpenAI-compatible API
The /v1/chat/completions request/response contract that vLLM, TensorRT-LLM and most serving engines implement, letting a model or engine be swapped behind an SAP integration layer without touching the calling code.
GGUF
GPT-Generated Unified Format — the container format llama.cpp uses for quantised model weights, with multiple precision variants (Q2_K through Q8_0) inside one file (see C243).

Sources

  1. Kwon et al. — Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM paper, SOSP 2023)
  2. NVIDIA TensorRT-LLM documentation
  3. ggerganov/llama.cpp GitHub repository
  4. Zheng et al. — SGLang: Efficient Execution of Structured Language Model Programs
  5. SAP Help Portal — SAP Autonomous Suite documentation
  6. SAP News Center — How SAP and NVIDIA Advance Enterprise AI Transformation (NIM inference optimisation, SAP-ABAP-1, 2026-03-17)
  7. GitHub SAP-samples — btp-generative-ai-hub-use-cases: bring-your-own OSS LLM on SAP AI Core (Ollama, LocalAI, llama.cpp, vLLM, custom HF Transformers server)
  8. GitHub SGLang — sgl-project/sglang (RadixAttention engine)
  9. SAP Help Portal — Choose a Resource Plan for training/inference in SAP AI Core
  10. OpenAI — Chat Completions API reference (the compatibility contract vLLM/TensorRT-LLM implement)
  11. NVIDIA Newsroom — SAP and NVIDIA to Accelerate Generative AI Adoption Across Enterprise Applications
  12. GitHub NVIDIA — TensorRT-LLM

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →