Latency vs Throughput Trade-offs — TTFT, TPOT, SLA Design
As of 2026-09-27
Latency vs Throughput Trade-offs: what is the difference?
Chat, code-completion, batch jobs and voice agents each need a different TTFT/TPOT target — mixing a latency-sensitive chat endpoint with a throughput-hungry batch job guarantees the batch job tanks chat p99.
What it is
LLM serving has three orthogonal performance metrics that an SLA must address separately. Time-To-First-Token (TTFT) measures the wait from request submission to the first emitted token — dominated by the prefill phase, in which the model processes the full input prompt in one parallel pass. Time-Per-Output-Token (TPOT, also written as inter-token latency, ITL) measures the gap between successive emitted tokens — dominated by the decode phase, in which the model generates one token per forward pass. End-to-end latency is TTFT + (output tokens × TPOT). Throughput (tokens/sec, aggregated across concurrent requests) is the system-wide capacity number.
The core trade-off: larger continuous batches raise throughput by amortising decode-time memory bandwidth across more requests in parallel, but they raise TTFT for any request that joins a batch already in progress (it waits for the next scheduling tick) and they raise TPOT (more requests share the same GPU clock per token). Smaller batches give better single-request latency but waste GPU compute when traffic is low. Production serving stacks (vLLM, TensorRT-LLM, SGLang) all implement continuous batching with paged attention to reduce the trade-off — but they cannot eliminate it.
Why it matters
- Larger continuous batches raise throughput but also raise TTFT (new requests wait for the next scheduling tick) and TPOT (more requests share the same GPU clock per token) — the trade-off can't be eliminated, only reduced.
- Concrete SLA targets differ by workload: chat wants p95 TTFT under 500ms, code-completion under 200ms, batch summarisation ignores TTFT entirely to maximise batch size, voice agents want sub-100ms TPOT but tolerate 500-800ms TTFT.
- The architect's fix is splitting workloads across separate endpoint tiers pointing at the same model file — one low-latency small-batch endpoint, one high-throughput large-batch endpoint.
Key points
- Three orthogonal metrics: TTFT (prefill-bound), TPOT/ITL (decode-bound), throughput (system-wide tokens/sec aggregated across concurrent requests).
- End-to-end latency = TTFT + (output_tokens × TPOT) — both must appear in any LLM SLA.
- Continuous batching raises throughput but raises TTFT (requests wait for next scheduling tick) and TPOT (requests share clock); paged attention reduces but does not eliminate the trade-off.
- Chat target: p95 TTFT < 500ms, TPOT 30-50ms. Code-completion: TTFT < 200ms. Batch jobs: throughput-first, TTFT irrelevant. Voice agents: TPOT < 100ms.
- Split workloads across endpoint tiers — low-latency + small-batch for chat, high-throughput + large-batch for offline; never mix on one endpoint.
- Benchmark at p50/p95/p99 — averages hide the long-tail latency that kills user experience.
Terms used on this page
- TTFT
- Time-To-First-Token — time from request submit to first emitted token; bounded by prefill compute over the full input prompt.
- TPOT / ITL
- Time-Per-Output-Token or Inter-Token Latency — time between successive emitted tokens during decoding; bounded by memory bandwidth at autoregressive generation.
- Continuous batching
- Serving technique where new requests join an in-flight batch at each decode step, instead of waiting for batch completion; standard in vLLM, TensorRT-LLM, SGLang.
- Paged attention
- Memory-management technique borrowed from virtual memory; KV-cache is stored in fixed-size pages to reduce fragmentation when serving variable-length sequences.
- Prefill
- The compute phase where the model processes the entire input prompt in one parallel forward pass before generating the first output token; the phase that determines TTFT.
- Decode
- The compute phase where the model generates output tokens one at a time, each requiring a separate forward pass; the phase that determines TPOT, and the phase continuous batching and PagedAttention primarily optimise.
- Tail latency (p95 / p99)
- The latency experienced by the slowest 5% (p95) or 1% (p99) of requests; the metric that actually reflects user-visible performance under load, since an average or median figure can look healthy while a meaningful share of requests are unacceptably slow.
- Endpoint tiering
- Running two or more separate serving deployments of the same model file, each configured (max-batch-size, reserved capacity) for a different SLA — the standard fix for a workload mixing interactive and batch traffic.
Sources
- NVIDIA — TensorRT-LLM Best Practices
- MLPerf Inference — Latency Metrics Specification
- Anyscale — LLM Serving Performance Deep Dive
- vLLM Docs — Scheduler configuration reference
- NVIDIA Triton Inference Server — TensorRT-LLM backend documentation
- Microsoft Learn — Provisioned throughput units (PTU) for Azure OpenAI
- Databricks Docs — Mosaic AI Model Serving
- Snowflake Docs — Cortex Analyst
- NVIDIA Developer Blog — Mastering LLM Techniques: Inference Optimization
- Databricks Blog — LLM inference performance engineering: best practices
- Anyscale Blog — Achieve 23x LLM Inference Throughput & Reduce p50 Latency (continuous batching)
- Hugging Face Docs — Text Generation Inference (TGI)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.