AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

Latency vs Throughput Trade-offs — TTFT, TPOT, SLA Design

Latency vs Throughput Trade-offs — TTFT, TPOT, SLA Design — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-09-27

Latency vs Throughput Trade-offs: what is the difference?

Chat, code-completion, batch jobs and voice agents each need a different TTFT/TPOT target — mixing a latency-sensitive chat endpoint with a throughput-hungry batch job guarantees the batch job tanks chat p99.

What it is

LLM serving has three orthogonal performance metrics that an SLA must address separately. Time-To-First-Token (TTFT) measures the wait from request submission to the first emitted token — dominated by the prefill phase, in which the model processes the full input prompt in one parallel pass. Time-Per-Output-Token (TPOT, also written as inter-token latency, ITL) measures the gap between successive emitted tokens — dominated by the decode phase, in which the model generates one token per forward pass. End-to-end latency is TTFT + (output tokens × TPOT). Throughput (tokens/sec, aggregated across concurrent requests) is the system-wide capacity number.

The core trade-off: larger continuous batches raise throughput by amortising decode-time memory bandwidth across more requests in parallel, but they raise TTFT for any request that joins a batch already in progress (it waits for the next scheduling tick) and they raise TPOT (more requests share the same GPU clock per token). Smaller batches give better single-request latency but waste GPU compute when traffic is low. Production serving stacks (vLLM, TensorRT-LLM, SGLang) all implement continuous batching with paged attention to reduce the trade-off — but they cannot eliminate it.

Why it matters

  • Larger continuous batches raise throughput but also raise TTFT (new requests wait for the next scheduling tick) and TPOT (more requests share the same GPU clock per token) — the trade-off can't be eliminated, only reduced.
  • Concrete SLA targets differ by workload: chat wants p95 TTFT under 500ms, code-completion under 200ms, batch summarisation ignores TTFT entirely to maximise batch size, voice agents want sub-100ms TPOT but tolerate 500-800ms TTFT.
  • The architect's fix is splitting workloads across separate endpoint tiers pointing at the same model file — one low-latency small-batch endpoint, one high-throughput large-batch endpoint.

Key points

  • Three orthogonal metrics: TTFT (prefill-bound), TPOT/ITL (decode-bound), throughput (system-wide tokens/sec aggregated across concurrent requests).
  • End-to-end latency = TTFT + (output_tokens × TPOT) — both must appear in any LLM SLA.
  • Continuous batching raises throughput but raises TTFT (requests wait for next scheduling tick) and TPOT (requests share clock); paged attention reduces but does not eliminate the trade-off.
  • Chat target: p95 TTFT < 500ms, TPOT 30-50ms. Code-completion: TTFT < 200ms. Batch jobs: throughput-first, TTFT irrelevant. Voice agents: TPOT < 100ms.
  • Split workloads across endpoint tiers — low-latency + small-batch for chat, high-throughput + large-batch for offline; never mix on one endpoint.
  • Benchmark at p50/p95/p99 — averages hide the long-tail latency that kills user experience.

Terms used on this page

TTFT
Time-To-First-Token — time from request submit to first emitted token; bounded by prefill compute over the full input prompt.
TPOT / ITL
Time-Per-Output-Token or Inter-Token Latency — time between successive emitted tokens during decoding; bounded by memory bandwidth at autoregressive generation.
Continuous batching
Serving technique where new requests join an in-flight batch at each decode step, instead of waiting for batch completion; standard in vLLM, TensorRT-LLM, SGLang.
Paged attention
Memory-management technique borrowed from virtual memory; KV-cache is stored in fixed-size pages to reduce fragmentation when serving variable-length sequences.
Prefill
The compute phase where the model processes the entire input prompt in one parallel forward pass before generating the first output token; the phase that determines TTFT.
Decode
The compute phase where the model generates output tokens one at a time, each requiring a separate forward pass; the phase that determines TPOT, and the phase continuous batching and PagedAttention primarily optimise.
Tail latency (p95 / p99)
The latency experienced by the slowest 5% (p95) or 1% (p99) of requests; the metric that actually reflects user-visible performance under load, since an average or median figure can look healthy while a meaningful share of requests are unacceptably slow.
Endpoint tiering
Running two or more separate serving deployments of the same model file, each configured (max-batch-size, reserved capacity) for a different SLA — the standard fix for a workload mixing interactive and batch traffic.

Sources

  1. NVIDIA — TensorRT-LLM Best Practices
  2. MLPerf Inference — Latency Metrics Specification
  3. Anyscale — LLM Serving Performance Deep Dive
  4. vLLM Docs — Scheduler configuration reference
  5. NVIDIA Triton Inference Server — TensorRT-LLM backend documentation
  6. Microsoft Learn — Provisioned throughput units (PTU) for Azure OpenAI
  7. Databricks Docs — Mosaic AI Model Serving
  8. Snowflake Docs — Cortex Analyst
  9. NVIDIA Developer Blog — Mastering LLM Techniques: Inference Optimization
  10. Databricks Blog — LLM inference performance engineering: best practices
  11. Anyscale Blog — Achieve 23x LLM Inference Throughput & Reduce p50 Latency (continuous batching)
  12. Hugging Face Docs — Text Generation Inference (TGI)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →