AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

Continuous Batching — How vLLM Hits 10× Throughput

Continuous Batching — How vLLM Hits 10× Throughput — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-09-27

What is Continuous Batching?

Static batching wastes 60-90% of GPU compute because one long request in a batch forces the GPU to keep running padded slots for requests that finished early — continuous batching frees each slot the instant its request ends.

What it is

Continuous batching — also called iteration-level or in-flight batching — is the scheduling technique that lets an LLM inference engine start processing a new request the moment an older one finishes, rather than waiting for an entire batch to complete together. It is the single largest reason engines such as vLLM and TensorRT-LLM reach roughly ten to twenty-four times higher throughput than a naive text-generation loop running on the same GPU hardware, and it is worth understanding even for a non-engineer evaluating an SAP AI deployment, because it directly drives the cost-per-token and the response latency a consultant will be asked to explain to a client.

Why it matters

  • The scheduler reasons in iterations, not requests: it repacks the batch every forward pass, admitting a new pending request the moment any slot frees.
  • Combined with PagedAttention, continuous batching keeps GPU utilization at 60-90% across requests of wildly different lengths, versus the massive waste of static batching.
  • Throughput improves 5-10× at the same hardware budget, with P99 tail latency improving even more since short requests no longer queue behind long ones.

Key points

  • Iteration-level scheduling — new requests admitted to the batch as soon as older ones finish, not at batch boundaries.
  • Naive static batching wastes 60-90% of GPU compute on real workloads with variable output length.
  • Combined with PagedAttention (C246), continuous batching keeps GPU at 60-90% effective utilisation across diverse-length requests.
  • 10-24x throughput improvement over HuggingFace generate() baseline on the same hardware.
  • P99 tail latency improves because short requests do not queue behind long ones.
  • Implementation complexity is why most teams adopt vLLM or TensorRT-LLM rather than build their own scheduler.
  • Complementary to speculative decoding: continuous batching is for throughput (high traffic); speculative decoding is for latency (single user).
  • Practical benchmark: HuggingFace generate() at 20-50 tokens/s versus vLLM at 800-2,000 tokens/s on the same H100 with 32 concurrent requests.
  • Moving nightly SAP document processing to vLLM: GPU hours reduced by up to 10× — from $3.10/h to $0.31/h for the same job.
  • Single-user conversational Joule gains little from continuous batching — speculative decoding or a faster model tier are the right levers.

Terms used on this page

Continuous batching
Scheduling LLM inference at the iteration level so new requests are admitted into the batch as soon as older requests finish, rather than waiting for the whole batch to complete.
Static batching
The naive approach where a batch of requests is fixed at admission time and all step through forward passes together until the longest one finishes.
Iteration-level scheduling
Synonym for continuous batching; emphasises that scheduling decisions happen at every forward-pass iteration, not at request boundaries.
Head-of-line blocking
The latency problem in static batching where short requests must wait for the longest request in their batch to finish; eliminated by continuous batching.
Speculative decoding
A complementary latency optimisation where a small draft model proposes several tokens ahead and the large model verifies them in a single forward pass; reduces per-token latency for a single request, whereas continuous batching raises aggregate throughput across concurrent requests — production engines typically apply both.
Max batch size (max_num_seqs)
The scheduler configuration parameter capping how many sequences the engine can process concurrently per iteration; set too low, it wastes available GPU capacity; set too high relative to KV-cache memory, it risks page eviction and latency spikes.

Sources

  1. Yu et al. — Orca: A Distributed Serving System for Transformer-Based Generative Models (OSDI 2022, the iteration-batching paper)
  2. Kwon et al. — Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023)
  3. NVIDIA TensorRT-LLM in-flight batching documentation
  4. Anyscale — How continuous batching enables 23x throughput in LLM inference
  5. Databricks — LLM inference performance engineering best practices
  6. NVIDIA Docs — TensorRT-LLM
  7. vLLM Docs — Scheduler configuration reference
  8. vLLM Documentation — project home
  9. NVIDIA Technical Blog — TensorRT-LLM accelerates encoder-decoder models with in-flight batching
  10. NVIDIA Triton Inference Server — TensorRT-LLM backend documentation
  11. arXiv — SGLang: Efficient Execution of Structured Language Model Programs (2312.07104)
  12. LMSYS Org — Fast and Expressive LLM Inference with RadixAttention and SGLang

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →