Continuous Batching — How vLLM Hits 10× Throughput
As of 2026-09-27
What is Continuous Batching?
Static batching wastes 60-90% of GPU compute because one long request in a batch forces the GPU to keep running padded slots for requests that finished early — continuous batching frees each slot the instant its request ends.
What it is
Continuous batching — also called iteration-level or in-flight batching — is the scheduling technique that lets an LLM inference engine start processing a new request the moment an older one finishes, rather than waiting for an entire batch to complete together. It is the single largest reason engines such as vLLM and TensorRT-LLM reach roughly ten to twenty-four times higher throughput than a naive text-generation loop running on the same GPU hardware, and it is worth understanding even for a non-engineer evaluating an SAP AI deployment, because it directly drives the cost-per-token and the response latency a consultant will be asked to explain to a client.
Why it matters
- The scheduler reasons in iterations, not requests: it repacks the batch every forward pass, admitting a new pending request the moment any slot frees.
- Combined with PagedAttention, continuous batching keeps GPU utilization at 60-90% across requests of wildly different lengths, versus the massive waste of static batching.
- Throughput improves 5-10× at the same hardware budget, with P99 tail latency improving even more since short requests no longer queue behind long ones.
Key points
- Iteration-level scheduling — new requests admitted to the batch as soon as older ones finish, not at batch boundaries.
- Naive static batching wastes 60-90% of GPU compute on real workloads with variable output length.
- Combined with PagedAttention (C246), continuous batching keeps GPU at 60-90% effective utilisation across diverse-length requests.
- 10-24x throughput improvement over HuggingFace generate() baseline on the same hardware.
- P99 tail latency improves because short requests do not queue behind long ones.
- Implementation complexity is why most teams adopt vLLM or TensorRT-LLM rather than build their own scheduler.
- Complementary to speculative decoding: continuous batching is for throughput (high traffic); speculative decoding is for latency (single user).
- Practical benchmark: HuggingFace generate() at 20-50 tokens/s versus vLLM at 800-2,000 tokens/s on the same H100 with 32 concurrent requests.
- Moving nightly SAP document processing to vLLM: GPU hours reduced by up to 10× — from $3.10/h to $0.31/h for the same job.
- Single-user conversational Joule gains little from continuous batching — speculative decoding or a faster model tier are the right levers.
Terms used on this page
- Continuous batching
- Scheduling LLM inference at the iteration level so new requests are admitted into the batch as soon as older requests finish, rather than waiting for the whole batch to complete.
- Static batching
- The naive approach where a batch of requests is fixed at admission time and all step through forward passes together until the longest one finishes.
- Iteration-level scheduling
- Synonym for continuous batching; emphasises that scheduling decisions happen at every forward-pass iteration, not at request boundaries.
- Head-of-line blocking
- The latency problem in static batching where short requests must wait for the longest request in their batch to finish; eliminated by continuous batching.
- Speculative decoding
- A complementary latency optimisation where a small draft model proposes several tokens ahead and the large model verifies them in a single forward pass; reduces per-token latency for a single request, whereas continuous batching raises aggregate throughput across concurrent requests — production engines typically apply both.
- Max batch size (max_num_seqs)
- The scheduler configuration parameter capping how many sequences the engine can process concurrently per iteration; set too low, it wastes available GPU capacity; set too high relative to KV-cache memory, it risks page eviction and latency spikes.
Sources
- Yu et al. — Orca: A Distributed Serving System for Transformer-Based Generative Models (OSDI 2022, the iteration-batching paper)
- Kwon et al. — Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023)
- NVIDIA TensorRT-LLM in-flight batching documentation
- Anyscale — How continuous batching enables 23x throughput in LLM inference
- Databricks — LLM inference performance engineering best practices
- NVIDIA Docs — TensorRT-LLM
- vLLM Docs — Scheduler configuration reference
- vLLM Documentation — project home
- NVIDIA Technical Blog — TensorRT-LLM accelerates encoder-decoder models with in-flight batching
- NVIDIA Triton Inference Server — TensorRT-LLM backend documentation
- arXiv — SGLang: Efficient Execution of Structured Language Model Programs (2312.07104)
- LMSYS Org — Fast and Expressive LLM Inference with RadixAttention and SGLang
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.