Run SGLang Offline Batch Inference Directly from Python
In this article
- Prerequisites
- Step 1: Imports and the execution contract
- Step 2: Define the batch workload
- Step 3: Initialise the engine
- Step 4: Run generation and consume outputs
- Step 5: Protect the spawn-based entry point
- Architecture tradeoffs
- What to watch out for
- Where to go next
- FAQ
- How do I run SGLang offline batch inference without an HTTP server?
- What command-line arguments can I pass to offline_batch_inference.py?
- How do I stop SGLang from using all GPU memory for the KV cache?
- Does offline batch inference support sampling parameters like temperature and top_p?
- Can I use SGLang offline inference with tensor parallelism across multiple GPUs?
SGLang has rapidly become one of the most compelling inference engines for large language models because its RadixAttention mechanism can reuse shared prefix state across large prompt sets. While a persistent HTTP server is the usual serving pattern for interactive applications, many model workflows—synthetic data generation, document scoring, evaluation suites, and reinforcement learning rollouts—do not need an always-on endpoint. For those jobs, offline batch inference is often the better deployment shape.
Running offline batch inference removes the network hop, HTTP serialization, and asynchronous polling that a serving API adds. You submit a list of prompts directly to the in-process engine, and the runtime spends its time on token generation rather than request handling. The pattern is particularly valuable when you need to process a large corpus overnight and want the GPU to remain saturated with forward passes.
This guide follows the official offline_batch_inference.py example from SGLang. The code is available in the SGLang repository and is provided under the Apache-2.0 licence.
Prerequisites
Before you run the script, confirm that the environment can load and execute the engine:
- A Linux machine or container with Python 3.8 or newer.
- An NVIDIA GPU with enough VRAM for your chosen model weights plus KV cache, and a CUDA driver compatible with your SGLang build.
- The
sglangPython package installed, including its SRT runtime dependencies. - Hugging Face credentials configured with
huggingface-cli loginif you plan to download gated model weights such asmeta-llama/Llama-3.1-8B-Instruct.
Step 1: Imports and the execution contract
The script starts with the same usage note that ships in the source example. It imports argparse for CLI parsing, msgspec for typed configuration serialization, the top-level sglang module, and SGLang's ServerArgs configuration class.
"""
Usage:
python3 offline_batch_inference.py --model meta-llama/Llama-3.1-8B-Instruct
"""
import argparse
import msgspec
import sglang as sgl
from sglang.srt.server_args import ServerArgs
ServerArgs is the central configuration object for the engine. Instead of building a separate offline-only config schema, the example reuses the same class that the SGLang runtime server accepts. That means the script can consume the same --model, tensor-parallelism, memory, and scheduling flags you already use when starting an SGLang server.
Step 2: Define the batch workload
The main function receives a fully populated ServerArgs object. In a production pipeline you might read prompts from a Parquet file, a job queue, or a feature store; the example defines a small in-memory list so the instruction is easy to run.
def main(
server_args: ServerArgs,
):
# Sample prompts.
prompts = [
"Hello, my name is",
"The president of the United States is",
"The capital of France is",
"The future of AI is",
]
# Create a sampling params object.
sampling_params = {"temperature": 0.8, "top_p": 0.95}
The script passes every prompt to the engine in one call, which is the essential difference from a serving loop where each request arrives independently. The sampling_params dictionary controls generation for the whole batch. A temperature of 0.8 and top_p of 0.95 produce varied text; if you are extracting structured JSON or running classification, set temperature to 0 for greedy decoding.
Step 3: Initialise the engine
The next line constructs the engine from the parsed server arguments.
# Create an LLM.
llm = sgl.Engine(**msgspec.structs.asdict(server_args))
msgspec.structs.asdict converts the typed ServerArgs object back into a dictionary of valid engine keyword arguments. sgl.Engine then loads the model weights into VRAM and prepares the runtime for generation. This is the expensive initialization step; it runs before any prompt is processed and can take much longer than the generation stage for small prompt sets.
Step 4: Run generation and consume outputs
llm.generate blocks until every prompt in the list has completed, either by reaching an end-of-sequence token or by hitting the configured maximum token limit.
outputs = llm.generate(prompts, sampling_params)
# Print the outputs.
for prompt, output in zip(prompts, outputs):
print("===============================")
print(f"Prompt: {prompt}\nGenerated text: {output['text']}")
Each element in outputs is a dictionary. The example prints the text field, but the same object can also carry metadata such as token usage, finish reason, and requested scoring information. Zipping prompts with outputs maps each input to its generated continuation in the original order.
Step 5: Protect the spawn-based entry point
The __main__ block is placed after main in the source file, but it is the actual program entry point. It builds an ArgumentParser, registers SGLang's CLI flags, parses them, and passes the resulting ServerArgs to main.
# The __main__ condition is necessary here because we use "spawn" to create subprocesses
# Spawn starts a fresh program every time, if there is no __main__, it will run into infinite loop to keep spawning processes from sgl.Engine
if __name__ == "__main__":
parser = argparse.ArgumentParser()
ServerArgs.add_cli_args(parser)
args = parser.parse_args()
server_args = ServerArgs.from_cli_args(args)
main(server_args)
Do not remove the if __name__ == "__main__": guard. SGLang uses subprocesses for parts of its runtime, and the source comment warns that a fresh spawned interpreter can re-import and re-execute the script. Without the guard, each child process can spawn further engine processes, leading to recursive process creation and an exhausted machine.
Architecture tradeoffs
The table below separates the three execution modes that a local SGLang setup can use.
| Execution Model | Throughput | Latency Profile | Networking Overhead | Best Use Case |
|---|---|---|---|---|
| Offline Batching (SGLang Engine) | Highest for bulk workloads | High time-to-first-token | None (in-memory) | Bulk text generation, evaluation suites, document expansion |
| API Server (HTTP/REST) | High | Low time-to-first-token | Moderate | User-facing chat applications, distributed agent networks |
| Streaming Server (Server-Sent Events) | Moderate | Ultra-low time-to-first-token | High | Real-time typing interfaces, interactive dashboards |
When you need maximum tokens per second for a known set of prompts, offline batching is the most direct model. The server modes add the request/response machinery that is necessary for concurrent clients but not useful for a job that already owns the entire prompt list.
What to watch out for
Local GPU inference at scale exposes failure modes that managed APIs hide.
Out-of-Memory (OOM) Exhaustion: The engine reserves VRAM for model weights and then allocates a KV cache pool. If you run other processes on the same GPU, limit the cache memory explicitly through the CLI. Otherwise, a competing allocation can trigger a CUDA OOM fault mid-generation.
Tensor Parallelism Constraints: If you use --tp to split a model across multiple GPUs, every participating GPU must be visible to the process. Slow interconnects can reduce the expected speedup because tensor-parallel collectives synchronize during each forward pass. Batch-size tuning and KV-cache placement then operate across the whole parallel group.
Long-Context and RAG Pressure: Dense retrieval and long context windows fill the RadixAttention cache much faster than short completions. This is a different memory pattern from the standard instruct workload covered in our look at /news/perplexity-pplx-embed-v2-context-9b-preview-rag. Watch the model's maximum context length when you batch large documents, because each sequence's attention history can consume a large block of the cache.
Tokenizer Mismatches: SGLang expects the tokenizer files to match the model weights. If you point --model at a directory with the wrong tokenizer configuration, generation can complete without raising an error while producing incoherent text. Verify that the model directory contains a valid tokenizer_config.json for the exact revision you intend to load.
Where to go next
After you have reproduced this script, replace the four sample prompts with a larger dataset from disk. Then test the effect of larger batch sizes, a lower sampling temperature, and a mixture-of-experts model pointer to see how the scheduler adapts. Once the offline path is stable, you can embed the same sgl.Engine call inside a data-processing job and keep the GPU busy without running an HTTP listener at all.
FAQ
How do I run SGLang offline batch inference without an HTTP server?
Use the Python sgl.Engine API directly. The example script parses ServerArgs from CLI flags and calls llm.generate(prompts, sampling_params) in-process. This bypasses REST or gRPC serialization and is suitable for bulk jobs that do not need concurrent client request handling.
What command-line arguments can I pass to offline_batch_inference.py?
Because the script calls ServerArgs.add_cli_args(parser), it accepts the same server arguments as the SGLang runtime, including --model, --tp, and memory-related flags. For example, you can run python3 offline_batch_inference.py --model meta-llama/Llama-3.1-8B-Instruct to select an instruct model.
How do I stop SGLang from using all GPU memory for the KV cache?
Pass --mem-fraction-static with a value below the default to reserve a portion of VRAM for the cache. The model weights are loaded first, and the remaining memory is available for the KV cache pool. Setting the value too high can cause out-of-memory failures when the GPU also has container or desktop overhead.
Does offline batch inference support sampling parameters like temperature and top_p?
Yes. The example passes a sampling dictionary with temperature 0.8 and top_p 0.95 directly to the generate call. The engine applies those values to every prompt in the batch; use separate calls or per-item overrides if you need different values.
Can I use SGLang offline inference with tensor parallelism across multiple GPUs?
Yes, pass --tp 2 or --tp 4 through the CLI and ServerArgs.from_cli_args will preserve it. All GPUs must be visible to the process, and communication overhead can reduce the benefit if the interconnect is slow. Batch sizing and KV cache allocation then apply across the tensor-parallel group.
Frequently asked questions
How do I run SGLang offline batch inference without an HTTP server?
Use the Python sgl.Engine API directly. The example script parses ServerArgs from CLI flags and calls llm.generate(prompts, sampling_params) in-process. This bypasses REST or gRPC serialization and is suitable for bulk jobs that do not need concurrent client request handling.
What command-line arguments can I pass to offline_batch_inference.py?
Because the script calls ServerArgs.add_cli_args(parser), it accepts the same server arguments as the SGLang runtime, including --model, --tp, and memory-related flags. For example, you can run python3 offline_batch_inference.py --model meta-llama/Llama-3.1-8B-Instruct to select an instruct model.
How do I stop SGLang from using all GPU memory for the KV cache?
Pass --mem-fraction-static with a value below the default to reserve a portion of VRAM for the cache. The model weights are loaded first, and the remaining memory is available for the KV cache pool. Setting the value too high can cause out-of-memory failures when the GPU also has container or desktop overhead.
Does offline batch inference support sampling parameters like temperature and top_p?
Yes. The example passes a sampling dictionary with temperature 0.8 and top_p 0.95 directly to the generate call. The engine applies those values to every prompt in the batch; use separate calls or per-item overrides if you need different values.
Can I use SGLang offline inference with tensor parallelism across multiple GPUs?
Yes, pass --tp 2 or --tp 4 through the CLI and ServerArgs.from_cli_args will preserve it. All GPUs must be visible to the process, and communication overhead can reduce the benefit if the interconnect is slow. Batch sizing and KV cache allocation then apply across the tensor-parallel group.
Related Guides

Chain of Verification with SGLang: Reduce LLM Hallucinations
Build a factored Chain-of-Verification pipeline with SGLang that runs draft, verify, refine, and summarize as independent LLM calls.
PagedAttention vs RadixAttention: How LLMs Tame the KV Cache
Two architectures attack LLM KV cache inefficiency from opposite angles — memory fragmentation and redundant prefill. Here is how each works.
Architect's Liquid Inference Auctions Every LLM Request
Architect's Liquid Inference auctions every LLM request across providers, locking the max price before the first token and paying 20% referral fees.