Model or dataset
sgl-project/mini-sglang avatar
sgl-project/mini-sglang

Mini-SGLang: A Compact LLM Inference Engine for Learning and Production

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

5,186 stars878 forksPythonMIT

At a glance

What is it?
Mini-SGLang is a Python LLM inference framework of about 5,000 lines that implements the core optimizations of SGLang in a readable, modular codebase. It serves as both a working inference engine and a transparent reference for researchers who want to understand how modern LLM serving systems are built.
Who is it for?
Mini-SGLang is the right tool for engineers who want to understand LLM serving internals from a codebase small enough to read in full. It runs a working OpenAI-compatible server with meaningful throughput optimizations.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 135 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What mini-sglang is for and who should use it

Large language model serving systems like SGLang, vLLM, and TensorRT-LLM are production-grade but large. SGLang alone spans tens of thousands of lines across multiple subsystems. Engineers wanting to understand or modify one subsystem (say, the KV cache or the scheduler) face a large surface area before getting to the part they care about.

Mini-SGLang addresses this by reimplementing the core of SGLang in roughly 5,000 lines of Python. The README describes it as designed to demystify the complexities of modern LLM serving systems. The codebase is fully type-annotated and modular, which makes it easier to isolate any one component and trace how a request flows from the API endpoint to the model and back.

The pyproject.toml lists the project under development status Alpha, meaning the API and internal structure may change. It is not intended for production deployment but for researchers adding new optimizations, engineers learning how inference batching works, and educators who want a codebase they can walk through with students. A working OpenAI-compatible server is a byproduct of the same code, so experiments can run against real models.

Key optimizations: Radix Cache, chunked prefill, and overlap scheduling

Mini-SGLang implements several optimizations that matter for throughput and latency at scale. The README describes each one.

Radix Cache reuses the key-value cache across requests that share a common prefix. This is significant for multi-turn conversations, system prompts, or few-shot examples where many requests begin with identical tokens. Without prefix reuse, each request recomputes the attention for those tokens. With Radix Cache, the computation happens once and the cached activations are reused.

Chunked prefill breaks the prefill phase (processing the input tokens) into fixed-size chunks rather than processing the entire input in one pass. Long-context requests can consume most of the GPU memory in a single prefill step, which blocks other requests. Chunking the prefill reduces peak memory usage and allows the scheduler to interleave prefill and decode work.

Overlap scheduling hides the CPU overhead of request scheduling by running it in parallel with GPU computation. Without overlap, the GPU waits while the CPU decides which requests to batch next. The README notes that setting `MINISGL_DISABLE_OVERLAP_SCHEDULING=1` disables this for ablation studies.

Tensor parallelism splits model weights across multiple GPUs so that models too large for a single GPU can still run. Mini-SGLang supports this via the `--tp` flag.

Installing mini-sglang and starting the server

Mini-SGLang runs on Linux only, for x86_64 and aarch64. It does not support Windows or macOS because it depends on Linux-specific CUDA kernels. Windows users can use WSL2.

The README recommends `uv` for installation:

bash
git clone https://github.com/sgl-project/mini-sglang.git
cd mini-sglang && uv venv --python=3.12 && source .venv/bin/activate
uv pip install -e .

Python 3.10 or higher is required. The NVIDIA CUDA Toolkit must be installed and its version must match the driver version, which can be checked with `nvidia-smi`. The project uses JIT-compiled CUDA kernels from `sgl-kernel`, `flashinfer`, and `sgl_kernel`, so the CUDA environment must be complete.

Once installed, launching an OpenAI-compatible server for a small model is a single command:

bash
python -m minisgl --model "Qwen/Qwen3-0.6B"

For a larger model requiring multiple GPUs with tensor parallelism:

bash
python -m minisgl --model "meta-llama/Llama-3.1-70B-Instruct" --tp 4 --port 30000

This deploys the model across 4 GPUs and listens on port 30000. Standard curl or any OpenAI-compatible client can send requests to this endpoint.

Docker deployment and persistent caches

The repository includes a Dockerfile that builds a two-stage image using a CUDA 12.8.1 base. The builder stage installs Python, `uv`, and the project dependencies. The runtime stage runs as a non-root user named `minisgl`.

Running the Docker image requires the NVIDIA Container Toolkit. A basic server startup:

bash
docker run --gpus all -p 1919:1919 \
    minisgl --model Qwen/Qwen3-0.6B --host 0.0.0.0

For faster subsequent startups, the README recommends using Docker volumes for the HuggingFace model cache, the TVM FFI cache, and the FlashInfer JIT compilation cache:

bash
docker run --gpus all -p 1919:1919 \
    -v huggingface_cache:/app/.cache/huggingface \
    -v tvm_cache:/app/.cache/tvm-ffi \
    -v flashinfer_cache:/app/.cache/flashinfer \
    minisgl --model Qwen/Qwen3-0.6B --host 0.0.0.0

Without persistent caches, the container recompiles the CUDA kernels and downloads the model on every restart, which adds significant startup time. The Dockerfile sets `HF_HOME`, `TVM_FFI_CACHE_DIR`, and `FLASHINFER_WO` environment variables to point to the cache directories inside the container.

Platform constraints and what mini-sglang cannot do

The Linux-only restriction is the most significant constraint. The README is explicit that Windows and macOS are not supported because `sgl-kernel` and `flashinfer` require Linux-specific CUDA kernels. ARM-based Macs and any machine without an NVIDIA GPU cannot run mini-sglang natively.

The pyproject.toml requires `torch<2.10.0`, which pins the PyTorch version. This means mini-sglang cannot immediately benefit from PyTorch releases at or above 2.10.0 without a dependency update. The same file requires `transformers>=4.56.0,<=4.57.3`, another narrow version range that constrains which model architectures the project can load.

Mini-SGLang does not implement all features of full SGLang. The README positions it as a compact implementation rather than a feature-complete one. Teams needing speculative decoding, vision-language model support, or the full range of quantization backends should use SGLang or vLLM directly. The architecture document at `docs/structures.md` describes the scope of what is implemented.

Mini-SGLang versus SGLang and vLLM

Full SGLang is the production system that mini-sglang is distilled from. SGLang includes the same Radix Cache, chunked prefill, and overlap scheduling, plus a much larger feature set. The benchmark commands in the README show both systems being tested with the same model and configuration, which confirms that mini-sglang is a direct implementation of the same core algorithms rather than a simplified approximation.

The benchmark setup in the README uses a 4xH200 GPU configuration connected by NVLink with the Qwen3-32B model. The mini-sglang launch command for that benchmark is:

bash
python -m minisgl --model "Qwen/Qwen3-32B" --tp 4 --cache naive

Compared to the full SGLang command, mini-sglang uses `--cache naive` rather than the default Radix Cache, which is the ablation condition for the benchmark. The `--model-source modelscope` flag provides an alternative download source for users with restricted HuggingFace access.

vLLM is another popular LLM inference engine with a different architecture and a larger feature set than mini-sglang. vLLM is also not a learning resource: it has a larger codebase and no specific focus on readability. For the purpose of understanding LLM serving, mini-sglang is deliberately the more accessible starting point.

Editorial conclusion

Mini-SGLang is the right tool for engineers who want to understand LLM serving internals from a codebase small enough to read in full. It runs a working OpenAI-compatible server with meaningful throughput optimizations. It is not a replacement for production SGLang: the README explicitly positions it as a reference, and the pyproject.toml lists its development status as Alpha. Teams building production inference infrastructure should start with SGLang or vLLM. Those learning how LLM serving works, or implementing research modifications, should clone mini-sglang and read the architecture document at docs/structures.md first.

Frequently asked questions

Does mini-sglang support Windows or macOS?

The README states that mini-sglang supports Linux only (x86_64 and aarch64). Windows and macOS are not supported because the required CUDA kernels from sgl-kernel and flashinfer are Linux-specific. The README suggests WSL2 on Windows or Docker as alternatives.

What is Radix Cache in mini-sglang?

Radix Cache is an optimization that reuses the KV cache for shared prefixes across requests. When multiple requests begin with the same system prompt or few-shot examples, the attention computation for those tokens is done once and cached, reducing computation for subsequent requests.

How does mini-sglang differ from full SGLang?

Mini-SGLang is a compact reimplementation of SGLang's core in about 5,000 lines of Python, designed for readability and learning. Full SGLang is the production system with a larger feature set. The pyproject.toml lists mini-sglang's development status as Alpha and it is not intended for production deployment.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. sgl-project/mini-sglang on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sgl-project-mini-sglang.svg)](https://hysenlabs.com/projects/sgl-project-mini-sglang)
Community notes

Community notes