tiny-vllm: a C++ and CUDA inference engine you build yourself
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
At a glance
- What is it?
- tiny-vllm is a teaching implementation of an LLM inference server in C++ and CUDA, shipped with a full course that derives each kernel from scratch. It is a learning resource first and a production server second.
- Who is it for?
- Adopt tiny-vllm if you want to read and modify a complete inference engine in C++ and CUDA, or if you teach a course that needs one. Do not adopt it as a serving layer for production traffic: the README frames it as a younger sibling of vLLM and as a learning tool, and it documents no HTTP API, no deployment story and no throughput numbers.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What tiny-vllm is for, and who it is not for
The README states the goal plainly: you build a high performance LLM inference engine in C++ and CUDA, described as a younger and smaller sibling of vLLM. The repository holds two things, a full source tree for the inference server and a course that walks through implementing it. The stated audience is people on a learning path and lecturers who want a teaching resource.
The problem it addresses is a gap in how inference engines are usually encountered. A production engine like vLLM is large enough that reading it end to end is impractical, and the papers behind its techniques (FlashAttention, PagedAttention) assume familiarity with the systems layer underneath. tiny-vllm closes that gap by implementing one model architecture completely, in a codebase sized for a course.
If you want a server to put behind a load balancer, this is the wrong project. The README never describes an HTTP endpoint, a client library, authentication, or a deployment artifact. It describes an engine and the lessons that produce it.
The engine pipeline: Safetensors to sampled token
The feature list in the README is a checklist of what the engine does, and it reads as a data flow. A real model is loaded from Safetensors, specifically Llama 3.2 1B Instruct. The full forward pass runs in two phases, prefill and decode. All computation happens in CUDA kernels. A KV cache stores attention keys and values across steps, which is what makes decode cheap relative to recomputing the whole sequence.
On top of that sit the scheduling and attention optimizations. Static batching comes first, then continuous batching, which lets requests join and leave a running batch instead of waiting for a fixed group to finish. Attention uses an online softmax formulation in the style of FlashAttention, and the KV cache is paged, with a PagedAttention CUDA kernel to read it. The course table of contents mirrors this order, moving from floating point representation and GPU versus CPU memory, through embeddings, RMSNorm, RoPE, residual connections and cublasGemmEx, then into prefill versus decode, KV cache, attention, GQA, softmax, causal mask, argmax and the feed forward network, before reaching batching and paging.
The README also notes a column-major to row-major transposition trick and a section on buffer reuse, which are the kind of details that separate a working kernel from a fast one. It does not publish latency or throughput figures for any of this.
Building tiny-vllm and running a first inference
The repository root contains CMakeLists.txt, build.sh, run.sh, test.sh, check.sh and full_test.sh, plus src/, include/, python/ and tests/. The README does not spell out a step by step install, so the build path is the scripts and the CMake project rather than a documented command sequence. Start by checking the CMake configuration, since that is where the CUDA toolkit and architecture requirements will be expressed.
cat CMakeLists.txtReading that file tells you which CUDA version and compute capability the project expects before you spend time compiling. Once the requirements look satisfiable, the build script is the entry point.
./build.shThe README does not document the arguments build.sh accepts or the directory it writes artifacts to, so treat the script itself as the reference. The same applies to running: run.sh is the launcher present at the root, and the README does not describe its flags or its expected model path.
./run.shThe model the README names as loadable is Llama 3.2 1B Instruct in Safetensors format, so any weights you point the engine at should match that architecture. For verification there are three more scripts, test.sh, check.sh and full_test.sh, and a tests/ directory. The README does not describe what each covers, so running full_test.sh is the broadest signal available without reading the test sources.
Where tiny-vllm stops short
The most concrete limitation is the model surface. The README names exactly one model, Llama 3.2 1B Instruct, loaded from Safetensors. There is no mention of a general architecture registry, of quantization formats, or of loading arbitrary Hugging Face checkpoints. If your work involves Mistral, Qwen, or a quantized variant, nothing in the README suggests tiny-vllm will load it.
The second limitation is operational. Continuous batching and PagedAttention are implemented, but an inference server in the deployment sense needs more than that: request admission, cancellation, timeouts, metrics, and a network protocol. The README describes none of these, and the top level entries contain no server configuration files or container definitions. The homepage field is empty, so there is no hosted documentation to fall back on.
The third is scope by design. This is a course. Sections are ordered for learning, not for minimizing time to a running service. A reader who wants a working endpoint today will spend that day on floating point representation and RMSNorm instead. That is the point of the project, and it is also the reason it does not substitute for vLLM in a deployment.
How tiny-vllm differs from vLLM and from a Python reference
The README positions tiny-vllm explicitly as a smaller sibling of vLLM, and the checklist overlaps with what vLLM is known for: continuous batching and PagedAttention both appear in both projects. The difference is the boundary of the codebase. vLLM is a serving system with a broad model zoo and a Python-centric extension path. tiny-vllm is C++ and CUDA end to end, with every kernel written out and explained, and with a single supported model. If you want to serve many models to many users, vLLM is the tool the README itself points at. If you want to understand why PagedAttention exists, by writing the kernel, tiny-vllm is the one that shows you.
The other natural comparison is a Python reference implementation of a transformer, the kind people write to learn attention. Those are readable but they do not teach the systems layer: no KV cache management, no batching scheduler, no memory layout decisions, no kernel engineering. tiny-vllm covers cublasGemmEx, the column-major to row-major transposition, buffer reuse and parallel reduction in CUDA, which are exactly the parts a Python reference omits. The trade is that you cannot run a Python reference at useful speed, and you cannot run tiny-vllm without a CUDA-capable GPU and a build toolchain.
Maintenance, licence and what upgrading costs
The repository is not archived, and the last push was on 2026-08-23, which is recent. There are no releases, so there is no versioned artifact to pin and no changelog describing breaking changes. Upgrading, in practice, means pulling the main branch and rebuilding, then running the test scripts to see whether your environment still matches. Because the project is a course, the sections and the source move together; a change to a kernel is likely to come with a change to the chapter that explains it, which is helpful for readers and awkward for anyone treating the code as a stable dependency.
The licence is Apache-2.0, which is a permissive licence that generally allows use, modification and distribution, including in commercial settings, provided the licence terms are met. This is not legal advice. The practical implication for a reader is that you can lift a kernel or a design idea into your own project, but you should read the LICENSE file at the root for the exact terms, including the patent grant and the notice requirements. Note that the model weights you load are a separate matter: Llama 3.2 1B Instruct ships under its own terms from its distributor, and the Apache-2.0 licence on this repository does not extend to those weights.
Editorial conclusion
Adopt tiny-vllm if you want to read and modify a complete inference engine in C++ and CUDA, or if you teach a course that needs one. Do not adopt it as a serving layer for production traffic: the README frames it as a younger sibling of vLLM and as a learning tool, and it documents no HTTP API, no deployment story and no throughput numbers. Before you start, verify that your GPU and CUDA toolkit satisfy what the CMakeLists.txt and build.sh expect, and confirm that you can obtain Llama 3.2 1B Instruct in Safetensors format, since that is the only model the README names as loadable.
Frequently asked questions
What is tiny-vllm used for?
The README describes it as a learning tool and a teaching resource: a full C++ and CUDA inference server plus a course that walks through building it. It is meant for people who want to understand how an LLM inference engine works from the kernels up.
Can I use tiny-vllm in C++?
Yes. The engine itself is written in C++ and CUDA, with sources under src/ and headers under include/, built through the CMakeLists.txt at the repository root. There is also a python/ directory, but the README presents the inference server as a C++ and CUDA project.
Which model does tiny-vllm load?
The README names Llama 3.2 1B Instruct, loaded from Safetensors format. It does not describe support for any other architecture or for quantized checkpoints.
Does tiny-vllm need a CUDA GPU?
Yes. The README states that all computation runs in CUDA kernels, and the course covers CUDA kernel engineering directly. The CMakeLists.txt and build.sh are where the toolkit requirements are expressed, and the README does not document a CPU-only path.
How do I build and run tiny-vllm?
The repository root contains build.sh and run.sh, along with test.sh, check.sh and full_test.sh for verification. The README does not document their arguments or output locations, so the scripts and CMakeLists.txt are the reference.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jmaczan-tiny-vllm)