Model or dataset
Wenyueh/MinivLLM avatar
Wenyueh/MinivLLM

minivLLM's demo runs a randomly initialized Qwen3 on two prompts repeated thirty times

Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation

1,059 stars177 forksPythonApache-2.0

At a glance

What is it?
Wenyueh/MinivLLM reimplements vLLM's paged and flash attention as a teaching project, and its pyproject declares a dependency on vllm>=0.15.0 that the requirements list leaves out. Setup.py, pyproject, and the README each state a different Python version and a different dependency set.
Who is it for?
minivLLM fits someone learning how an inference engine is put together, or someone who wants a readable paged attention and block scheduling implementation to read next to a real one. Three things to check before you build on it.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 43 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The main demo builds a small Qwen3 with random initialization

The quickstart runs five commands:

bash
# Install uv package manager
curl -LsSf https://astral.sh/uv/install.sh | sh

# Sync dependencies
uv sync

# Run the main inference engine
uv run python main.py

# Run prefilling benchmark
uv run python benchmark_prefilling.py

# Run decoding benchmark
uv run python benchmark_decoding.py

What main.py actually does is worth stating plainly, because it decides how to read every number the project produces. It creates a small version of Qwen3 with random initialization. It creates 60 chat prompts, which are 2 base prompts repeated 30 times each. It processes them through the engine with batch processing, uses paged attention and KV cache management, and generates up to 256 tokens per prompt with temperature sampling.

So the weights are noise and the prompts are two. Nothing here measures model quality, and nothing here measures a diverse generation distribution either. What it does exercise is the engine path: batching, block allocation, the scheduler, prefill, and decode. Read the timings as engine timings, which is the point of the project, and not as a serving benchmark.

pyproject declares vllm>=0.15.0, and the requirements list does not mention it

The stated requirements are Python from 3.11 up to but not including 3.12, a CUDA capable GPU, and three dependencies named transformers, torch, and xxhash, managed by uv.

The project metadata lists four. Alongside transformers, torch, and xxhash it declares vllm with a floor of 0.15.0.

That fourth entry is the interesting one, because this repository describes itself as a custom implementation of the vLLM inference engine with its own paged attention and flash attention. Installing it therefore pulls in the engine it is replicating. Nothing in the documentation says what that dependency is for: it is not named as a comparison baseline in the benchmark sections, which compare PyTorch and Triton implementations against each other, and it is not named as a source of schemas or kernels. No explanation appears.

Two consequences. A uv sync in this repository installs a full inference stack, which is a different act from installing a study of one. And anyone benchmarking against real vLLM has it available, whether or not the repository intends that.

setup.py and pyproject disagree about Python and about dependencies

Two build paths exist. pyproject.toml uses the setuptools backend, sets package-dir to src, and finds packages under src. setup.py passes the same layout to setup() by hand.

They do not agree on the interpreter. pyproject sets requires-python to 3.11 or newer and below 3.12. setup.py sets python_requires to exactly 3.11.14. So the metadata path accepts any 3.11 release while the setup.py path accepts one patch version. A machine on 3.11.9 satisfies the first and is rejected by the second.

They do not agree on dependencies either. pyproject lists transformers, torch, xxhash, and vllm. setup.py passes a single requirement, torch. An install driven by setup.py gets torch and nothing else, so a script that imports transformers or xxhash fails on that path while working under uv sync.

Both declare the distribution name as myvllm at version 0.1.0, with the description A custom vLLM project, so the package identity is stable even when the metadata is not. uv.lock is committed, which means the uv path has a resolvable answer even though the setuptools path does not.

The dev tool set is declared twice with different floors

The development dependencies appear in two separate places in pyproject.toml, and they are not identical.

The first is an optional extras group named dev, holding pytest, black, and isort with no version constraints at all. The second is a dependency group also named dev, holding pytest at 7.0 or newer, black at 23.0 or newer, and isort at 5.0 or newer.

Two blocks with the same name and different contents is a duplication that tools handle differently. pip understands the optional extras group and ignores the dependency group unless it is installed as a group. uv reads dependency groups. So depending on which installer you use, you can get an unpinned or a floored set of the same three tools, and if you use both, the unpinned extras group is the one that wins.

This is a small thing in a repository whose subject is a GPU inference engine, but it is a fair sample of how the packaging metadata was assembled: copied between files rather than derived from one source.

Five Python files at the root, and the script guide covers three

The top level of the repository holds main.py, main_llama32.py, benchmark_prefilling.py, benchmark_decoding.py, and benchmark_tps.py, alongside HowToApproachvLLM.md, its Chinese counterpart, two READMEs, assets/, src/, tests/, uv.lock, pyproject.toml, and setup.py.

The section titled What Each Script Does walks through three of them: main.py, benchmark_prefilling.py, and benchmark_decoding.py. The other two are absent from it. main_llama32.py is a second engine entry point with a name that says which model family it targets, and benchmark_tps.py is a third benchmark, whose name says tokens per second.

That is the shape of the gap. The tokens per second measurement is the number most people want from an inference engine, and it is the one the guide does not explain. The second model entry point is the one that would let you compare against something other than the randomly initialized small Qwen3, and it is undocumented beyond its filename.

There is also a tests/ directory and no test command anywhere in the quickstart or the script guide.

The prefilling baseline is capped at 128 tokens by shared memory

The prefilling benchmark compares three attention implementations over the phase that processes input prompts.

The PyTorch standard version materializes the full attention matrix and uses memory that grows with the square of the sequence length. A naive Triton kernel does the same quadratic work but is limited by shared memory constraints to 128 tokens. The flash attention version uses linear memory, computing an online softmax and processing attention in blocks.

The number to hold onto is 128. One of the three kernels has a hard ceiling at that sequence length, and it is the one the others are being compared against. Any measurement of how much flash attention wins over the naive path is bounded by the naive path's own limit, so the interesting comparisons for long prompts are PyTorch standard against flash attention, with the Triton kernel dropping out of the range where the quadratic approaches hurt most.

The decoding benchmark is a different comparison and is not memory-shape bound. It compares a loop based PyTorch implementation using a paged KV cache, a vectorized PyTorch implementation with batch gathering and masking, and a custom Triton kernel written for paged attention decode.

The structure block names myvllm/ as the root and ends mid-word

The project structure is drawn as a tree rooted at myvllm/, containing src/myvllm/ with models/, engine/, layers/, and utils/.

So the distribution name, the package directory, and the drawn root are the same word, myvllm, while the repository is called MinivLLM. That is a deliberate and reasonable choice for a package, and it is worth naming because a clone lands in a directory called MinivLLM while the import name and the tree root are myvllm.

The engine/ entry is described in detail and is the most useful part of the block: sequence definition for input prompts, block management for KV cache management for GPU, a scheduler for iteration based scheduling of sequences, a runner for the actual implementation of running prefilling and decoding, and an engine for the generation API interface. That is a map of the engine worth reading before the code.

The block does stop partway through the last line. The utils/ entry reads Utility helpers and inference context manageme, and the word is cut off there, with no further entries shown after it.

Editorial conclusion

minivLLM fits someone learning how an inference engine is put together, or someone who wants a readable paged attention and block scheduling implementation to read next to a real one. Three things to check before you build on it. Whether the numbers from benchmark_prefilling.py and benchmark_decoding.py matter to you, since the model has random weights and the prompt set is two base prompts repeated thirty times. Which Python pin you are held to, because pyproject allows all of 3.11 while setup.py names one patch release. And why vllm>=0.15.0 is in your dependency tree at all, since nothing in the documentation says what it is for. Anyone wanting production throughput numbers or model quality from this repository should look at vLLM itself.

Frequently asked questions

What is minivLLM, and what does it implement itself?

It is a custom implementation of the vLLM inference engine, based on Nano-vLLM, with self-contained paged attention and flash attention. It ships benchmark scripts for flash attention during prefilling and paged attention during decoding.

Does minivLLM download real model weights for its demo?

No. main.py creates a small version of Qwen3 with random initialization and builds 60 chat prompts from 2 base prompts repeated 30 times each, generating up to 256 tokens per prompt with temperature sampling. The output measures the engine path rather than model quality.

How do I set up minivLLM locally?

Install uv with curl -LsSf https://astral.sh/uv/install.sh | sh, then run uv sync, then uv run python main.py. The benchmarks are uv run python benchmark_prefilling.py and uv run python benchmark_decoding.py. For multi-GPU, change world_size to a value above 1 in the config in main.py.

What Python version does minivLLM require?

pyproject.toml requires Python 3.11 or newer and below 3.12, and the requirements say the same. setup.py differs, setting python_requires to exactly 3.11.14. A CUDA capable GPU is also required.

Which attention implementations does minivLLM compare?

For prefilling it compares PyTorch standard with quadratic memory, a naive Triton kernel also with quadratic memory that is limited by shared memory to 128 tokens, and flash attention with linear memory using an online softmax. For decoding it compares a loop based PyTorch version, a vectorized PyTorch version, and a Triton paged attention kernel.

Does minivLLM depend on vLLM itself?

pyproject.toml declares vllm>=0.15.0 among its dependencies, alongside transformers, torch, and xxhash. The requirements section names only the latter three, and the documentation does not say what the vLLM dependency is used for.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Wenyueh/MinivLLM on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/wenyueh-minivllm.svg)](https://hysenlabs.com/projects/wenyueh-minivllm)