# PegaInfer runs on Rust and CUDA, and still inherits NVIDIA Dynamo's package metadata

> Nineteen Rust crates carry the serving stack, a pinned dynamo commit and a local kvbm fork sit underneath it, and the README's own numbers change shape from panel to panel.

**pegainfer-project/pegainfer** — Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

- Repository: https://github.com/pegainfer-project/pegainfer
- Website: https://pegainfer.org/
- Stars: 715 · Forks: 111
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/pegainfer-project-pegainfer

## The workspace package block still reads Dynamo Inference Framework

The root Cargo.toml gives the workspace a `[workspace.package]` block that belongs to NVIDIA Dynamo rather than to PegaInfer. It sets `authors` to NVIDIA Inc. with the `sw-dl-dynamo@nvidia.com` contact address, `description` to `Dynamo Inference Framework`, `homepage` and `repository` to the `ai-dynamo/dynamo` project, `keywords` to llm, genai, inference, nvidia and distributed, and `version` to `1.2.0`. A comment above the block explains that these values are inherited by dynamo-ported crates that pull fields through `edition.workspace = true`.

That inheritance is the part worth watching. Any crate that takes its fields from the workspace table publishes carrying Dynamo's description, contact address and homepage, while the repository record describes the project as a pure Rust and CUDA inference engine with an OpenAI-compatible API. Three notice files sit at the root next to the project's own LICENSE: NOTICE, NOTICE_DYNAMO and THIRD_PARTY_LICENSES_DYNAMO.txt. Apache-2.0 is the only license named in the workspace metadata and the only one the repository record reports, so there is no second license here to reconcile, and the upstream attribution is left explicit rather than folded into the project name.

## Two workspace dependencies are pinned to one dynamo commit

The upstream dependencies are not version ranges. `dynamo-kv-hashing` and `dynamo-tokens` both come from `git = "https://github.com/ai-dynamo/dynamo"` at `rev = "364cc8aa543d97f0563e17de3069d98ea051f33f"`. A comment in the same table says every ai-dynamo/dynamo dependency has to stay on that same rev, because crates from one git checkout unify their shared internal deps through `dynamo-tokens`, and mixing revs would split type identity.

The third entry works differently. `kvbm-logical` is a path dependency on `kvbm/kvbm-logical`, marked in the members list as a fork of the upstream kvbm under Apache-2.0, so the workspace mixes one pinned external commit with a vendored fork held inside the tree.

Past those three, the dependency table is conventional: `axum` for the HTTP layer, `async-nats` at 0.45.0 with the service feature, `anyhow`, `async-stream`, `async-trait`, `bincode` at 2.0.1, `bindgen` at 0.72.1 for generating bindings, and a `bytes` entry declared at 1.10.1. The table ends mid-entry on that `bytes` declaration, with its feature list only half written, so the rest of the third-party set is not visible here. The last recorded push for the repository is 2026-09-30, which is newer than both published releases, so the tree has moved past the 0.1.1 artifact.

The scale explains why the pinning rule matters. The members list holds nineteen `pegainfer-*` crates, from `pegainfer-core`, `pegainfer-kernels` and `pegainfer-server` through one crate per model line: `pegainfer-qwen3`, `pegainfer-qwen35`, `pegainfer-gemma4`, `pegainfer-glm52`, `pegainfer-k3`, `pegainfer-kimi-k2` and `pegainfer-deepseek-v2-lite`. `default-members` is `["pegainfer-server"]`, the resolver is `3` and the workspace edition is `2024`. A clean build resolves that external commit before it compiles a server that ships a single model line by default.

## Release 0.1.0 shipped under a different project name

The release list carries two tags and two different names. `v0.1.0` is titled `OpenInfer 0.1.0` and dates from 2026-06-13; `v0.1.1` is titled `PegaInfer v0.1.1` and dates from 2026-08-26. The rename sits between the two releases.

Not every reference followed it. The Slack badge near the top of the README still points at a join.slack.com invite under the `t/openinferhq` path, so the community link a reader follows from the current README belongs to the older name. The last recorded push for the repository is 2026-09-30, after both tags.

This shows up at install time rather than in the source. The installer selects the latest release by default, so a fresh install gets the 0.1.1 artifact named PegaInfer. `PEGAINFER_VERSION` selects an exact version instead, and it is the only lever for reaching the 0.1.0 build. Anyone pinning a version to match an older tutorial or an existing deployment script is pinning the artifact that still carries the OpenInfer name, which also means the older README instructions for that release refer to the same installer path.

## The install path is a curl pipe and the binary floor is driver 580

The documented install for the prebuilt binary is one line that pipes a remote script into a shell:

```bash
curl -fsSL https://raw.githubusercontent.com/pegainfer-project/pegainfer/main/install.sh | bash
```

Weights are a separate step. The README asks you to download Qwen3-4B into `models/Qwen3-4B`, so that directory has to exist before the server starts:

```bash
pegainfer --model-path models/Qwen3-4B
```

The server listens on port 8000, and if the command is not on your shell's path:

```bash
export PATH="$HOME/.local/bin:$PATH"
```

The Qwen3-only release bundles CUDA 13 and cuBLAS. It requires Linux x86_64, an NVIDIA GPU with compute capability 8.x to 12.x, driver 580 or newer, glibc 2.35 or newer and OpenSSL 3. A source build states a much lower floor, R545 with CUDA 12.3, and notes that newer toolkits and model-specific kernels can require a newer driver. A card that satisfies the source build floor does not necessarily satisfy the prebuilt binary. The install path itself offers no checksum and no signature step before bash runs the script; `PEGAINFER_VERSION` chooses which version lands on disk, not whether its contents can be verified.

## The Windows Qwen3.5 example installs Triton and then sets the TileLang variable

Qwen3.5 uses Triton AOT kernels and K3 kernel generation goes through TileLang, so both model lines need a Python interpreter at build time even though the engine itself is Rust and CUDA. On Linux the Qwen3.5 path is a venv, Triton and the interpreter variable:

```bash
uv venv
uv pip install triton
export PEGAINFER_TRITON_PYTHON=.venv/bin/python
cargo run --release --features qwen35 -- --model-path models/Qwen3.5-4B
```

The Windows path is where the two variables collide. That example installs Triton, pins a Windows-specific build of it below version 3.7, and then assigns the TileLang variable rather than the Triton one:

```powershell
uv venv .venv --python 3.12
uv pip install "triton-windows<3.7"
$env:PEGAINFER_TILELANG_PYTHON = ".venv\Scripts\python.exe"
cargo run --release --features qwen35 -- --model-path models/Qwen3.5-4B
```

Read literally, a Qwen3.5 build on Windows is handed the TileLang interpreter with no Triton interpreter set anywhere in the block. The environment table documents the two variables separately, `PEGAINFER_TRITON_PYTHON` for Qwen3.5 Triton AOT compilation and `PEGAINFER_TILELANG_PYTHON` for K3, so they are not interchangeable by design. Treat that block as needing a correction before it works on Windows.

## Only qwen3 is compiled in, and the model table ends inside the Qwen3.8 row

Only `qwen3` is enabled by default, and the prebuilt binary ships with that same single feature. Other model lines require `--features <feature>` at build time. At launch `--model-path` selects a checkpoint and its `config.json` identifies the family, so the family is decided by the checkpoint rather than by a command-line switch.

The table that enumerates the lines names two of them completely. Qwen3 covers dense 0.6B to 32B with full attention and GQA, serving greedy and sampling, tensor parallel, prefix cache and KV offload, with DFlash and DSpark limited to the 4B size. Qwen3.5 covers dense 0.8B to 27B on gated DeltaNet plus full attention, text-only BF16 with build-time Triton, and it carries Qwen3.8-27B on the same feature.

The table then ends inside that last row. The link text for the Qwen3.8 support record stops partway through a docs path, and no row for Kimi-K2, K3, GLM or Gemma follows, even though the project description says the engine serves Qwen3 to Kimi-K2 and the workspace carries `pegainfer-kimi-k2`, `pegainfer-k3`, `pegainfer-glm52` and `pegainfer-gemma4` crates. What those lines support in terms of attention style, precision or parallelism is not written down here.

## Every performance panel changes the hardware, the revision and the comparison

The performance section splits its numbers into four panels, and no two of them share a setup. The Qwen3 4B panel is a speculative decoding comparison against DSpark on single-request greedy decoding over ShareGPT and SPEED-Bench coding. The Qwen3.5 9B and 27B panel is a GH200 concurrency sweep at revision `ffb959c4`, using random 1,024-token prompts and 128-token outputs. The GLM-5.2 panel contrasts a co-located EP4 setup on 4 GPUs with a disaggregated TP4 prefill plus EP4 decode on 8 GPUs total. The Gemma 4 26B-A4B panel derives ratios from reported median end-to-end latencies, and its source is a comment on issue 758 rather than a document in the tree.

The detail panel for Qwen3 is the one with a fixed side-by-side: one RTX 5090, BF16, TP1, PegaInfer at revision `70888b2` against vLLM 0.24.0. Resident memory loaded and idle reads 771 MB against 3814 MB, cold startup to HTTP ready 2.99 s against 70.0 s, and warm compile cache startup about 3.0 s against 32.7 s. Three caveats sit in the same paragraph: the vLLM side carries a version with no revision, its figure sums a process tree while PegaInfer runs as one process, and the run is separate from the DSpark panel. The Gemma panel repeats the pattern, using revision `e7a41975` for its BF16 comparison and `ea02a9f7` for the default-FP8 one, with BF16 KV in both of its own comparisons.

## The repository root carries four agent-tool directories beside the Rust build files

Next to Cargo.toml, Cargo.lock and LICENSE, the root of the tree holds `.agents/`, `.claude/`, `.codex/` and `.cursor/`, with AGENTS.md and CLAUDE.md at the top level. Alongside them sit configuration files whose names indicate their job: `.pre-commit-config.yaml`, `rustfmt.toml`, `taplo.toml`, `typos.toml`, `.editorconfig`, `.gitignore`, `.dockerignore` and `.gitmodules`. No visible text ties the four agent-tool directories to a documented workflow, and the root listing gives no hint which one is authoritative.

The rest of the tree follows the crate layout. `scripts/`, `tools/`, `deploy/`, `docker/`, `test_data/` and `tests/` hold tooling and fixtures, and `docs/` holds the per-model reports the performance section links to, including the Qwen3 serving report behind the footprint figures. `pegainfer-server` is the entrypoint, model crates carry the model implementation and diagnostics, and `cargo run --release -- --help` prints the CLI that got compiled in, which differs depending on which features you built. `install.sh` sits at the root as well, which is the file the documented install command fetches.

## Conclusion

PegaInfer is worth evaluating if you serve Qwen3 on Linux x86_64 with a recent driver and want an OpenAI-compatible endpoint from one Rust process. Two things to check first. Read Cargo.toml: the workspace package metadata still carries NVIDIA Dynamo's description, contact address and homepage, and two dependencies are pinned to one external dynamo commit, so registry metadata will not tell you which project you installed. Then match your hardware and driver to the artifact you actually plan to run, since the prebuilt binary asks for driver 580 while a source build states a floor of R545, and the Qwen3.5 and K3 lines add Python, Triton and TileLang at build time. The performance panels work as a map of the work, not as a comparison you can quote elsewhere.

## FAQ

### Does PegaInfer need Python or PyTorch at runtime?

Not for the default Qwen3 build, which the README says needs no Python, and no PyTorch anywhere in the serving stack. Python comes back at build time for other feature lines: Qwen3.5 uses Triton AOT kernels and reads the interpreter from PEGAINFER_TRITON_PYTHON, while K3 reads PEGAINFER_TILELANG_PYTHON for TileLang kernel generation.

### What hardware does the prebuilt PegaInfer binary require?

The Qwen3-only release bundles CUDA 13 and cuBLAS and requires Linux x86_64, an NVIDIA GPU with compute capability 8.x to 12.x, driver 580 or newer, glibc 2.35 or newer and OpenSSL 3. Model weights are not included and are downloaded separately into models/Qwen3-4B.

### Which models can PegaInfer serve out of the box?

Only the qwen3 feature is enabled by default, and the prebuilt binary ships with it. Other lines are reached with --features at build time. The model table describes Qwen3 dense 0.6B to 32B and Qwen3.5 dense 0.8B to 27B, with Qwen3.8-27B on the same feature, and stops partway through the Qwen3.8 row, so the remaining lines are not enumerated there.

### How is PegaInfer licensed and what does it borrow from upstream?

Apache-2.0, named both in the workspace package metadata and by the repository record, with NOTICE, NOTICE_DYNAMO and THIRD_PARTY_LICENSES_DYNAMO.txt at the root. Two workspace dependencies come from ai-dynamo/dynamo at the pinned rev 364cc8aa543d97f0563e17de3069d98ea051f33f, and kvbm-logical is a local path fork of that project's kvbm.

### How does PegaInfer compare to vLLM on memory and startup time?

One measurement is given, for Qwen3-4B on a single RTX 5090 in BF16 with TP1, PegaInfer revision 70888b2 against vLLM 0.24.0: 771 MB against 3814 MB resident when loaded and idle, 2.99 s against 70.0 s cold startup to HTTP ready, and about 3.0 s against 32.7 s with a warm compile cache. The vLLM figure sums its process tree, and this run is separate from the DSpark panel.

## Sources

- [License: Apache-2.0](https://github.com/pegainfer-project/pegainfer/blob/main/LICENSE)
- [pegainfer-project/pegainfer on GitHub](https://github.com/pegainfer-project/pegainfer)
- [Project website](https://pegainfer.org/)
- [README](https://github.com/pegainfer-project/pegainfer/blob/main/README.md)
- [Releases](https://github.com/pegainfer-project/pegainfer/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pegainfer-project-pegainfer
