Model or dataset
weicj/vLLM-2080Ti-Definitive avatar
weicj/vLLM-2080Ti-Definitive

The default branch is the line the README calls unmaintained

The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-request decode with support of FP8 weight ( Join Discord :https://discord.gg/VFqVVySdMS )

1,084 stars148 forksPythonApache-2.0

At a glance

What is it?
This is a vLLM fork aimed at dual RTX 2080 Ti and other SM75 cards, with launcher profiles, four launch modes and its own baseline pinned to an upstream commit. Two things in the repository contradict the marketing: the default branch is the legacy line, and the description's throughput claim is less than half the table's.
Who is it for?
Judgment: this fork is a serious piece of hardware enablement, and its value is concentrated in things a general vLLM install cannot give you: preserved SM75 source changes, a validated profile per hardware layout, and a launcher that knows which topology was actually benchmarked. Read it that way rather than as a faster vLLM.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The default branch is the legacy line, and the docs point elsewhere

The repository's default branch is named `vllm-2080ti-definitive-0.1.x`, while the README's own branch link points at `main` and the support section describes the 0.1.x line as no longer actively maintained. So a plain clone lands on the older line while every baseline, profile and instruction in the document describes 0.2.x. The version story inside the 0.2 line is a rolling baseline rather than a fixed one: the current baseline is v0.2.2-post3, and the release history shows three post-releases of the same base version inside a week, v0.2.2-post1 on 25 September, post2 on 26 September and post3 on 1 October 2026. Alongside that sits an upstream pin, commit b23433088b described as v0.29.1rc0-33, which is what makes the fork reproducible: a reader can say exactly which upstream tree the SM75 changes were applied to.

The description claims 100+ tok/s and the table claims 220

The project description says the runtime delivers maximum 100+ tok/s single-request decode, while the highlights table reports 220.84 tok/s for a 4K prompt and 209.35 tok/s for a 32K prompt on two RTX 2080 Ti cards with Qwen3.8 27B in NVFP4 and an FP8 KV cache at 256K context. The second row is four Tesla T10 cards with the same model in FP8 and an FP16 KV cache, at 191.89 and 189.38 tok/s. So the description appears to describe an older or different configuration rather than the current table, and anyone quoting a throughput number should quote the table with its configuration attached. The table also carries its own caveat directly underneath, which is the more important line: both rows are single-request tests using DFlash2 with the default K=7, on synthetic text inputs with high speculative-hit rates, and real-task throughput depends on draft acceptance and may not reach these figures. That is an unusually direct admission that the headline numbers are an upper bound set by the draft model's accuracy.

The comparison table checks out, and bandwidth is the weakest ratio

The argument for the hardware is a table against a single RTX 3090 Ti, and the arithmetic is internally consistent. Two 2080 Ti cards give 8,704 dedicated FP32 datapaths against 5,376, a 1.62x ratio, and 136 SMs against 84, the same 1.62x, which follows from both cards sharing the same per-SM count. Tensor Cores are the outlier at 1,088 against 336, a 3.24x ratio, the largest margin in the table. Dense FP16 matrix throughput is 228 TFLOPS against 160, or 1.43x, and total VRAM is 44 GB against 24 GB at 1.83x. Total memory bandwidth is the thinnest claim, 1,232 GB/s against 1,008, only 1.22x, and it is the number that matters most for single-request decode at long context. Read together, the table argues convincingly for capacity and tensor throughput and much less strongly for the bandwidth that decoding actually depends on.

The environment floor moved, and the old line is the fallback

The 0.2.x target environment is Ubuntu 26.04 or later, Linux kernel 7 or later, GCC/G++ 15, CUDA 13.0, PyTorch 2.13 and Python 3.12. Everything older is redirected to the 0.1.x line, with CUDA 12.8, PyTorch 2.11, older kernels and GCC 12, 13 or 14 all named as reasons to step back, and that line is stated as no longer actively maintained. So a user on a current LTS with a CUDA 12 toolchain is explicitly in the unsupported case, and the documentation is clear about it rather than quietly hoping. The stack-level additions the fork brings are named too: Marlin, FlashInfer and FlashQLA, TurboQuant with INT8 KV cache, MTP and DFlash2, and CUDA Graph support. The hardware target list is wider than the two 2080 Ti cards, covering Tesla T10, T40 and T4, TITAN RTX, and Quadro RTX 6000 and 8000, with validated tensor parallel sizes of 2 and 4.

The launcher owns the machine and the profile owns only the route

The division of responsibility is stated precisely, and it is the thing to understand before using either. Profiles live at a flat path, `profiles/<hardware>/<model>/<weight>/<route>.env`, and a profile selects only route parameters. The launcher owns GPU selection, the port, target and draft model paths, the chat template and the reasoning defaults, so a profile cannot accidentally change which port the service listens on. Four launch modes exist: normal for stable daily deployment, fast for validated routes and the default, aggressive for the highest performance with increased quality risk, and safe as a conservative fallback for troubleshooting and compatibility. The interactive launcher applies a profile, configures the GPU and TP/PP topology, starts the service with health and smoke checks and can stop a running service. Automation goes through environment variables such as MODEL_DIR, PROFILE, MODE, GPU_DEVICES, TP_SIZE and NON_INTERACTIVE, with `--print-config` available to preview a route without launching:

bash
MODEL_DIR=/path/to/checkpoint \
PROFILE=2x2080Ti/qwen27b/w8a16/mtp4-fp8kv-1x262K-text-only.env \
MODE=fast GPU_DEVICES=1,5 TP_SIZE=2 \
NON_INTERACTIVE=1 ./launcher.sh

Note that the GPU list and the tensor parallel size are two independent inputs here, which is how a four-card host can be restricted to a two-card route.

Two weight routes share one profile path, and one model name does not match

The tested checkpoint table has four rows, and two details in it are worth reading closely. The NVFP4 route and the INT4 W4A16 route are different weight formats from different sources, one from unsloth and one from RedHatAI, yet both point at the same profile path, `qwen27b/w4a16`. That is defensible if the profile describes the activation scheme rather than the weight format, but it means the profile name alone does not tell you which checkpoint it expects. The 35B row is named Qwen3.x 35B while linking a card called Qwen3.6-35B-A3B-FP8, so the family label and the specific card disagree in the same row. The hardware story has a similar honesty note: four-GPU Tesla T10 profiles target TP=4, and routes with a noted image-semantic failure are left explicitly marked as candidates rather than presented as validated.

NVLink is recommended and PCIe is only a floor

The hardware question and answer section is where the project is most careful about what it has and has not proven. NVLink is recommended, PCIe peer-to-peer is the baseline requirement, and narrow PCIe links without NVLink are described as not a proven substitute for the validated topology. The instruction is to confirm peer-to-peer and benchmark the actual host topology before treating a configuration as a deployment route, which is the correct framing for a project whose whole value is a list of topologies that were actually measured. The same caution extends to hardware outside the target list: other Turing cards need independent validation across VRAM capacity, PCIe and NVLink topology, model head dimensions, KV cache dtype and CUDA Graph behaviour. Five separate things to re-validate is a fair estimate of how much of a serving stack is specific to the cards it was tuned on.

Updating preserves your local environments and your local profiles

Two shell helpers carry the whole lifecycle, and the update path is the interesting one. A fresh checkout is cloned and built with `./build.sh`. An existing checkout runs `./update.sh`, which is documented as preserving local environments, dependency caches, logs, results and the `profiles/local` directory, then comparing the `VERSION` file with the latest release, downloading the matching source archive and offering to run the build again afterwards. So a user who has local modifications, or who has written their own route under `profiles/local`, keeps them across an upgrade. That is a meaningful design decision for a repository whose artefacts are large compiled builds on a specific CUDA and PyTorch combination, where a re-clone would mean a full rebuild. The trade-off is that the update is archive-based rather than a merge, so local edits outside the preserved paths have to be reapplied by hand.

Editorial conclusion

Judgment: this fork is a serious piece of hardware enablement, and its value is concentrated in things a general vLLM install cannot give you: preserved SM75 source changes, a validated profile per hardware layout, and a launcher that knows which topology was actually benchmarked. Read it that way rather than as a faster vLLM. The two claims to discount are the description's throughput figure, which is less than half what the highlights table reports and therefore describes a different configuration, and the synthetic-input caveat under that table, which says the numbers come from speculative decoding with a high draft acceptance rate on synthetic text. Before cloning, check which branch you actually get, because the repository's default branch is the 0.1.x line that the support section describes as no longer maintained while all the documented baselines live on main. And treat the T10 image-semantic routes as candidates, since the documentation marks them as such rather than validating them.

Frequently asked questions

What is vLLM-2080Ti-Definitive?

A hardware-focused fork of vLLM for dual RTX 2080 Ti and other SM75 GPUs, including Tesla T10, T40, T4, TITAN RTX and Quadro RTX cards. It preserves the SM75 source changes, launcher profiles and validation evidence, and adds Marlin, FlashInfer, TurboQuant with INT8 KV cache, MTP and DFlash2, and CUDA Graph support.

Which branch should I clone for the current vLLM-2080Ti work?

main, where the documented 0.2.x baseline v0.2.2-post3 lives. The repository's default branch is named for the 0.1.x line, which the support section describes as no longer actively maintained and which is the fallback for CUDA 12.8, PyTorch 2.11, older kernels or GCC 12 to 14.

How fast is inference on two RTX 2080 Ti cards?

The highlights table reports 220.84 tok/s for a 4K prompt and 209.35 tok/s for a 32K prompt with Qwen3.8 27B in NVFP4 at 256K context and an FP8 KV cache. Those are single-request tests using DFlash2 with default K=7 on synthetic text with high speculative-hit rates, and real-task throughput depends on draft acceptance.

What launch modes does the vLLM-2080Ti launcher have?

Four: normal for stable daily deployment, fast for validated routes and the default, aggressive for the highest performance with increased quality risk, and safe as a conservative fallback. The profile file selects only route parameters, while the launcher owns GPU selection, port, model paths, chat template and reasoning defaults.

Does vLLM-2080Ti-Definitive work without NVLink?

Peer-to-peer PCIe is the baseline requirement and NVLink is recommended. Narrow PCIe links without NVLink are described as not a proven substitute for the validated topology, and the documentation says to confirm peer-to-peer and benchmark the actual host before treating a setup as a deployment route.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. weicj/vLLM-2080Ti-Definitive on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/weicj-vllm-2080ti-definitive.svg)](https://hysenlabs.com/projects/weicj-vllm-2080ti-definitive)