# Qwen3.8-27B on a consumer card: what the launcher picks, and what happens when VRAM runs out

> A Windows and Linux serving kit that picks an EXL3 quant to fit the card it finds, installs its own Python environment, and serves an OpenAI-compatible endpoint with a chat UI on port 3080. The failure mode it documents best is not an error at all: too little free VRAM makes the driver page the model into system RAM and run it many times slower.

**MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install** — Qwen3.8-27B on 16-32 GB Nvidia GPUs one-click install for Windows / Linux

- Repository: https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install
- Website: https://x.com/MiaAI_lab
- Stars: 544 · Forks: 56
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/miaai-lab-qwen3-8-27b-16gb-nvidia-gpus-one-click-install

## Too little free VRAM pages the model into system RAM instead of failing

The most useful thing in the kit's documentation is a failure mode that produces no error. `windows\start.bat` checks free VRAM immediately before loading a model, and it wants the `GPU_MEM_GB` budget from `.env` plus a little margin. If that much is not free, it lists the programs holding VRAM, browsers, games, Discord and other AI tools by name, and waits: Enter re-checks, `c` continues anyway, `q` quits, and after 120 seconds it continues on its own. Continuing anyway is the trap. The README says to take the check seriously on Windows, because with too little free VRAM the driver pages the model into system RAM instead of failing, and it then runs many times slower. So a working chat window is not evidence that the model fits. The same free VRAM reading is the first item in the troubleshooting table for a model that loads but crawls, alongside lowering `CONTEXT_SIZE` or `GPU_MEM_GB`.

## PyPI has no exllamav3 1.4.4, the GitHub release does

Normally nothing needs compiling. The engine's own GitHub release publishes a wheel per combination of CUDA line, torch version and Python, and the launcher installs the one that matches the virtual environment it just created, roughly a 100 MB download with no compiler involved. The version gap is stated plainly in the configuration file: PyPI is what has no 1.4.4, because it jumps from 1.4.2 to 1.4.5, and the release has it. A source build is the fallback for the cases no wheel covers, which the kit names as an unbuilt CUDA line, aarch64 on GB10 or Spark, or a torch newer than the release. Setting `EXL3_REPO` forces the source path from a git URL or a local path on the machine, and the example given is `EXL3_REPO=git+https://github.com/turboderp-org/exllamav3.git@v1.4.4`, with a warning not to copy a Linux or Spark path onto Windows. Inside the engine repository itself, the install builds from the local checkout instead.

## cu128 is the default because the engine release builds cu128 and cu132 only

`TORCH_INDEX_URL` overrides the PyTorch wheel index, and the default is cu128, which the configuration file calls deliberate. The reason is narrow and mechanical: the engine release builds for cu128 and cu132 only, so a torch from any other CUDA line means there is no prebuilt engine and the kit has to compile. cu128 is also what covers Blackwell, and it needs driver 570 or newer. GB10 and Spark are the exception, using cu130 and compiling either way. The torch version matters as much as the index, since the newest torch is not always one the engine publishes a wheel for, so the launcher picks the version too. `tools/wheels.py` reads that version out of the same table it uses to choose the wheel, which is the mechanism that stops the two choices from drifting apart. You can run that table yourself with `.venv\Scripts\python.exe tools\wheels.py`. The driver and CUDA requirements in the prerequisites table follow from the same choice rather than being independent.

## Node is only needed for the chat UI, and Python is the only manual install

The prerequisites are narrower than the model name suggests. Python 3.11 or newer, 64-bit, is the single thing you install by hand. Node 22.19 or newer is needed for the chat UI only, since that is the version the current `dsh` wants; an older LTS such as 20 raises `EBADENGINE` and the chat UI may fail, while without Node the `/v1` endpoint still serves and the launcher states what is missing. CUDA Toolkit, Visual Studio Build Tools and Git are all listed as not needed, because the engine arrives as a prebuilt wheel and compiling is the fallback. Everything the kit installs stays inside its own folder, `.venv/`, `models/`, `logs/`, `apps/` and `.dsh/`, so nothing lands in the system Python and no administrator rights are required. Disk is the number to plan around: 9.7 to 22.9 GB per quant, plus several GB for the Python environment and PyTorch. The card itself needs 12 GB of VRAM or more and compute capability 7.5 or newer.

## The 2.0 bpw quant is the only weight set not from turboderp

The kit began as a 16 GB recipe and the 2.0 bpw quant is still that floor. It is also the one artifact here that is not turboderp's own upload: it comes from `Mia-AiLab/Qwen3.8-27B-EXL3-2.0bpw` under the filename `SC_2.00bpw_H3_V3`. Everything from 2.5 bpw upward is pulled from `turboderp/Qwen3.8-27B-exl3` by revision, so the provenance of the weights is split between two uploaders depending on which end of the size range you land on. Which quant you actually get is decided by your VRAM at setup, and you can change it at any time. The disk footprint follows the same ladder, from 9.7 GB at the small end to 22.9 GB at the large one. The README also states plainly that 16 GB is the card size the kit was built around, while the requirement line admits 12 GB, which leaves the smallest supported card running the floor quant by definition.

## START-HERE.bat installs and starts, --no-start installs only

On Windows there are two entry points answering two different questions, and the split is deliberate:

```
windows\START-HERE.bat              install, then start what was installed
windows\START-HERE.bat --no-start   install only, for fetching a second size
```

The one comment holding an em dash in the original has been reduced to a comma here; the two commands are unchanged. `START-HERE.bat` opens a page in the browser and does the whole install there, showing what it found on the card, offering the model sizes that fit it, then installing and downloading with a progress bar and a live log while nothing is asked in the console. When the download finishes it loads that model and hands the page over to the chat. The download is resumable, so closing the window or rebooting costs nothing, and a half-downloaded model is labelled as such in the menus. Nothing offers to start a model until every weight file is on disk. Setting `SETUP=console` in `.env` brings back the older console questions. The daily script is the other half:

```
windows\start.bat          start a model that is already here
windows\start.bat setup    go to setup instead (same as START-HERE.bat)
```

It never downloads anything.

## The console window is the server, and the chat listens on 127.0.0.1:3080

`windows\start.bat` picks a model when more than one size is on disk, Enter takes the one that ran last, and it starts on its own after 45 seconds so an unattended machine still comes up. It then opens the chat UI at `http://127.0.0.1:3080/` over an OpenAI-compatible endpoint. The console window it opened is the server, so closing that window stops the model, which is worth knowing before you tidy up the taskbar. A tray component called Simplex adds an icon with Open Simplex, Restart the model, Show the Simplex folder, View the log and Quit Simplex, switchable off with `TRAY=no`. Every launch writes a full transcript into `logs\`, so a crash that scrolls past is still readable, and `windows\simplex.bat logs` prints the tail. The first successful launch also adds Start-menu and desktop shortcuts, which `SHORTCUTS=no` suppresses. Stopping has three shapes:

```
windows\stop.bat                 stop both the model and the chat UI
windows\stop.bat --harness-only  leave the model loaded, close the UI
windows\stop.bat --server-only   leave the UI running, unload the model
```

## Images: off means the vision tower lost the argument with CONTEXT_SIZE

The troubleshooting table is short and each row names a key. A Ready box reading Images: off means the vision tower did not fit alongside the context, and the fix is to lower `CONTEXT_SIZE` and restart, or pick a smaller quant, which puts the same trade-off in two places: what fits in VRAM is shared between the weights and the context window. A model that loads but crawls points back at free VRAM or at `CONTEXT_SIZE` and `GPU_MEM_GB`. A window that closed before you read the error leaves the answer in `logs\`, newest file first. Simplex failing to find Python means installing 64-bit Python 3.11 or newer from python.org and ticking Add python.exe to PATH, then running the file again. Starting over is `reset_new_user.bat` in the kit root, which deletes the weights, the venv, `.env`, the logs and the shortcuts while keeping every tracked file, and asks you to type `RESET` before it does anything. The Linux half of the kit lives in the `linux/` folder, and its instructions say to run the scripts from the kit root because they find their own way, where that section stops.

## Conclusion

Adopt this kit if you have a 12 GB or larger NVIDIA card with compute capability 7.5 or newer and want Qwen3.8-27B served locally without assembling an environment yourself, because the quant choice, the engine wheel and the torch version are picked for you from tables in the repository. Do not adopt it on a card you share with a browser or a game, since the kit's own warning is that the driver will page the model into system RAM rather than fail. Before running `windows\START-HERE.bat`, read the free VRAM step in `windows\start.bat`, decide whether cu128 is right for your card, and note that PyPI has no exllamav3 1.4.4, so a `pip install` path would land on 1.4.2 or 1.4.5 instead.

## FAQ

### What does this Qwen3.8-27B kit need before it will run?

An NVIDIA card with 12 GB of VRAM or more and compute capability 7.5 or newer, driver 570 or newer, and 64-bit Python 3.11 or newer, which is the only thing installed by hand. Node 22.19 or newer is needed for the chat UI, and without it the `/v1` endpoint still serves. CUDA Toolkit, Visual Studio Build Tools and Git are not required.

### Which quant of Qwen3.8-27B will the kit download for my card?

The choice is made from your VRAM at setup and can be changed at any time. The 2.0 bpw quant is the 16 GB floor and comes from `Mia-AiLab/Qwen3.8-27B-EXL3-2.0bpw`, while everything from 2.5 bpw up is pulled from `turboderp/Qwen3.8-27B-exl3` by revision. Disk use runs from 9.7 GB to 22.9 GB per quant.

### Does the kit need a compiler or the CUDA Toolkit?

No. The engine's GitHub release publishes a wheel per CUDA line, torch version and Python combination, and the launcher installs the matching one, about a 100 MB download with no compiler. Compiling is the fallback for an unbuilt CUDA line, aarch64 on GB10 or Spark, or a torch newer than the release, and `EXL3_REPO` forces that path.

### Why does my Qwen3.8-27B model load but run slowly on Windows?

The README points at free VRAM. `windows\start.bat` checks it against the `GPU_MEM_GB` budget from `.env` before loading and lists what is holding it, and with too little free VRAM the driver pages the model into system RAM instead of failing, which makes it run many times slower. Lowering `CONTEXT_SIZE` or `GPU_MEM_GB` is the other lever.

### Where does this kit put the model server and how do I stop it?

It serves an OpenAI-compatible endpoint with the chat UI at `http://127.0.0.1:3080/`, and the console window it opened is the server, so closing that window stops the model. `windows\stop.bat` stops both the model and the UI, `--harness-only` leaves the model loaded, and `--server-only` unloads the model while the UI keeps running.

## Sources

- [Issues](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install/issues)
- [License: MIT](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install/blob/main/LICENSE)
- [MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install on GitHub](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install)
- [Project website](https://x.com/MiaAI_lab)
- [README](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/miaai-lab-qwen3-8-27b-16gb-nvidia-gpus-one-click-install
