# club-3090: Docker recipes for serving Qwen3.6 on one or two RTX 3090s

> club-3090 collects validated docker compose configs, patches and helpers for running modern LLMs on RTX 3090s across vLLM, llama.cpp and ik_llama. It is a practical answer to a narrow question: which config actually survives long-context, tool-using workloads on 24 GB cards.

**noonghunna/club-3090** — Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1 and 2 cards.

- Repository: https://github.com/noonghunna/club-3090
- Stars: 2,288 · Forks: 148
- Language: Python
- License: Apache-2.0
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/noonghunna-club-3090

## The problem club-3090 solves for 24 GB Ampere owners

Running a 27B model on a single RTX 3090 is not a question of whether it loads. It is a question of what breaks first: the KV cache at long context, the prefill pass on a large prompt, or the tool-calling loop when a 25K-token tool return comes back. club-3090 is a repository of compose files and helper scripts that encode answers to those questions for specific model and card combinations, rather than a framework you build on.

The intended audience is narrow and stated plainly in the README: people with one or two RTX 3090s who want to run modern LLMs at home, in a homelab, or as a dev backend. The repo is model-agnostic in structure but ships curated configs today for Qwen3.6-27B and related models at 1 and 2 card counts, with vLLM, llama.cpp and ik_llama as the engines. If you do not own this hardware class, most of the value is in the documentation rather than the compose files.

## Two routes, one API: how the vLLM and llama.cpp composes differ

The central design decision is that club-3090 does not pick a winner between engines. It ships two complementary routes and tells you to choose by what your workload breaks on. The vLLM dual route targets maximum throughput, with the README citing up to 127 TPS on code workloads and 4 concurrent streams at 262K context, plus vision, tools, MTP and streaming. The llama.cpp single route targets robustness: 200K context on one 3090, described as max-safe with margin, and stress-tested against prefill cliffs and 25K-token tool returns.

Both routes expose a drop-in OpenAI-compatible API on localhost:8020. That is the architectural contract: whatever engine sits underneath, your client points at the same port and speaks the same protocol. The composes are the unit of delivery, and scripts/switch.sh handles taking one down before bringing another up.

The trade-off is explicit rather than hidden. The llama.cpp single route is slower, roughly 51 to 60 TPS with Q4_K_M and MTP according to the README, but it does not crash on real-world tool-using agents. The vLLM dual route is faster and feature-complete but needs two cards to escape the long-context cliff described below. There is no config in the repository that gives you both on one card.

## Installing club-3090 and serving a first model

The README assumes Linux or macOS. On Windows, the documented path is WSL2 first, via docs/WSL_SETUP.md, because native Windows runs only the upstream llama.cpp binary and none of this repository's tooling.

Start by cloning and letting the interactive setup script pick and verify a model. The script asks which model you want and where weights should live, and it SHA-verifies the download. Profile compatibility tooling needs PyYAML, which Ubuntu LTS usually provides as python3-yaml.

```bash
git clone https://github.com/noonghunna/club-3090.git
cd club-3090
bash scripts/setup.sh
```

To skip the prompts, export MODEL_DIR and pass the model name. Note the variable is singular; the .env.example warns that MODELS_DIR is silently ignored.

```bash
MODEL_DIR=/scratch/models bash scripts/setup.sh qwen3.6-27b
```

Launching is a single script. According to the README, launch.sh calls switch.sh to bring the old variant down and the new one up, then runs verify-full.sh so you know the endpoint is serving cleanly before a client connects.

```bash
bash launch.sh
```

If you prefer a terminal UI, the serve cockpit installs from the checkout. With uv it is one command; with plain pip you install the core package first.

```bash
uv pip install -e tools/serve-cockpit
c3
```

On first run, press S to set the model directory and HuggingFace token, Ctrl+S to save, then r to browse the catalog and serve a variant. After a git pull, the install must be re-run to pick up new dependencies and UI changes.

## The single-card long-context cliff and the retired escape hatch

The most important limitation is documented in the README rather than buried: on 24 GB single-card vLLM, there is an open issue the project calls Cliff 2, a GDN prefill OOM at single prompts above roughly 50K tokens. The stated workaround is the dual-card route with TP=2, which escapes it.

The former single-card escape, the llamacpp/default variant, was retired on 2026-08-12 and is now available only with --force. The README is direct about the consequence: on one card there is no longer a cliff-immune Qwen path. That is a real regression for anyone who adopted the repo for exactly that configuration, and it is the kind of detail that a summary page would omit.

If your workload involves long single prompts on one 3090 and you cannot add a second card, this repository currently points you at a route it has withdrawn. The diagnosis lives in docs/CLIFFS.md, and that file is the right place to start before assuming a config will hold.

## Beyond 3090s: 4090, 5090 and multi-GPU configs

The tooling is calibrated for 3090s, but the README states the configs are class-aware, and contributors have benchmarked 4090 and 5090 rigs. Per-class caveats are documented: the 4090 has tighter idle VRAM, and the 5090 has a 32 GB envelope. Cross-rig benchmark rows live in the FAQ rather than in the main tables.

For three or more GPUs of any class, docs/MULTI_CARD.md covers tensor-parallel scaling math, derivation from dual.yml, and which TP values are valid. That page is the honest boundary of the project's support: the composes are written for one and two cards, and beyond that you are deriving your own configuration from the dual-card baseline.

SGLang is listed as evaluated but currently blocked on Ampere, with details in docs/engines/SGLANG.md. That is a useful signal for anyone who assumed all major serving engines were interchangeable on this hardware.

## Universal pull versus hand-written composes

Since v0.8.0, with extensions in v0.8.2, the repository ships a universal pull flow. It evaluates any safetensors Hugging Face repository against the KV-cache math, gives a one-line fit verdict with --recommend, and when a pull hard-blocks, lets you send a redacted diagnostic back in one consented step with --submit-last. Architecture coverage widens each release.

This is a different approach from tools that simply download weights and let you discover the fit problem at load time. The verdict is described as honest about confidence rather than binary, which matters because the answer for a given model and card count is often 'it fits, but not at the context length you want.'

The limitation is that the pull flow does not replace a curated compose. It tells you whether a model is plausible on your hardware; it does not give you the tuned flags, the MTP setting, or the per-workload pitfalls that the model-specific pages under models/<name>/ carry. For a model outside the supported list, you are evaluating first and configuring second.

## Alternatives, licensing and what maintenance actually looks like

The obvious alternative is running the engines directly. vLLM and llama.cpp both ship their own documentation and their own docker images, and you can write a compose file for either in an afternoon. The difference in approach is that club-3090 does not abstract the engines away; it records which flags, quantizations and card counts produced a working result for a specific model. That record is the product. If you enjoy deriving KV-cache budgets and testing prefill behavior yourself, the upstream projects give you the same engines with fewer opinions layered on top.

A second alternative is a hosted API. The repository includes docs/COMPARISONS.md covering cost crossover and when each option wins, so it addresses that comparison itself rather than pretending local always wins.

On maintenance: the repository is not archived, and the last push was on 2026-07-13, with v0.10.2 released the same day. That is roughly two months before the current date, so the project is not dormant, but the release cadence visible in the changelog is the thing to check rather than any general claim about activity. The licence is Apache-2.0, which permits commercial use and modification with the usual notice and patent-grant conditions; that is a description of the licence text, not legal advice, and you should read LICENSE before shipping anything derived from it. The practical upgrade cost is low if you use the cockpit, since the README notes that a git pull requires re-running the install to pick up new dependencies and UI changes.

## Conclusion

Adopt club-3090 if you already own one or two RTX 3090s and want a working OpenAI-compatible endpoint on localhost:8020 without deriving KV-cache math yourself. Skip it if you run a different accelerator class or want a single engine abstraction; the recipes are calibrated for 3090s and the composes are class-aware rather than generic. Before committing, read docs/CLIFFS.md for the single-card prefill cliff, check that the retired llamacpp/default variant is not in your path, and confirm your MODEL_DIR name is singular, since MODELS_DIR is silently ignored.

## FAQ

### Is an RTX 3090 still good in 2026?

club-3090 is built entirely around the premise that it is, with production-ready configs for Qwen3.6-27B on one and two 3090s. The README does not make a general claim about the card's standing; it documents what runs on it today.

### Does the RTX 3090 support FP16?

The repository material does not state an FP16 capability answer for the 3090. It discusses quantization choices such as Q4_K_M, IQ4_KS and AWQ in docs/QUANTIZATION.md, which is the closest documented ground.

### How many CUDA cores does an RTX 3090 have?

The repository material does not give a CUDA core count. Hardware questions are routed to docs/HARDWARE.md, which covers 4090, NVLink and power caps rather than core specifications.

## Sources

- [Official README](https://github.com/noonghunna/club-3090#readme)
- [Project repository](https://github.com/noonghunna/club-3090)
- [Release notes](https://github.com/noonghunna/club-3090/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/noonghunna-club-3090
