# AirLLM's layer streaming: 4GB of VRAM, 360GB of disk, and a pinned transformers

> AirLLM runs 70B and larger models by swapping layers in and out of a small GPU, which moves the bottleneck from VRAM to host disk and to dependency versions. Here is what each supported model actually costs in disk, host RAM, and an exact transformers release.

**lyogavin/airllm** — AirLLM 70B inference with single 4GB GPU

- Repository: https://github.com/lyogavin/airllm
- Stars: 35,279 · Forks: 3,728
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/lyogavin-airllm

## First inference writes the whole model out, layer by layer

The quickstart is two steps. Install the package:

```bash
pip install airllm
```

Then hand a Hugging Face repo id to `AutoModel.from_pretrained` and the model is used like a regular transformer model:

```python
from airllm import AutoModel

MAX_LENGTH = 128
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")
```

The same one line handles larger models, and the commented variants in that example climb to `deepseek-ai/DeepSeek-V3` at roughly 12GB. What you cannot skip is the setup cost hiding under the line. The note beside the example says that during inference the original model is first decomposed and saved layer-wise, and that you need enough room in the Hugging Face cache directory. The init call also accepts `layer_shards_saving_path` if you want to choose where those layers land. For Qwen3.8-Flash-Next the disk figure is about 360GB of checkpoint, with `delete_original=True` available to reclaim the originals after the split. So the first prompt is not really a prompt. It is a long write to disk.

## The VRAM figure is the one published, and the host figure is the one you budget

Every headline number here is a VRAM number. 70B on a single 4GB card without quantization, distillation, or pruning. Kimi K3 at 2.8T parameters under 4GB. Qwen3.8-Flash-Next at 125B in 5.95GB, measured end to end on one RTX 4090. Qwen3.8-27B in 3.33GB on one RTX 3090. Kimi K3 in 3.72GB on one RTX 6000 Ada. DeepSeek-V3 at 671B in about 12GB. None of those figures include the host, and the host is where the weight lives. For Qwen3.8-Flash-Next, whose n-gram embedding table runs to about 51B entries, that table is file-mapped on the host and a 64GB machine is stated to be enough, while the decoder layers stream. The consequence is that AirLLM does not erase the memory requirement, it relocates it onto disk bandwidth and host RAM. A machine with 4GB of VRAM and a slow or small SSD is still the wrong machine, and nothing in the project predicts which one that is.

## The supported model list cannot be installed into one environment

This is the sharpest conflict in the repository. Kimi K3 support needs `transformers` 4.56.x, because the model's remote code does not load on 5.x. Qwen3.8-27B support needs `transformers` 5.8+. Qwen3.8-Flash-Next needs a build with `qwen4_exp` in tree, which the project installs straight from the GitHub main branch:

```bash
pip install git+https://github.com/huggingface/transformers.git
```

One virtual environment therefore cannot hold both a Kimi K3 and a Qwen3.8-27B at the versions they require. The pinned dependency file does not resolve this, because it points at the transformers repository with no version constraint, and holds `peft` at v0.3.0 and `accelerate` at v0.20.3. What you get is a package that installs without complaint and then fails at the first forward pass for whichever model does not match your build. Pick the model first, then build the environment around that model's line.

## The 3x compression speedup is opt-in and drags in a second dependency

Block-wise compression is a separate switch. Enabling it takes three steps: put `bitsandbytes` in place with `pip install -U bitsandbytes`, keep airllm past 2.0.0 with `pip install -U airllm`, and pass the argument at init:

```python
model = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct",
                     compression='4bit' # specify '8bit' for 8-bit block-wise quantization
                    )
```

The project puts the ceiling at up to 3x inference speed with what it calls almost ignorable accuracy loss, and it links a paper for why block-wise rather than plain quantization: ordinary quantization has to compress weights and activations both, which makes accuracy harder to hold and leaves it exposed to outliers. Watch the pinning, because the file and the instructions disagree. `requirements.txt` holds `bitsandbytes==0.39.0` while the setup steps tell you to upgrade it, and compression is the only consumer of that dependency here. The consequence is that the fastest configuration is the one your pinned file is least likely to match.

## Usage is a Python import, and the repository is a notebook first

There is no command line entry point in the quickstart. Everything is a Python call, and the repository shows that: its primary language is a Jupyter Notebook, and the one example file it carries is `examples/inferrence.ipynb`, misspelled in the tree itself. The package code sits in `air_llm/`, and around it are `training/`, `rlhf/`, `eval/`, `scripts/`, `data/`, and `anima_100k/`, which is more surface than a weekend team will touch. The init call also accepts a local path instead of a repo id, and the prose still describes the older class name `AirLLMLlama2` even though `AutoModel` is what the example imports. What the project does not hand you is a served endpoint, so putting AirLLM behind an application means writing that glue yourself, and the glue inherits the layer-by-layer load on every request.

## Training keeps the adapters resident and streams everything else

The September update carries the same idea past inference: stream the frozen weights one layer at a time and keep the adapters on the GPU. Two figures arrive with it. Qwen3.8-Flash-Next at 125B trains under 6GB on an RTX 3060 Ti, and Qwen3.8-27B trains in about 2GB at sequence length 512. The v4.0.0 tag is named for the first of those. The supporting packages are already in the pinned file, with `peft` for the adapters and `wandb` for tracking runs, joined by `evaluate`, `scikit-learn`, and `sentencepiece`. The training note does not give a figure for any context longer than 512 tokens, and it does not say how many optimizer states stay resident. Sequence length is the number to ask about before you book GPU time, because it is the axis the one published figure is tied to.

## Version numbers track model support, so a pin is a model whitelist

The release list reads like a model list. v3.2.0 in August 2026 added Qwen3.8-27B on 3.33GB, v3.3.0 ten days later added Qwen3.8-Flash-Next on 5.95GB, and v4.0.0 in early September is named for training Qwen3.8-Flash-Next under 6GB. Before that, v3.0 in June 2026 brought FP8 support and DeepSeek-V3 at about 12GB, and the 2024 line moved from CPU inference and non-sharded models in v2.10.1 to Qwen2.5 in v2.11.0, with MacOS support arriving earlier as v2.8.2. The repository is not archived and the last push landed on 2026-09-29. Three tags in three weeks is a quick cadence for a package where each release also implies a different set of supported models and dependency lines. The practical consequence is that pinning the version pins your model list, and upgrading to reach a new model means rechecking the CUDA and transformers requirements from the section above.

## Conclusion

Take it when a model that will not fit is the only option and you have the disk and a Linux box with the right CUDA build. Skip it when latency matters, when your environment already pins transformers for something else, or when a 360GB checkpoint does not fit on the host. Before you start, read the model-specific dependency lines, because the newest supported models ask for transformers versions that contradict each other.

## FAQ

### What is AirLLM?

A pip package that reduces inference memory usage so a 70B model runs on a single 4GB GPU card, with no quantization, distillation, or pruning involved. It is licensed Apache-2.0 and its code lives in the air_llm directory.

### Does AirLLM really work?

The project reports end to end measurements rather than estimates: 5.95GB for Qwen3.8-Flash-Next on one RTX 4090, 3.33GB for Qwen3.8-27B on one RTX 3090, and 3.72GB for Kimi K3 on one RTX 6000 Ada. On the training side it reports Qwen3.8-Flash-Next under 6GB on an RTX 3060 Ti.

### How to setup AirLLM?

Run `pip install airllm`, then pass a Hugging Face repo id or a local path to `AutoModel.from_pretrained`. During inference the model is first decomposed and saved layer-wise, so the Hugging Face cache directory needs room before anything is generated.

### How much RAM does it take to run a 70B model with AirLLM?

The host figure given is for Qwen3.8-Flash-Next, whose roughly 51B n-gram embedding table is file-mapped on the host, where a 64GB machine is stated to be enough while decoder layers stream. That same model needs about 360GB of checkpoint disk, with delete_original=True reclaiming the originals after the split.

### Is AirLLM slow?

The project publishes two speed claims and no throughput number: a 10% improvement from prefetching added in v2.5, and up to 3x from block-wise compression at 4bit or 8bit. Since layers stream from disk, disk speed and the pinned dependency versions both land in your latency budget.

## Sources

- [Official README](https://github.com/lyogavin/airllm#readme)
- [Project repository](https://github.com/lyogavin/airllm)
- [Release notes](https://github.com/lyogavin/airllm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lyogavin-airllm
