Model or dataset
MakazhanAlpamys/Soup avatar
MakazhanAlpamys/Soup

Soup: fine-tune an 8B LLM on a 4 GB laptop GPU from one YAML file

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

7,189 stars1,142 forksPythonApache-2.0

At a glance

What is it?
Soup is a Python CLI that wraps QLoRA fine-tuning and DPO behind a single YAML config, with an opt-in layer-streaming mode that keeps the frozen base out of VRAM. The promise is real but narrow, and the docs themselves flag the parts still in beta.
Who is it for?
Soup fits engineers with one consumer GPU who want a LoRA or QLoRA run configured in YAML and started with a single command, and who are willing to read the changelog before trusting a number. It does not fit teams that need a stable API surface: pyproject.toml declares Development Status 3 - Alpha, and the v0.74.0 notes list a breaking change to soup serve.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Soup targets: QLoRA setup work, not model quality

Most of the work in a small fine-tuning run is not the training loop. It is deciding a batch size that fits, picking a quantization scheme, wiring PEFT adapters onto the right modules, and keeping the frozen base in a dtype that does not blow up memory. Soup's pitch is that this layer is configuration, not code. You write a YAML file, run one command, and the CLI handles GPU detection, batch sizing and quantization.

The audience is specific. The README's headline claim is an 8B model on a 4 GB laptop GPU, and the repository topics include consumer-gpu and low-vram. That is a person with one gaming laptop or a single rented card, not a cluster. The README also states that even experienced teams spend 30-50% of their time fighting infrastructure, which is the argument for the abstraction, though that figure is the project's own framing and not something a reader can check from the repository.

The scope is broader than SFT. The topics list dpo and qlora alongside sft, and the release notes for v0.74.0 mention preference losses arriving in v0.72.3. So the intended workflow covers both supervised fine-tuning and preference post-training behind the same config format.

How layer streaming keeps the frozen base out of VRAM

The mechanism is stated plainly in the README: layer streaming keeps the frozen base out of VRAM and feeds it to the GPU one decoder layer at a time. The video caption in the README describes the concrete layout for Llama-3.1-8B-Instruct with NF4, LoRA, batch 1 and sequence length 512: a 3.60 GB base store pinned in RAM across 32 layers, plus two 113 MB VRAM buffers, reaching a measured peak of 3.32 GB.

That is the whole idea. The weights live in host memory, and only the layer currently being computed occupies device memory. It is a memory-for-bandwidth trade, and the README is explicit that the trade has a cost: the tok/s figure of 119.6 was measured on v0.72.2, before the v0.73.0 correctness repair that cost 4.8% at 32B, and it has not been re-run on a 4 GB card since. Anyone quoting 119.6 tok/s today is quoting a pre-repair number.

The correctness argument is the interesting part. The README says the streamed run is bit-exact against a normal resident run, and that this was reproduced independently on an H100 at 113.00 tok/s in the same 3.32 GB. There is a paper under a Zenodo DOI, and a proof notebook that caps the process to 4 GB and then asserts bit-identity. Streaming is opt-in via stream_layers: true and the README labels it BETA.

Installing soup-cli and running a first chat-template fine-tune

The core package is deliberately light. According to pyproject.toml, a bare pip install soup-cli pulls the CLI, config system and data tools, but no PyTorch; the training stack (torch, transformers, peft, trl, datasets, bitsandbytes, accelerate) lives in the [train] extra. The README gives the install line with that extra.

bash
pip install "soup-cli[train]"
soup init --template chat
soup train

The first command installs the training dependencies. The second writes a starter config from the chat template, and the third starts the run. The README does not print the generated YAML, so what you get from soup init --template chat has to be read from the file it creates.

Python version is bounded on purpose. pyproject.toml sets requires-python to >=3.10,<3.13, and the comment explains why: without a ceiling, pip on 3.13+ resolves torch wheels the project has not run, and the failure surfaces as a loader crash inside c10.dll or libc10.so before any Soup code executes. There is a test, tests/test_requires_python_bound.py, that derives the bound from the CI matrix.

For a container instead of a local install, the repository ships a Dockerfile and a docker-compose.yml. The image installs from PyPI rather than local source, and the compose file reserves all NVIDIA GPUs with ipc: host.

yaml
services:
  soup:
    image: ghcr.io/makazhanalpamys/soup:latest
    volumes:
      - .:/workspace
    ipc: host

Layer streaming is not on by default. To use it you set stream_layers: true in the config, and the README marks the feature beta.

The fp32 base bug and what v0.74.0 actually changed

The v0.74.0 release title is blunt: the frozen base was loaded in fp32 the whole time. According to the release notes, every SFT load silently upcast the frozen base to fp32 on all three load paths, meaning a base that never receives an optimizer step was materialised at twice its checkpoint precision. Measured on an H100 with Llama-3.1-8B plus LoRA, peak went from 48,241 MiB to 18,658 MiB, a 2.59x reduction, byte-identical across three repeats.

That is worth pausing on. A memory bug of this size sat in the default path, and the fix changes peak VRAM by more than a factor of two on an unchanged config. A trainable base still loads fp32 deliberately, so the fix is scoped to the frozen case.

The same release fixes a second memory-adjacent failure: the free Colab and Kaggle tiers could not stream at all. T4, P100, V100 and GTX 16xx crashed layer streaming because peft creates LoRA adapters in the checkpoint dtype while the fp16 GradScaler needs fp32 gradients. That is the exact hardware the headline claim is aimed at, which suggests the 4 GB story was validated on the RTX 3050 path before it held on the free tiers.

There is also a security note: four SSRF bypasses of the same shape (abbreviated, decimal, hex and octal IPv4 spellings such as 127.1, 2130706433, 0x7f000001 and 0177.0.0.1) reached the telemetry and webhook guard and, through a path the first fix never touched, the OTLP tracing validator. And soup serve now exits 2 when bound to a non-loopback host without --tool-auth-token, instead of printing a warning. That is a breaking change for anyone who was running the server on a public interface.

Where Soup is the wrong tool

The declared torch floor is broken. The README's known-limitation note says the declared torch>=2.5.0 floor does not work with trl>=0.29: at torch 2.5.1, trl cannot import. A fresh install resolves a newer torch, so the pinned floor is a trap for anyone building a reproducible environment from the declared constraints rather than from a lockfile.

Layer streaming is beta and the headline benchmark is stale. The 119.6 tok/s figure predates the v0.73.0 correctness repair and has not been re-run on a 4 GB card. If your decision depends on a throughput number at 4 GB, the repository does not currently give you a current one.

The project also labels itself alpha. pyproject.toml carries Development Status 3 - Alpha, and v0.74.0 shipped a breaking change to soup serve alongside the memory fix. If you need a training stack whose CLI contract is frozen, this is not it yet.

Finally, layer streaming is a memory technique, not a speed technique. It trades host memory and data movement for VRAM, and the release notes record a 4.8% throughput cost at 32B from a correctness repair. If you have enough VRAM to hold the frozen base, the resident path is the simpler one, and the README presents streaming as opt-in for that reason.

How Soup differs from a hand-written PEFT script

The real alternative is not another CLI. It is a training script you write yourself with transformers, peft and trl, which is exactly the stack Soup wraps. The difference is where the decisions live. In a hand-written script, batch size, quantization config, target modules and dtype handling are Python you control and can debug line by line. In Soup, they are defaults the CLI picks, and the v0.74.0 fp32 bug is the cost of that arrangement: a default you did not write silently doubled the memory of the frozen base on every SFT load path.

The counter-argument is also visible in the same release. Layer streaming is not something a typical hand-written PEFT script does. Feeding a frozen base to the GPU one decoder layer at a time, with a bit-exactness assertion and a proof notebook that caps the process at 4 GB, is real engineering that a solo script author would have to build from scratch. If your constraint is a 4 GB card, the abstraction is buying you something you would not otherwise have.

So the split is by constraint. If you have VRAM and want full control of the training loop, a script is the lower-risk path. If you do not have VRAM, Soup is offering a mechanism, not just a wrapper.

Licence, maintenance and the cost of upgrading

Soup is Apache-2.0, declared in pyproject.toml and present as a LICENSE file at the repository root, with a NOTICE file alongside it. The Dockerfile installs from PyPI rather than local source, so a container built from this repository tracks the latest published release. For a team that needs to pin a version, that is a detail to check before building the image.

On maintenance, the last push to the repository was on 2026-09-09, and the repository is not archived. Releases have been frequent: v0.73.2 on 2026-08-15, v0.73.3 on 2026-08-18, v0.74.0 on 2026-09-04, with pyproject.toml already at version 0.75.0. The v0.74.0 notes state that 116 of the 120 merged pull requests in that release came from outside the maintainer, by 25 people, which is a real signal about where the work is coming from.

Upgrade cost is the part to plan for. The project is alpha and the release notes contain a breaking change (soup serve exiting 2 on a non-loopback bind without --tool-auth-token), a dependency-surface change (Transformers 5.x, TRL 0.29, PEFT 0.20), and a memory fix that changes peak VRAM by 2.59x on unchanged configs. A version bump can move memory, dependencies and CLI exit codes at once. Two extras, [train] and [mlx], previously declared ranges that could not be satisfied together; v0.74.0 states that pip install "soup-cli[train,mlx]" resolves again.

Editorial conclusion

Soup fits engineers with one consumer GPU who want a LoRA or QLoRA run configured in YAML and started with a single command, and who are willing to read the changelog before trusting a number. It does not fit teams that need a stable API surface: pyproject.toml declares Development Status 3 - Alpha, and the v0.74.0 notes list a breaking change to soup serve. Before adopting, verify two things on your own hardware. First, run the Colab notebook at notebooks/proof-4gb.ipynb, which caps the process at 4 GB and asserts a streamed model is bit-identical to a normal one. Second, check your torch version against the known limitation in the README: the declared torch>=2.5.0 floor does not work with trl>=0.29.

Frequently asked questions

What is Soup and what does it do?

Soup is a Python CLI, published as soup-cli, that configures and runs LLM fine-tuning and post-training from a single YAML file. It wraps the torch, transformers, peft and trl stack behind commands like soup init and soup train.

How do I install Soup?

The README gives pip install "soup-cli[train]" to include the training stack; a bare soup-cli install is the light CLI without PyTorch. The package requires Python >=3.10,<3.13.

Can Soup really fine-tune an 8B model on a 4 GB GPU?

The README claims an 8B model on a 4 GB laptop GPU using layer streaming, measured at 3.32 GB peak on an RTX 3050 Laptop 4 GB. That throughput figure was measured on v0.72.2, has not been re-run on a 4 GB card since, and layer streaming is opt-in and labelled beta.

Official sources

  1. License: Apache-2.0
  2. MakazhanAlpamys/Soup on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/makazhanalpamys-soup.svg)](https://hysenlabs.com/projects/makazhanalpamys-soup)