# GuppyLM: a 9M-parameter fish that teaches you how an LLM is built

> GuppyLM is a from-scratch transformer trained on 60K synthetic fish conversations. It is an educational artifact, not a chatbot, and the README is upfront about that.

**arman-bd/guppylm** — A ~9M parameter LLM that talks like a small fish.

- Repository: https://github.com/arman-bd/guppylm
- Stars: 3,840 · Forks: 358
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/arman-bd-guppylm

## What GuppyLM actually is, and who it is for

GuppyLM is a language model that plays a fish. The README describes it as a tiny model that speaks in short, lowercase sentences about water, food, light, and tank life, and states plainly that it does not understand human abstractions like money, phones, or politics. That is the whole product surface. There is no tool calling, no retrieval, no system prompt negotiation.

The intended audience is people who have used an LLM API and want to see the machinery underneath. The README's opening claim is that the project exists to show that training your own language model is not magic, and that one Colab notebook and five minutes are enough. Everything in the repository supports that framing: a data generator, a tokenizer step, a vanilla transformer, a training loop, and an inference path, each in its own file. If your goal is a capable assistant, this is the wrong repository. If your goal is to stop treating a transformer as a black box, the scope is unusually complete for its size.

## The architecture is deliberately boring

The README's architecture table lists 8.7M parameters, 6 layers, hidden dimension 384, 6 attention heads, a 768-wide feed-forward network with ReLU, a 4,096-token BPE vocabulary, a 128-token maximum sequence length, LayerNorm, learned positional embeddings, and an LM head weight-tied to the embedding matrix. The README calls it a vanilla transformer and says explicitly that there is no GQA, no RoPE, no SwiGLU, and no early exit.

That list of omissions is the interesting part. Modern small models usually adopt grouped-query attention to shrink the KV cache and rotary embeddings to generalize past the training length. GuppyLM takes neither. Learned positional embeddings mean the model has no positional signal at all beyond position 128, and the README's own note about interactive chat confirms the consequence: the conversation grows and quickly runs into the 128-token limit, reducing quality. Weight tying between the embedding and the output head is the one efficiency choice present, and at this scale it saves roughly 1.5M parameters, which is not trivial when the total is 8.7M. The design reads as a teaching default rather than a tuned one, and the README presents it that way.

## Chatting with Guppy locally

The Quick Start section gives a local path. The dependencies are torch and tokenizers, and the entry point is the guppylm package as a module. Run the install line first, then the chat command with no arguments to enter interactive mode.

```bash
pip install torch tokenizers
python -m guppylm chat
```

What you should see is a You> prompt. Type a line and Guppy answers in lowercase, in one or two short sentences. The README shows a cat prompt answered with a line about hiding behind the plant, and a rain prompt answered with a line about rain being the best thing about outside. Because the model is small and the context is capped at 128 tokens, the README warns that quality degrades as the interactive conversation grows. For a cleaner signal, use the single-prompt form, which prints one response and exits:

```bash
python -m guppylm chat --prompt "tell me a joke"
```

The README's sample output for a joke request is a fish-themed pun about a wall. Note that the local path assumes the model weights are available; the README's other routes are the browser demo, which downloads a quantized ONNX model of roughly 10 MB and runs inference through WebAssembly, and the use_guppylm.ipynb Colab notebook, which the README says downloads the pre-trained model from HuggingFace.

## Training your own copy in Colab

The training route is the train_guppylm.ipynb notebook. The README's instructions are three steps: set the runtime to a T4 GPU, run all cells, and then either upload the result to HuggingFace or download it locally. The notebook is described as downloading the dataset, training the tokenizer, training the model, and testing it.

The data comes from the arman-bd/guppylm-60k-generic dataset on HuggingFace, with 60,000 samples split into 57,000 train and 3,000 test, 60 categories, and a JSON shape of input, output, and category. Loading it is a two-line exercise, and the README shows the first training row as a greeting exchange.

```python
from datasets import load_dataset
ds = load_dataset("arman-bd/guppylm-60k-generic")
print(ds["train"][0])
```

If you want the training run to publish somewhere other than a manual upload, the repository ships a .env.example with three keys: HF_TOKEN, HF_REPO, and HF_DATASET, pre-filled with placeholder values such as hf_your_token_here and your-username/guppylm-9m-chat. The README does not document which script reads that file or whether the notebook consumes it, so treat the environment file as a hint about intended configuration rather than a documented interface. There is also a Makefile with a single notebook target that runs python3 tools/make_colab.py.

## Where GuppyLM breaks down

The 128-token context is the most visible failure mode, and the README names it rather than hiding it. Any multi-turn conversation drifts out of distribution once the accumulated history passes that boundary, and there is no truncation or summarization strategy described to manage it. The same limit rules out anything that needs a long document in the prompt.

The second limit is knowledge. The model was trained on 60K synthetic conversations generated by template composition across 60 topics. The README's own sample dialogue makes the boundary explicit: asked about the meaning of life, Guppy answers food. That is charming in a demo and useless in a support bot. Anything outside the fish persona will produce confident nonsense, because the training distribution contains nothing else.

The third limit is evaluation. The repository has an eval_cases.py file of held-out test cases, but the README does not report scores, a baseline, or a comparison against any other model. There is no published number that tells you whether a training run went well. If you need a measurable quality bar before adopting a model, this project does not supply one.

## How it compares with nanoGPT and a pretrained small model

The closest comparison is nanoGPT, the widely used minimal GPT training repository. Both are single-file-scale educational transformers, but they differ in what they leave to the reader. nanoGPT trains on a plain text corpus you supply and focuses on the training loop and sampling code. GuppyLM adds the pieces before and after training: generate_data.py produces the conversation corpus from templates across 60 topics, dataset.py handles batching, eval_cases.py holds test cases, and the package exposes a chat entry point. The trade-off is that GuppyLM's data is synthetic and narrow by construction, while nanoGPT's corpus is whatever you bring. If you want to learn how a training loop works on real text, nanoGPT is the more direct instrument. If you want to see a full pipeline including data synthesis and an inference interface, GuppyLM covers more ground.

The other comparison is a pretrained small model pulled from HuggingFace. A model in the few-hundred-million-parameter range will answer general questions, follow instructions, and handle far more than 128 tokens. It will also teach you nothing about tokenizer training or positional embeddings, because you never see them. GuppyLM is not competing on capability and the README does not pretend otherwise.

## Licence, maintenance, and what a fork costs you

The README carries an MIT badge that links to a LICENSE file at the repository root, so the project presents itself as MIT-licensed. The repository metadata does not list a licence, which means the authoritative answer is the text of that file, not the badge. Read it before you redistribute weights or the dataset, and note that the HuggingFace model and dataset are separate artifacts with their own pages.

On maintenance: the repository is not archived, and the last push was on 2026-04-15. The project is small and mostly static, so a quiet period is not alarming on its own, but it does mean that issues filed against the current dependency floors may sit unanswered. The requirements.txt pins torch>=2.0.0, tokenizers>=0.19.0, tqdm>=4.65.0, numpy>=1.24.0, and datasets>=2.14.0, all as lower bounds with no upper bound. A future major release of any of those libraries could break the training loop without a corresponding commit here. For an educational notebook you run once, that risk is small. For anything you intend to keep running, pin the versions yourself.

## Conclusion

Adopt GuppyLM if you want a complete, readable pipeline that goes from generated text to trained weights in one notebook, and if you are fine with a model whose vocabulary tops out at 4,096 BPE tokens and 128-token context. Do not adopt it as a production assistant or as a base for domain fine-tuning; an 8.7M-parameter vanilla transformer with learned positional embeddings has no room for that. Before you commit, check the LICENSE file at the repository root, since the README badge says MIT but the repository metadata does not list a licence, and confirm that the HuggingFace model arman-bd/guppylm-9M still resolves.

## FAQ

### What does LLM stand for in the context of GuppyLM?

LLM stands for large language model, the general class of model GuppyLM belongs to. GuppyLM is a very small member of that class at 8.7M parameters, trained from scratch on 60K synthetic conversations.

### What are the four types of LLM?

The README does not define a taxonomy of LLM types, so this cannot be answered from the repository. What it does state is that GuppyLM is a vanilla transformer with no GQA, no RoPE, no SwiGLU, and no early exit.

### What does LLM mean in AI?

In AI, LLM refers to a large language model, a model trained to predict text. GuppyLM is one such model, scaled down to 8.7M parameters and 128-token sequences so it can be trained in a single Colab notebook.

### What is LLM in ChatGPT?

The README does not describe ChatGPT or its architecture, so no comparison can be made from it. It only positions GuppyLM against the idea that training a language model requires a PhD or a GPU cluster.

## Sources

- [arman-bd/guppylm on GitHub](https://github.com/arman-bd/guppylm)
- [Issues](https://github.com/arman-bd/guppylm/issues)
- [README](https://github.com/arman-bd/guppylm/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/arman-bd-guppylm
