Model or dataset
rasbt/reasoning-from-scratch avatar
rasbt/reasoning-from-scratch

reasoning-from-scratch: a code-first path from a pretrained LLM to an RL-trained reasoner

Implement a reasoning LLM in PyTorch from scratch, step by step

5,242 stars820 forksJupyter NotebookApache-2.0

At a glance

What is it?
Sebastian Raschka's repository turns a Qwen3 base model into a small reasoning model across eight chapters of notebooks, covering inference-time scaling, GRPO and distillation. It is a teaching codebase tied to a book, not a library you drop into production.
Who is it for?
Adopt it if you want to read and run the mechanics of inference-time scaling, GRPO and distillation on a Qwen3 base model, and you accept that the code follows a book rather than a stable API. Do not adopt it if you need a supported inference server or an RL training framework for production workloads; nothing in the repository claims that role.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What reasoning-from-scratch actually teaches, and who the notebooks are written for

The repository is the official code for Build a Reasoning Model (From Scratch), published by Manning. Its stated scope is narrow and explicit: it starts from a pretrained open-source base LLM, Qwen3, and adds reasoning behaviour on top of it. Inference-time scaling, reinforcement learning and distillation are the three families of method the README names. The book does not re-implement the base model itself; for that the README points readers to the author's earlier book on building a large language model.

That split matters when you decide whether to clone it. If your question is how a reasoning model is trained, the notebooks are the material. If your question is how a transformer is written from scratch, this repository assumes you already have that ground and only shows the chapter on Qwen3 source code as an appendix. The intended reader is someone who learns by running code and reading the diff between one chapter and the next, not someone shopping for a dependency.

The chapter sequence: inference-time scaling first, then reinforcement learning, then distillation

The table of contents in the README lays out a progression that is worth reading before you touch any notebook. Chapter 2 generates text with a pretrained LLM. Chapter 3 evaluates reasoning models and brings in sympy as a verifier. Chapters 4 and 5 stay at inference time: scaling, then self-refinement. Chapter 6 moves to training with reinforcement learning, chapter 7 improves GRPO, and chapter 8 distils a reasoning model for cheaper inference.

The ordering is the argument. A reader sees that you can get measurable behaviour change without touching weights at all, and only then meets the training loop. Each chapter folder holds a main notebook plus an exercise-solutions notebook, and the appendices cover Qwen3 source, larger LLMs, batching, evaluation approaches and a chat interface built with chainlit. The mental-model diagram in the README summarises the techniques, but the notebooks are where the data flow lives: prompts in, generated traces out, scored by a verifier, then fed back into training in the later chapters.

Installing reasoning-from-scratch and running the first chapter notebook

The README gives a shallow clone as the way to get a copy, and the repository also ships a pyproject.toml, so the package installs with pip. The declared Python range is >=3.10,<3.15 and torch is pinned to >=2.10,<3, which is tighter than most PyTorch projects and is the first thing to check against your environment.

bash
git clone --depth 1 https://github.com/rasbt/reasoning-from-scratch.git
cd reasoning-from-scratch
pip install -e .

Installing in editable mode gets you jupyterlab, tokenizers, sympy and matplotlib alongside torch. Two optional groups exist: an extra group for the MMLU datasets used in Appendix F and chainlit for the chat UI in Appendix G, with chainlit marked for Python below 3.14. The dev group carries pytest, ruff, huggingface-hub, safetensors and transformers.

From there, open the chapter 2 notebook, which the README lists as ch02/01_main-chapter-code/ch02_main.ipynb, and run it. That chapter's job is generating text with the pretrained model, so the observable result is generated output from a base LLM before any reasoning method has been applied. The README states the code uses GPUs automatically when they are present and that the main chapters are designed to run on consumer hardware within a reasonable timeframe. It does not promise a specific runtime, and you should not assume one.

Where the repository stops short: book-shaped code, consumer hardware, no server

This is a book companion, and the code is organised around chapters rather than around a stable public API. The package name reasoning_from_scratch appears in pyproject.toml with a dynamically read version, and requirements.txt pins reasoning-from-scratch >= 0.1.2, but nothing in the README describes a compatibility policy or a deprecation process. If you build on internal notebook functions, expect to follow the book's revisions rather than a semver contract.

The hardware claim is deliberately hedged. The README says the main chapters are designed to mostly run on consumer hardware within a reasonable timeframe, and that the code uses GPUs automatically if available. It does not say which GPU, how much memory, or how long any chapter takes. The appendices that scale to larger LLMs and to batching exist precisely because the main path is small by design. A small-but-functional model for educational purposes is the stated goal, so a reader who wants a competitive reasoning model is looking at the wrong repository.

There is also no inference server here. Appendix G builds a chat interface with chainlit, but that is a demonstration surface, not a deployment target. No serving layer, no batching scheduler beyond the appendix on throughput-oriented execution, no observability. Treat it as a lab.

How this differs from TRL and from the LLMs-from-scratch repository

The nearest functional overlap is Hugging Face TRL, which provides maintained trainers for reinforcement learning from human feedback and related methods. The difference is one of intent. TRL is a library you call; reasoning-from-scratch is code you read and modify, and the README frames the whole project as learning how a reasoning LLM works by adding the capability yourself step by step. GRPO appears here as a chapter you implement and then improve in the next chapter, not as a trainer you configure.

The second comparison is internal to the author. Build a Large Language Model (From Scratch) covers how a conventional base LLM is implemented, and the README explicitly calls this book standalone and focused on reasoning methods. So the two repositories are complements, and the README says so: this one starts from Qwen3 weights and adds reasoning on top. If you already own the earlier book's code, you still need this one for inference-time scaling, GRPO and distillation.

Licence, maintenance and the cost of keeping up

The licence is Apache-2.0, declared both in the LICENSE file and in the pyproject classifier, with the source headers carrying a copyright notice for Sebastian Raschka. Apache-2.0 is permissive and includes a patent grant; the practical implication for a reader is that reusing notebook code in your own project is allowed under the usual attribution terms. That is a statement about the licence text, not legal advice, and the book itself is a separate commercial product with its own terms.

The last push to the default branch was on 2026-09-06, and the most recent release is v1.0 from 2026-05-18. The table of contents is marked in progress, so chapter content can still change. Upgrade cost is dominated by the torch pin: torch>=2.10,<3 and Python >=3.10,<3.15 mean a future Python release will lock you out until the pin moves. The repository ships a uv.lock alongside requirements.txt, so a uv-based install gives you a resolved environment; the README itself only shows the git clone, and points to chapter 2 for guidance on installing Python and managing packages.

Editorial conclusion

Adopt it if you want to read and run the mechanics of inference-time scaling, GRPO and distillation on a Qwen3 base model, and you accept that the code follows a book rather than a stable API. Do not adopt it if you need a supported inference server or an RL training framework for production workloads; nothing in the repository claims that role. Before you start, check the pyproject.toml Python and torch bounds, then open ch02/01_main-chapter-code/ch02_main.ipynb and confirm the model download works on your machine.

Frequently asked questions

How do you build a reasoning model from scratch with this repository?

Clone the repository, install it with pip, then work through the chapter notebooks in order. Chapters 4 and 5 cover inference-time scaling and self-refinement, chapter 6 covers reinforcement learning, chapter 7 improves GRPO, and chapter 8 covers distillation.

Can ChatGPT do reasoning?

The README states the methods in the book mirror approaches used in creating large-scale reasoning models such as DeepSeek R1 and GPT-5 Thinking. The repository itself does not evaluate ChatGPT or any hosted model; it works with a pretrained open-source Qwen3 base model.

What is a reasoning example in the context of reasoning-from-scratch?

The repository treats reasoning as something added on top of a pretrained base LLM through inference-time scaling, reinforcement learning and distillation. Chapter 3 evaluates reasoning models using a verifier that depends on sympy.

What are the 8 elements of reasoning?

The repository does not describe eight elements of reasoning. Its chapter sequence covers generating text with a pretrained LLM, evaluating reasoning models, inference-time scaling, self-refinement, reinforcement learning, GRPO improvements and distillation.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. rasbt/reasoning-from-scratch on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rasbt-reasoning-from-scratch.svg)](https://hysenlabs.com/projects/rasbt-reasoning-from-scratch)