# Tinker Cookbook: post-training recipes for models you fine-tune through the Tinker API

> Tinker Cookbook is a Python library of supervised fine-tuning, RL, DPO and distillation recipes built on top of the Tinker training SDK. It is for engineers who want working post-training loops without running the distributed training stack themselves.

**thinking-machines-lab/tinker-cookbook** — Post-training with Tinker

- Repository: https://github.com/thinking-machines-lab/tinker-cookbook
- Stars: 4,166 · Forks: 544
- Language: Python
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/thinking-machines-lab-tinker-cookbook

## What Tinker Cookbook solves, and who it is written for

Fine-tuning a language model normally means owning a training stack: checkpointing, gradient accumulation, distributed placement, and a sampler that reflects the weights you just updated. Tinker Cookbook takes the position that this is infrastructure work, not research work. It is a Python library of post-training implementations that talk to the Tinker API, where the README says you "send API requests to us and we handle the complexities of distributed training."

The audience is narrow on purpose. You need a Tinker account and an API key before any of it runs. Within that boundary the library is broad: the recipes folder covers chat supervised fine-tuning on conversational datasets such as Tulu3, math RL with verifiable rewards, code RL with sandboxed execution, DPO and a three-stage RLHF pipeline, on-policy and off-policy distillation, retrieval-augmented RL, multi-agent self-play and cross-play, audio SFT and RL, and VLM image classification. Each recipe ships with its own README describing implementation details, launch commands and expected results.

The distinction that matters is between the two libraries in the same repository. tinker is the SDK: you create a training client, call forward_backward, optim_step, save_state, load_state. tinker-cookbook is the layer above it that turns those primitives into loops you can run. If you want to understand the mechanics, the sl_loop.py and rl_loop.py recipes are minimal examples of the primitives; sl_basic.py and rl_basic.py are minimal examples of the configured recipes.

## The training loop the cookbook actually builds

The mechanism is visible in the README's own snippet. A ServiceClient is created, and from it you create a LoRA training client for a named base model with a rank. Every subsequent operation goes through that client: forward and backward, optimizer step, save state, load state. Sampling is not a separate service. You call save_weights_and_get_sampling_client, which returns a client whose weights match the training client's current state, and you sample from it. That round trip is the core of on-policy recipes: generate, score, update, regenerate.

The renderer layer is where the cookbook earns its name. Models do not share tokenization or prompt formatting, and the repository depends on tml-renderers for that work. For the Inkling family the README states that the appropriate renderer and tokenizer are selected automatically for any thinkingmachines/Inkling* model, including the :peft: long-context variants, so you pass a model name and the rendering path follows. The cookbook also carries tiktoken, which pyproject.toml annotates as required for the Kimi tokenizer.

Recipes are configured, not scripted from scratch. Each one is a directory with its own README, and the top-level recipes README is the index. That structure is the design bet: instead of a single training script with flags, you get a set of runnable configurations you copy and modify. The cost is that a recipe's assumptions about dataset format and reward function are baked in, and adapting one to a different task means reading its README rather than setting an option.

## Installing tinker-cookbook and running a first recipe

The README gives three steps, and the first two happen outside the terminal. Sign up for Tinker, then create an API key from the console and export it as TINKER_API_KEY. Nothing in the library works without that variable set, so do it before installing anything.

The installation command pulls the tinker SDK in as a dependency, so you do not install it separately:

```bash
uv pip install tinker-cookbook
```

If you want the nightly build instead of the latest stable release, the README gives this alternative, which installs from the nightly branch of the repository:

```bash
uv pip install 'tinker-cookbook @ git+https://github.com/thinking-machines-lab/tinker-cookbook.git@nightly'
```

Python 3.11 or newer is required, and torch>=2.10 is a hard dependency because tml_renderers uses libtorch's stable C ABI, which pyproject.toml notes was added in 2.10. If your environment pins an older torch, the install will fight you.

Many recipes need extras that are not in the base install. The optional dependency groups in pyproject.toml are named math-rl, audio, multiplayer-rl, modal, vector-search, cloud, wandb, tutorials and dev. A math RL run, for example, needs the math-rl group for math-verify, pylatexenc and sympy. The tutorials group pulls in marimo and matplotlib plus tinker_cookbook[math-rl]; the README says you can run any of the 20+ notebooks with:

```bash
marimo edit tutorials/101_hello_tinker.py
```

That opens the notebook in the marimo editor rather than executing it headlessly. Before running a full recipe, the cheapest check is the minimal example in tinker_cookbook/recipes/sl_basic.py, which the README describes as a minimal configuration for supervised learning. If that completes, your key, model name and rendering path are all working, and a failure inside a larger recipe is a data or reward problem rather than a setup problem.

## Where the API dependency becomes a real limitation

Every training run in this library leaves your machine. The Tinker service performs the forward and backward passes, and the cookbook's job is to orchestrate requests against it. That is the trade the project makes, and it has consequences the README does not soften.

You cannot train a model that Tinker does not serve. The README's example uses meta-llama/Llama-3.2-1B, and the Inkling documentation covers the thinkingmachines/Inkling family, but the set of available base models is defined by the service, not by the repository. If your work depends on a checkpoint that is not on that list, this library cannot help you regardless of how well a recipe matches your task.

You also cannot train offline, and you cannot keep training data on your own hardware. Datasets and prompts go to the API. For teams with data residency constraints or a policy against sending training corpora to a hosted service, that is a disqualifying property, not a configuration problem.

A third constraint is less obvious. Because weights and checkpoints live behind the service, the cookbook's persistence model is the save_state and load_state pair plus checkpoint archive retrieval. The README shows downloading weights through a REST client, calling get_checkpoint_archive_url_from_tinker_path on a sampling client's model path and writing the result to a tar.gz file. That is a real escape hatch for exporting a finished model, but it is a download step, not a local checkpoint directory you can mount, diff or resume from without the service.

## Tinker Cookbook against writing your own loop on the SDK

The honest alternative is not a different framework. It is using the tinker SDK directly, which the same repository ships and which the cookbook depends on. The difference is one of level, and it is worth being precise about what you give up in each direction.

The SDK gives you five primitives and nothing else: create a LoRA training client with a base model and rank, call forward_backward, call optim_step, save and load state, and get a sampling client from saved weights. That is enough to write a training loop, and sl_loop.py and rl_loop.py exist precisely as demonstrations of how little is required. What the SDK does not give you is a dataset pipeline, a renderer selection policy, a reward function interface, or a logging convention. You write those, and you own their bugs.

The cookbook gives you those pieces pre-assembled, at the cost of adopting its structure. A recipe directory with a README and a launch command is easier to run and harder to bend. If your task is close to one of the listed recipes (chat SFT, math RL, code RL, preference learning, distillation, tool use, multi-agent, audio, VLM classification) the cookbook saves real time. If your task is adjacent to all of them, you will spend the same effort reading recipe internals as you would have spent writing the loop, and you will have less freedom in how it is organized.

A middle path exists and the repository supports it: use the SDK primitives for the training step and borrow only the pieces of the cookbook you need, since the library is a dependency you can import from rather than a framework that owns your entry point.

## Maintenance, versioning and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-23. Releases follow a stable-plus-nightly pattern: v0.5.7 on 2026-09-03, v0.5.6 on 2026-09-01, and a nightly build tagged 0.5.8.dev5+g1e53aa3d1 on 2026-09-23. The gap between v0.5.6 and v0.5.7 is two days, which suggests releases track merged work closely rather than following a long cadence.

That pattern has an upgrade cost. A nightly channel exists and the README documents installing from it, which means the project expects some users to run ahead of the stable release. If you install from the nightly branch, you are tracking a moving target and should expect the version string to change under you. The safer default is the PyPI release, and the version constraint on transformers in pyproject.toml (>=4.57.6 with several excluded 5.x releases up to <=5.5.4) shows the maintainers are actively pinning around upstream breakage rather than letting resolution float.

Licensing is Apache-2.0, declared both in the repository and in pyproject.toml. That is a permissive licence with an explicit patent grant, and it imposes no copyleft obligation on your own code. It says nothing about the Tinker service itself: the licence covers this library, while access to training and sampling is governed by your Tinker account and whatever terms apply to it. The README does not document rollback or downgrade procedures for a bad upgrade, so pin a version in your own environment if reproducibility matters to you.

## Conclusion

Adopt Tinker Cookbook if you already have a Tinker API key and want a working supervised or reinforcement loop today instead of assembling one from primitives. Do not adopt it if you need to train a model the Tinker service does not serve, or if you cannot send training data to a hosted API. Before committing, verify that your base model appears in the Tinker model list, that torch>=2.10 resolves in your environment, and that your recipe's extra dependencies are installed.

## FAQ

### How much does Tinker cost to use?

The README does not state pricing. It links to a Models & Pricing page on the Tinker documentation site for per-model context lengths and pricing, and installation requires signing up for Tinker and creating an API key from the console.

### Is Tinker free to use?

The README does not say. It describes a sign-up flow and an API key created from the Tinker console, and directs readers to a separate Models & Pricing page, so cost information lives outside this repository.

### Can you provide a list of Tinker models?

The README names meta-llama/Llama-3.2-1B in its SDK example and the thinkingmachines/Inkling and thinkingmachines/Inkling-Small models, including :peft: long-context variants. It points to the Models & Pricing documentation page for the full set of per-model context lengths.

### How does Thinking Machines Tinker work?

You send API requests to the Tinker service and it handles distributed training. In the cookbook you create a LoRA training client for a base model, call forward_backward and optim_step, then get a sampling client from the saved weights and sample from it.

## Sources

- [Issues](https://github.com/thinking-machines-lab/tinker-cookbook/issues)
- [License: Apache-2.0](https://github.com/thinking-machines-lab/tinker-cookbook/blob/main/LICENSE)
- [README](https://github.com/thinking-machines-lab/tinker-cookbook/blob/main/README.md)
- [Releases](https://github.com/thinking-machines-lab/tinker-cookbook/releases)
- [thinking-machines-lab/tinker-cookbook on GitHub](https://github.com/thinking-machines-lab/tinker-cookbook)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/thinking-machines-lab-tinker-cookbook
