Model or dataset
google-deepmind/simply avatar
google-deepmind/simply

google-deepmind/simply: a minimal JAX codebase for LLM research

Minimal and scalable research codebase in JAX, designed for rapid iteration on frontier research in LLM and other autoregressive models.

575 stars69 forksPythonApache-2.0

At a glance

What is it?
Simply is a small JAX training and research codebase from Google DeepMind, built so that humans and coding agents can fork it and iterate on optimizers, losses and RL algorithms quickly. It is minimal by design, which is also its main constraint.
Who is it for?
Adopt Simply if you already know JAX and your work is changing the internals of a training loop: a new optimizer, a new loss, an RL algorithm. Do not adopt it if you want a supported training framework with a stable API, a model zoo, or a serving product; the README frames the project as a codebase to fork and hack, not as a platform.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Simply is for, and who it is actually aimed at

Simply is a research codebase in JAX for LLM and other autoregressive models. The README states the goal directly: it is "an environment where both humans and AI agents can rapidly iterate on frontier LLM research." That phrasing sets the audience. This is not a library you import into an existing service, and it is not a framework with a compatibility promise. It is a repository you fork, read and modify.

The design bet is that time-to-new-idea matters more than feature coverage. The README lists four properties, and the first two are about removing friction: quick to fork and hack, and minimal abstractions and dependencies. The claim attached to the second one is worth quoting because it is the whole pitch: learn JAX, and you are ready to read and hack the code. If you do not already know JAX, the onboarding cost lands on you, not on the project.

The other two properties are less conventional. Simply ships an agent harness under simply/agent/ with a Bash tool and context management, and vendors a separate Go harness called Amplio under amplio/ for long-horizon runs. The intended loop is that an LLM served by Simply can read the code, propose changes, run experiments and iterate, either autonomously or with a human steering. That is an unusual thing for a training codebase to include, and it explains why the repository carries AGENTS.md, CLAUDE.md and GEMINI.md at the top level alongside CONTRIBUTING.md.

The mechanism: JAX, Orbax, Grain and a config-driven entry point

The dependency list in pyproject.toml is short and readable, which is the point. JAX handles model and training, Orbax handles checkpoint management, and Grain handles the data pipeline. Around those sit einops, sentencepiece, tokenizers, tensorboard, tensorboardX and wandb. There is no configuration framework, no experiment tracker abstraction layer, and no plugin system.

The entry point is a module, not a script. The README runs experiments as python -m simply.main with an --experiment_config flag naming a config (lm_test, lm_no_scan_test) and an --experiment_dir flag naming where the run writes. That shape tells you how the codebase is organised: named configs select behaviour, and the directory is the run's identity. Because there is no registry or database, the experiment directory is the only record of a run beyond whatever you log to TensorBoard or Weights & Biases.

Two flags in the example commands reveal the debugging model. Setting JAX_DISABLE_JIT=True disables JIT, and the lm_no_scan_test config disables use_scan. In a JAX codebase, those two features are exactly what make stepping through a training loop in a normal Python debugger awkward, so turning them off for a small test run is the intended way to get printable arrays. The cost is obvious: without jit and scan you are not measuring anything about performance, only about correctness.

The serving stack is separate and talks gRPC. Its Python stubs are generated from simply/serving/*.proto rather than checked in, so a fresh clone cannot import the serving code until you run the generator. That is a deliberate choice to keep generated files out of the tree, and it adds one setup step that a checked-in stub would not.

Installing Simply and running a first local test

JAX installation is environment-specific, and the README defers to the JAX installation page for it. Pick the line that matches your hardware, then install Simply itself.

bash
# CPU:
pip install -U jax
# GPU:
pip install -U "jax[cuda13]"
# TPU:
pip install -U "jax[tpu]"

pip install .

The extras in pyproject.toml cover the optional pieces: tfds for TensorFlow Datasets, math-eval for simply/utils/math_eval.py, dev for pytest, serving for the gRPC stack, and agent for the built-in harness. If you plan to use the agent, install it explicitly.

bash
pip install ".[agent]"

Assets come from HuggingFace through a setup script. It writes models to ~/.cache/simply/models/ and datasets to ~/.cache/simply/datasets/, and the README notes that only a few datasets and models are included for testing at present.

bash
pip install huggingface_hub
python setup/setup_assets.py --models-only

With assets in place, the smallest useful run is the local test config. It writes into a temporary directory and logs to stderr, so you can confirm the pipeline works before touching an accelerator.

bash
EXP=simply_local_test_1; rm -rf /tmp/${EXP}; python -m simply.main --experiment_config lm_test --experiment_dir /tmp/${EXP} --alsologtostderr

If you want to step through the loop with a debugger, swap in the no-scan config and disable JIT first. Expect this run to be slow; it exists to make arrays printable, not to train anything.

bash
export JAX_DISABLE_JIT=True; EXP=simply_local_test_1; rm -rf /tmp/${EXP}; python -m simply.main --experiment_config lm_no_scan_test --experiment_dir /tmp/${EXP} --alsologtostderr

There is also a uv path that does environment and assets together: uv sync, then uv run python setup/setup_assets.py. The repository ships a uv.lock, so this route pins the dependency set for you.

The agent harness and Amplio, and what they cost you

Simply's built-in harness is invoked as a module with three arguments: a task file, an environment, and an LLM. The README's example points the environment at the current directory and the model at a Vertex AI Gemini endpoint through LiteLLM.

bash
python -m simply.agent.main \
    --task_file=simply/agent/example_tasks/code_stats.md \
    --env="Local:." \
    --llm="LiteLLM:vertex_ai/gemini-2.5-pro"

The task file is markdown in the repository's example_tasks directory, which is why python-frontmatter and jinja2 appear in the agent extra. The harness is small by description: a Bash tool and context management. Treat it as a starting point rather than a finished product.

Amplio is a different animal. It is vendored under amplio/, written in Go, and built with make build, which the README says needs Go (see amplio/go.mod) and Node.js 22 or newer. Running ./amplio serve prints a URL with an access token, and it reads system_llm_hq and system_llm_fast from <data-dir>/config.toml. Two model keys, one for a stronger model and one for a faster one, is a routing decision baked into the configuration surface.

The cost here is real. You now maintain two toolchains, Python and Go, plus a Node build for the web UI. The payoff Amplio claims is DB-first persistence, so a long run resumes after a crash. If your research runs are short and interactive, that payoff does not apply to you, and the second toolchain is pure overhead.

Where Simply is the wrong tool

The README's own framing is the first limitation: it is a codebase to fork and hack. There is no API stability statement, no deprecation policy, and no release history in the repository. Version 0.3.9 in pyproject.toml signals pre-1.0, and requires-python is >=3.12, so you cannot run it on an older interpreter without changing the project metadata.

Asset coverage is thin by the project's own admission. The README says only a few datasets and models are included for testing and that more will be added. If your work depends on a specific pretrained checkpoint or corpus, check the download script's contents before you plan around it; nothing in the README promises a particular model is there.

The serving stack has a setup trap. The stubs are generated, not checked in, so importing simply/serving/ fails until you install the serving extra and run python setup/gen_protos.py. A team that clones the repository and expects everything to import will hit this.

Finally, the minimal-abstractions stance cuts both ways. There is no abstraction layer between you and JAX, which is what makes the code readable, and also means there is no place to plug in a different backend, a different checkpoint format, or a different data loader without editing the codebase. If you need to train on a cluster your organisation already runs with its own scheduler and storage conventions, the GCloud and GKE guides in docs/ describe one path, and adapting the code to another is your work.

How it differs from a full training framework

The natural comparison is a framework such as Levanter or torchtitan, or a higher-level library such as Hugging Face Transformers' Trainer. The difference is not features; it is where the abstraction boundary sits.

A framework like Levanter or torchtitan gives you a training loop you configure and extend through defined hooks, with the project owning sharding, checkpointing and resumption behaviour. Simply inverts that. Orbax and Grain are used directly, and the training loop is code you read and edit. The README's stated aim is minimising the time to implement a new optimizer, loss or RL algorithm, and that aim is incompatible with a thick extension API, because the API would be the thing you fight.

Transformers' Trainer sits further away still: you supply a model and a dataset, and the library owns the loop. That is the right trade for fine-tuning an existing architecture. It is the wrong trade when the architecture or the update rule is the research question.

The agent angle has no direct equivalent in either. Levanter and torchtitan do not ship an agent harness or a vendored long-horizon runner. Simply does, and the README's prompt example has an agent read simply/utils/optimizers.py, run the Adam baseline on lm_test, then propose and benchmark new optimizers across iterations. Whether that loop produces useful research is an open question, but the scaffolding for it is in the repository rather than in a separate tool.

Maintenance, licensing and what an upgrade costs

The repository is not archived, and the last push was on 2026-09-09, which is recent. No releases are listed for this project, so the version in pyproject.toml is the only version signal available, and upgrades are a git pull plus a dependency reinstall rather than a package-manager bump.

That has a practical consequence. Because Simply is meant to be forked, upstream changes arrive as diffs against code you have already modified. There is no changelog and no migration guide in the repository, so a pull that touches the training loop or the config schema lands directly in your working tree. Teams that fork should expect to review upstream commits by hand rather than merge a version.

The dependency set is the other upgrade surface. JAX is pinned at >=0.10, and the gpu and tpu extras carry platform markers (sys_platform == 'linux'), so the accelerator path is Linux-only as declared. The tfds extra pulls a different TensorFlow build on macOS than elsewhere. None of these are unusual, but they mean your lockfile, not the project, decides what you actually run.

Licensing is Apache-2.0, which permits commercial and modified use and includes a patent grant. Amplio is vendored inside the repository, so its files ship under whatever the repository's licence states unless amplio/ carries its own notice; check that directory before redistributing. This is a description of the licence identifier, not legal advice.

Editorial conclusion

Adopt Simply if you already know JAX and your work is changing the internals of a training loop: a new optimizer, a new loss, an RL algorithm. Do not adopt it if you want a supported training framework with a stable API, a model zoo, or a serving product; the README frames the project as a codebase to fork and hack, not as a platform. Before committing, verify two things yourself: that your accelerator path matches one of the extras in pyproject.toml (gpu, tpu), and that the datasets and checkpoints under ~/.cache/simply/ cover what you need, since the README states only a few are included for testing.

Frequently asked questions

What does google-deepmind/simply mean as a project name?

It is the name of a minimal JAX research codebase for LLM and other autoregressive models, described in the README as an environment where humans and AI agents can rapidly iterate on frontier research. The name is a pun on the project's stated goal of minimal abstractions and dependencies.

When should I use google-deepmind/simply instead of a full training framework?

Use it when the research question is inside the training loop: a new optimizer, a new loss, or an RL algorithm. The README's aim is minimising the time to implement those ideas, which is why the codebase keeps abstractions and dependencies minimal rather than offering an extension API.

How do I install google-deepmind/simply?

Install JAX for your hardware first, since the README says that step is environment-specific, then run pip install . from the repository root. Optional pieces come from extras such as tfds, math-eval, serving, agent and dev, and assets are downloaded separately with python setup/setup_assets.py.

Does google-deepmind/simply include model checkpoints and datasets?

The setup script downloads them from HuggingFace into ~/.cache/simply/models/ and ~/.cache/simply/datasets/, and the locations can be changed with --models-dir, --datasets-dir or the SIMPLY_MODELS and SIMPLY_DATASETS environment variables. The README states that only a few datasets and models are included for testing at present.

Why does importing google-deepmind/simply's serving code fail after a fresh install?

The serving stack talks gRPC and its Python stubs are generated from simply/serving/*.proto rather than checked in. Install the serving extra and run python setup/gen_protos.py once after install to create them.

Official sources

  1. google-deepmind/simply on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-deepmind-simply.svg)](https://hysenlabs.com/projects/google-deepmind-simply)