# NanoJev: a 0.6B model that returns decision probabilities instead of generating tokens

> A small research project built on Qwen3-0.6B with decision heads bolted on, evaluated on four arcade tasks against the Jev model it imitates. The README leads with benchmark tables and the repository ships no releases at all.

**TianyuCodings/NanoJev** — A nano replica of Jev: parallel decisions, dynamic candidates, and an end-to-end training pipeline.

- Repository: https://github.com/TianyuCodings/NanoJev
- Stars: 2,506 · Forks: 258
- Language: Python
- License: MIT
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/tianyucodings-nanojev

## The interface: state, question, candidates

NanoJev describes itself as a nano replica of Jev, the decision model TypeScript introduced, and the design goal is stated in one line: states and questions in, complete probability distributions out, with zero output-token decoding.

That last phrase is the whole idea. Instead of asking a language model to generate an answer and parsing it, you hand the model a state, a question, and the candidate answers, and it returns a probability for each one. Nothing is decoded, so nothing needs parsing and nothing can fail to parse.

Five capabilities are listed. Parallel decisions batch independent states, questions and candidate paths through one backbone forward pass. Dynamic candidates let the Choice task return a distribution over between 2 and 255 supplied candidates using a shared scoring head. Boolean and ordered scores predict either a proposition's probability or a distribution and expectation over 2 to 10 ordered levels. Direct probabilities support ranking, selection or sampling without generating answer tokens. And one small backbone, Qwen3-0.6B with decision heads attached, is reused across all four game tasks.

Each head behaves differently: Choice uses set attention followed by a softmax, Boolean uses a sigmoid, and Score returns a probability-weighted level. That is a small amount of machinery on top of a base model, and it is the part worth understanding before judging any of the numbers.

## The benchmark table, read carefully

The headline result is a held-out table of successful episodes on a 274-case test set, using the same observation interface, candidate actions and seeded epsilon-greedy controller across systems. Test and OOD together contain 548 cases per model, and every evaluated trajectory passes independent simulator replay.

| Model | Maze | Snake | Basic | Predict Position |
|---|---:|---:|---:|---:|
| NanoJev | 4/10 | 8/8 | 128/128 | 27/128 |
| Jev | 7/10 | 8/8 | 56/128 | 11/128 |
| Untuned Qwen3-0.6B | 2/10 | 0/8 | 56/128 | 11/128 |

Two observations matter more than the wins. First, the model loses: on the maze, Jev scores 7 of 10 against NanoJev's 4 of 10, and the project says so plainly rather than dropping the column. Second, the large navigation showcases use different settings from this table. The maze figure of 225 attempts for NanoJev against 2,738 for Jev comes from combining local safety probabilities with the same exploration code and remembered open paths, which is a different configuration from the controlled evaluation.

The ViZDoom Predict Position column is the honest weak spot. At 27 of 128, NanoJev more than doubles Jev's 11 of 128 but remains far from solved, and the README's framing is that the policy is learning when to turn, wait and fire rather than that it has learned to play. Treat that column as progress evidence, not as a result.

One naming caveat: the benchmark table in the README reproduces the model name as NanoJev, while the repository is `TianyuCodings/NanoJev` and the README's own links point at a project directory of that name. That is a formatting inconsistency in the docs, not a different project.

## Dataset scale, splits and two names for one checkpoint

The dataset is 18,760 decision questions per target variant, including 16,333 ViZDoom questions, with matched hard-target and soft-target variants covering the same questions across train, dev, calibration, test and OOD splits.

| Task | All five splits | Training split |
|---|---:|---:|
| ViZDoom Predict Position | 11,173 | 6,788 |
| ViZDoom Basic | 5,160 | 3,054 |
| Maze | 1,469 | 653 |
| Snake | 958 | 403 |
| Total per variant | 18,760 | 10,898 |

Expert gameplay is included as well: 896 Predict Position episodes with 17,498 recorded decisions, of which 512 episodes are assigned to training. Of the 10,898 stored training questions, 10,893 pass the target-validity filter.

The checkpoint has two names, which the README explains in a table. `unified-games-v1` is the Hugging Face release tag identifying matching model and dataset snapshots, and you download with `revision="unified-games-v1"`. `hard_lr1e5` is the training experiment: hard one-hot action targets for Predict Position, backbone learning rate 1e-5, decision-head learning rate 1e-4. Both names refer to the same selected model, the step-400 checkpoint, used across the four demos. Training uses complete-question cross entropy.

That two-naming scheme is a genuine usability wrinkle. If you are reproducing the reported numbers, you need the tag for downloads and the experiment name for logs and configs, and nothing in the repository's file listing makes the pair obvious.

## A Python research repository with a Node harness bolted on

GitHub reports the language as Python and the tree supports it: `scripts/`, `configs/`, `docs/`, `research/`, `results/`, `web/` and `assets/`. There is no `requirements.txt`, though. Instead there are three separate requirement files split by task, which is a sensible way to handle a project where the game environments pull very different dependencies.

```text
requirements-shooting-demo.txt
requirements-toy.txt
requirements-vizdoom.txt
```

The `web/` directory then explains the presence of `package.json` and `package-lock.json` in a Python project. The Node side is the browser demo harness, not the model:

```json
"scripts": {
    "demo": "node --env-file=.env scripts/jev_probe.mjs",
    "demo:replay": "python3 -m http.server 8080 --bind 127.0.0.1 --directory web",
    "data:games": "python3 scripts/build_game_decisions.py"
}
```

The package requires Node 22 or newer and declares a single dependency, the `ai` package. The replay script serves the `web/` directory locally so the recorded browser replays of Jev, NanoJev and Untuned Qwen can be watched without the hosted site, which the README notes currently requires access.

One more detail worth flagging: the model and dataset are distributed on Hugging Face rather than through this repository, and the repository has no releases. With 2,393 stars, 246 forks and a last push on 2026-09-21, that is a deliberate distribution choice, but it means version pinning happens by Hugging Face revision rather than by release tag.

## The sibling project needs an LLM, this one does not

The README opens with a pointer to JevHarness, a separate project where an LLM builds task-specific decision harnesses using Jev, with optional refinement using rewards and execution traces, and it ships an interactive demo. The relationship is clear enough: JevHarness is the tool that builds harnesses, NanoJev is a small model trained to make decisions inside one.

The `.env.example` file in this repository contains a single variable, `AI_GATEWAY_API_KEY=`, which belongs to the harness-building path rather than to the decision model. The Node `demo` script loads it with `node --env-file=.env`. So the repository layout carries a trace of the LLM-assisted workflow even though the project's own headline is zero output-token decoding.

Neither fact contradicts the other: a local model can be evaluated without any gateway, and the gateway key is there for the demonstration scripts that build or refine harnesses. But if you clone this repository expecting a self-contained Python package, you will find a Node harness, three separate requirement files, no packaging metadata, and an env file whose only variable you do not need for the model itself.

The repository carries no GitHub topics and no license metadata beyond the MIT `LICENSE` file at the root, so discovery depends entirely on the model and dataset links in the README rather than on search metadata.

## Conclusion

NanoJev is worth reading as a shape of model rather than as a scoreboard. Its bet is that you supply the candidate answers and the model returns a distribution over them, which moves the cost from generation to enumeration and makes the output directly usable for ranking or thresholding. The claims on the README are specific enough to be checked, down to the per-split dataset counts and the fact that Jev beats NanoJev on the maze. The model and dataset are on Hugging Face under the `unified-games-v1` revision, and there is no release history in the repository, so treat the linked snapshots as the artifact of record.

## FAQ

### What is NanoJev and what does it do differently from a normal language model?

It is a Qwen3-0.6B backbone with decision heads attached that returns probability distributions rather than generating answer tokens. You supply a state, a question and the candidate answers, and it scores each candidate, so the output can be ranked, thresholded or sampled directly.

### How does NanoJev compare to Jev on its benchmark tasks?

Mixed, and the README reports it that way. NanoJev leads on ViZDoom Basic at 128 of 128 against 56, and on Predict Position at 27 of 128 against 11, ties on Snake at 8 of 8, and loses on the maze at 4 of 10 against Jev's 7 of 10.

### Where do I download the NanoJev model and dataset?

Both are on Hugging Face rather than in the GitHub repository, under the C-Tianyu organisation. The model and dataset pair is pinned with the revision unified-games-v1, which corresponds to the step-400 checkpoint from the hard_lr1e5 training run.

### Does the NanoJev repository have releases or a package I can pip install?

There are no releases and no packaging metadata. The Python side is a set of scripts with three separate requirements files split by task, so you work from a clone. The browser demos are Node scripts that require Node 22 or newer.

## Sources

- [Issues](https://github.com/TianyuCodings/NanoJev/issues)
- [License: MIT](https://github.com/TianyuCodings/NanoJev/blob/main/LICENSE)
- [README](https://github.com/TianyuCodings/NanoJev/blob/main/README.md)
- [TianyuCodings/NanoJev on GitHub](https://github.com/TianyuCodings/NanoJev)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tianyucodings-nanojev
