# The NeoHorse repository holds a report PDF and two example scripts, not the models

> NeoHorse is a family of open-weight language models for agent workflows, in two sizes, released under Apache 2.0 across three distribution channels. What is in the repository is a technical report, a figure directory, a subdirectory for the newer decision model, and two short example scripts. The interesting reading is in the evaluation table, where the headline average is computed over a subset that the project's own base model is not part of.

**TokenRhythm/NeoHorse** — NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.

- Repository: https://github.com/TokenRhythm/NeoHorse
- Stars: 1,525 · Forks: 19
- Language: Unknown
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/tokenrhythm-neohorse

## The repository has no code in it: a report PDF, a figure folder and two scripts

The top level of the tree is six entries. A licence, this README, a technical report as a PDF, a directory of figures, a directory of examples, and a directory named after the newer decision model. The examples directory holds exactly two files, one for chat and one for a tool call. That is the whole of the code in the repository, and the reason the project reports no detectable primary language. Everything you actually run lives elsewhere: the weights are on a model hub and on a second model hub, and the page links both for every checkpoint. The PDF being committed rather than only cited is a small choice worth noticing, since it means the evaluation method travels with the repository instead of behind a preprint link.

## The headline average is computed over a subset the base model is not in

The average column at the right of the comparison table is described precisely: it is the equal-weight mean of the six benchmark groups, and only models with results in all six groups are ranked by it. Anything missing is shown as a dash and goes unranked. Six models appear in the table. Four have complete results and therefore an average, and NeoHorse-Jev is the highest of those four at 77.70. The fifth row, the project's own base model, has dashes in three of the six columns and no average at all, so it is not part of the ranking the release is announced on. That is a defensible rule and the page states it outright rather than hiding the missing cells. It does mean the summary sentence and the base model are talking about different populations, which is worth keeping straight when you read them.

## The gain over the base model is measured only where the base model has numbers

The prose claim and the table agree, and both are narrower than they first look. The stated improvement is eleven and a half percentage points over the base model, and it is scoped to three benchmarks. Those three are exactly the ones where the base model has results. So the comparison is: on the subset the base model can be measured on, the decision model averages 83.26 percent against a base average in the low seventies. The three benchmarks where the base model has nothing are the three where the decision model posts its second-best and best-looking figures, including the highest single score anywhere in the table. Nobody can read a gain on those from this table, because the other side of it is blank.

## The six environments in the figure are not the six benchmarks in the table

Under the decision-model comparison there is a caption naming six environments, in order, left to right and top to bottom: two arcade-style games, a robot manipulation task, a tile game, a four-player bomb arena, and an autonomous driving task. None of those six appears in the benchmark table beside it. The table's six groups are a text suite, two named decision benchmarks, a text variant of one of them, and three data-oriented suites. So the repository contains two different evaluation stories, one measured on text benchmarks and one illustrated with interactive or embodied environments, and the page does not connect them. If you are deciding whether the decision model can drive your application, the environments are the more relevant picture and the table is the more comparable one, and you only get numbers for the second.

## Three generations of lineage, from one upstream base to a classifier

The download table has a base-model column, and reading it downward gives the lineage. The decision model is built on the smaller general model of the same size. The general models are post-trained from an upstream 4B-class family of causal language models. So there are three levels: an upstream base, a post-trained pair at two sizes, and a decision model derived from the smaller of the pair. Two size choices are offered rather than one, described as a lighter local footprint and a higher-capacity option on the same text-first serving interface. Both general sizes are also published as a low-bit GGUF package in several precisions, from full 16-bit down through three quantised tiers, plus a separate build for Apple silicon, which is the practical answer to whether you can run this on a laptop at all.

## The recursive improvement loop has been run once, and iterating it is future work

The page is careful about the phrase in the title. The general models are called an initial prototype on the path toward recursive self-improvement, and the last sentence of that section says extending the loop across successive iterations is the next step. Read the mechanism and the claim lines up. A routing harness assigns tasks to a pool of different models, records the tool interactions and their outcomes, estimates what capability the task demanded, and feeds that capability-level feedback into the next training mixture. An updated model can then go back into the harness, which closes one evaluation, selection and update cycle. The training side is described as two methods, a routing-guided curriculum pass over supervised examples and a routing-guided on-policy distillation step, both of which turn execution trajectories into training signal while preserving the execution and harness context. One cycle is a prototype. A loop that improves itself is not yet shown.

## The decision model is a classifier wearing a language model's clothes

The decision model does not generate an answer and parse it. It uses prefill-only inference, which means the questions and candidate answers are supplied in the prompt and nothing is generated back, and the output is read directly. Three named primitives cover the cases: one selects an action, one checks a condition, and one assigns a rating. Applications define both the questions and the possible answers, and use the results to route requests, pick tools, or drive a workflow. That design is why it can be scored on the text decision benchmarks at all: a free-form generation model would have to be parsed and graded, while this returns a distribution over options the application already supplied. The naming is unconventional, particularly the word used for the condition check, and the page does not explain why the three were named that way.

## Conclusion

NeoHorse is worth evaluating if you want a small open model for tool use and you are willing to treat the benchmark table as a starting point rather than a settled result. Three things to check before you commit. The average that carries the release is computed only over models with complete results, and the project's own base model is excluded from it, so read the per-benchmark columns rather than the summary. The improvement claim over that base is measured on the three benchmarks where the base has numbers at all. And the recursive improvement loop is described by its authors as a prototype run once, with iterating it named as the next step rather than a result.

## FAQ

### What is in the NeoHorse repository?

A licence, the README, a technical report as a PDF, a figures directory, an examples directory holding two scripts for chat and for a tool call, and a subdirectory for the decision model. The model weights themselves are hosted on a model hub and a second model hub rather than in the repository.

### What are the NeoHorse model sizes and where do they come from?

A 4B and a 9B pair post-trained from an upstream causal language model, plus a 4B decision model built on the smaller general model. The download table gives the base model for each checkpoint, so the lineage runs from the upstream base through the general pair to the decision model.

### How is the NeoHorse average benchmark score calculated?

It is the equal-weight mean of six benchmark groups, and only models with results in all six are ranked. The page states that missing results are shown as dashes and go unranked, which is why the project's own base model has no average despite scoring on three of the groups.

### What does NeoHorse-Jev use prefill-only inference for?

It turns application state into decisions and probabilities without generating text. Three primitives cover it: one selects an action, one checks a condition, and one assigns a rating. Applications supply the questions and the possible answers, then route requests or pick tools from the results.

### Can I run NeoHorse locally, and in what formats?

The weights are published as low-bit packages at 16-bit and at 8-bit, 5-bit and 4-bit, and there is a separate build for Apple silicon. Distribution runs through two model hubs, and the page says the smaller quantised versions exist to use less disk and memory so the models can run on your own hardware.

## Sources

- [Issues](https://github.com/TokenRhythm/NeoHorse/issues)
- [License: Apache-2.0](https://github.com/TokenRhythm/NeoHorse/blob/main/LICENSE)
- [README](https://github.com/TokenRhythm/NeoHorse/blob/main/README.md)
- [TokenRhythm/NeoHorse on GitHub](https://github.com/TokenRhythm/NeoHorse)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tokenrhythm-neohorse
