Model or dataset
TokenRhythm/NeoHorse avatar
TokenRhythm/NeoHorse

NeoHorse: Open-Weight Agent Models with a Routing-Based Path to Recursive Self-Improvement

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.

1,396 stars17 forksUnknownApache-2.0

At a glance

What is it?
NeoHorse is TokenRhythm's family of open-weight language models built for agent workflows, with a routing harness that feeds execution outcomes back into training. The 4B and 9B checkpoints are production-ready today; NeoHorse-Jev adds prefill-only decision inference on top.
Who is it for?
NeoHorse-1 is a practical choice for engineers who want an Apache 2.0 agent model at the 4B or 9B scale, with GGUF and MLX variants that reduce hardware requirements. NeoHorse-Jev adds structured decision output for routing or workflow control without requiring a separate classification model.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What NeoHorse Is and Who It Targets

NeoHorse is a family of open-weight language models from TokenRhythm, designed specifically for text-based agent harnesses. The repository contains two related model lines. NeoHorse-1 provides 4B and 9B causal language models post-trained from Qwen3.5 for tool use, coding, and instruction following. NeoHorse-Jev, released on 2026-09-24, is a decision model built on the NeoHorse-1-4B checkpoint that uses prefill-only inference to output structured decisions: a selected action (Choice), a boolean condition check (Noul), or a numerical rating (Score).

The target users are engineers building agent pipelines who want an open-weight model they can run locally, fine-tune, or embed in a harness without licensing restrictions. The Apache 2.0 license covers both model weights and the repository code. GGUF quantized variants (8-bit, 5-bit, and 4-bit) and MLX versions for Apple silicon are available on Hugging Face, which reduces the memory requirements enough to run on consumer hardware.

Routing Harness and the Recursive Self-Improvement Loop

The core design concept in NeoHorse-1 is what the README calls agentic post-training. A routing harness assigns tasks to a pool of models, records how each model handled tool calls and what outcomes it produced, estimates the capability demand of each task, and feeds capability-level feedback back into the next training mixture. Updated models can return to the harness, forming an evaluation-selection-update loop.

The README describes this as a prototype on the path toward recursive self-improvement (RSI): a system that uses its own execution history to improve itself. The current release is that prototype, not an autonomous RSI system. The next step described in the README is extending this loop across successive iterations.

Two training techniques support this design. Routing-guided curriculum SFT selects training data based on the harness's capability estimates. Routing-guided on-policy distillation uses execution trajectories as training signal while preserving execution and harness context. The README also mentions data quality steps including exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling.

NeoHorse-Jev: Prefill-Only Decision Output

NeoHorse-Jev is a separate model that builds on NeoHorse-1-4B to handle decisions rather than open-ended text generation. It operates through prefill-only inference, meaning the model scores possible answers rather than generating tokens autoregressively. Applications define a question and a set of possible answers; Jev returns a Choice (one selected answer), a Noul (a boolean), or a Score (a numeric rating).

The README reports evaluation results across six text benchmark groups. NeoHorse-Jev-4B scores 77.70 average across JevBench, Kev, OpenJev text, Nimble, VitaminC, and MASSIVE. Across the three benchmarks Nimble, VitaminC, and MASSIVE, it reaches 83.26 percent mean accuracy, which is 11.50 percentage points above NeoHorse-1-4B. The comparison table includes Open-Jev-9B, Kev-4B, and two Laya variants. Models without complete results across all six benchmark groups are not ranked by average.

The design trade-off is focus: Jev is purpose-built for classification and scoring tasks in agent routing and workflow control. It does not replace a general-purpose language model in the same pipeline.

Downloading and Running NeoHorse Models

The checkpoint variants and their distribution points are listed in the README. NeoHorse-1-4B and NeoHorse-1-9B are available on Hugging Face under the TokenRhythm organization. GGUF variants (NeoHorse-1-4B-GGUF and NeoHorse-1-9B-GGUF) include 16-bit BF16 weights and smaller 8-bit, 5-bit, and 4-bit quantized versions. Both models are also available on ModelScope for users in regions where Hugging Face is less accessible. MLX versions for Apple silicon are in the TokenRhythm MLX collection on Hugging Face.

The repository itself contains the technical report PDF at TechnicalReport_NeoHorse_v1.pdf, example scripts under examples/ (including examples/chat.py and examples/tool_call.py), and the Jev subdirectory with its own README covering deployment and evaluation. The README does not include install commands; the primary path to the models is through Hugging Face or ModelScope download pages linked in the README.

The 8-bit, 5-bit, and 4-bit GGUF quantizations reduce disk and memory footprint; the README does not specify what precision loss to expect at each quantization level. The repository has no GitHub releases. The last push was on 2026-09-26.

Limitations and Cases Where NeoHorse Is the Wrong Choice

NeoHorse-1 tops out at 9B parameters. For tasks that require a larger context window, higher reasoning capacity, or better performance on complex code generation, a 70B or larger model is more appropriate. The README does not document the training data composition in detail beyond mentioning decontamination and labeling steps, so reproducibility of the training pipeline is limited by what the technical report discloses.

NeoHorse-Jev uses prefill-only inference, which means it is only useful when the set of valid answers can be enumerated in advance. It cannot generate free text. Pipelines that need both decision logic and generative output must combine Jev with a separate generative model.

The routing harness that feeds into training is described as a prototype. The README does not document a ready-made harness that teams can deploy; it describes the design and training outcome. Engineers who want to replicate the self-improvement loop will need to build the harness themselves based on the technical report.

The RELATED SEARCHES for this project return results about a Hungarian equestrian supply store rather than the model, which indicates the project name overlaps with an unrelated business.

Comparison with Other Open-Weight Agent Models

The benchmark table in the Jev README positions NeoHorse-Jev-4B against Open-Jev-9B, Kev-4B, and Laya English. Open-Jev-9B is more than twice the parameter count and still scores lower on the average of all six benchmarks, according to the README's table. Kev-4B is also 4B scale and scores 74.25 average.

For the base NeoHorse-1 models (4B and 9B), the closest general category of alternative is Qwen3.5, which is the base model NeoHorse-1 is post-trained from. The NeoHorse post-training adds agent-specific capability: tool use, coding tasks, and the routing harness curriculum. Teams that only need a general instruction-following model may find the base Qwen3.5 adequate without the additional agent post-training.

Maintenance Status and License

The last push to the repository was on 2026-09-26, two days before the date of this article, which indicates active development. The project follows rapid release cadence: NeoHorse-1 was released on 2026-09-07 and NeoHorse-Jev followed on 2026-09-24.

All weights and code are released under the Apache 2.0 license, which permits commercial use, modification, distribution, and sublicensing. Attribution is required. The README does not mention any additional terms restricting model use beyond the Apache 2.0 text.

Editorial conclusion

NeoHorse-1 is a practical choice for engineers who want an Apache 2.0 agent model at the 4B or 9B scale, with GGUF and MLX variants that reduce hardware requirements. NeoHorse-Jev adds structured decision output for routing or workflow control without requiring a separate classification model. Teams that need a larger model, a model with a public API, or a model with well-documented RLHF training data should evaluate other options. The technical report at arxiv.org/abs/2609.08183 covers the routing harness design, training methodology, and benchmark protocols in detail.

Frequently asked questions

What is NeoHorse and what makes it different from other open-weight models?

NeoHorse is a family of open-weight agent models from TokenRhythm, post-trained from Qwen3.5 for tool use, coding, and agent workflows. Its distinguishing design is a routing harness that records execution outcomes and feeds capability-level feedback back into training, forming a prototype evaluation-selection-update loop toward recursive self-improvement.

What is NeoHorse-Jev and how does it differ from NeoHorse-1?

NeoHorse-Jev is a decision model built on NeoHorse-1-4B that uses prefill-only inference to return structured outputs: a chosen action (Choice), a boolean result (Noul), or a numeric rating (Score). Unlike NeoHorse-1, it does not generate free text; it is designed for routing and workflow control where the set of valid answers is defined in advance.

What quantization formats are available for NeoHorse-1?

NeoHorse-1-4B and NeoHorse-1-9B are available in GGUF format with 16-bit (BF16), 8-bit, 5-bit, and 4-bit quantized variants on Hugging Face. MLX versions for Apple silicon are also available in the TokenRhythm collection.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. TokenRhythm/NeoHorse on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tokenrhythm-neohorse.svg)](https://hysenlabs.com/projects/tokenrhythm-neohorse)