Model or dataset
trotsky1997/OpenFugu avatar
trotsky1997/OpenFugu

OpenFugu rebuilds Sakana's Fugu as a learned router, and grades its own evidence

Open reimplementation of Sakana Fugu — the 'one model to command them all' LLM orchestrator. Read → run → train → serve.

462 stars83 forksPythonApache-2.0

At a glance

What is it?
OpenFugu is an independent Python reconstruction of the mechanism behind Sakana AI's closed Fugu orchestrator: a small coordinator that routes each query to a pool of other people's models and returns one answer. Its most distinctive feature is not the routing but the bookkeeping, since every claim in the documentation carries an EXEC, CODE or DATA grade, and the headline routing number arrives with its own caveat attached.
Who is it for?
Adopt OpenFugu if you are studying learned routing as a technique, if you want a runnable reference for a linear-head router trained without gradients, or if you need the evidence grading as a model for your own reverse-engineering notes.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 110 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Fugu is described as a policy over models rather than a model

At the centre of the project is a reframing. Fugu is sold as a single model, and the project's position is that it is really a policy over models: a small coordinator that, per query, routes work to a pool of frontier LLMs and returns one answer. The user sees one model. Underneath there is a selection step and someone else's weights doing the actual work.

Two consequences follow, and they are the reason this kind of system is worth building. No worker weights are ever touched, so the whole thing is macro-level composition over other people's models rather than a new model. And because the coordinator is small, the thing you train is the routing decision, which is a far cheaper object than a model and can be trained without backpropagation through the workers.

Serving makes the arrangement concrete. `openfugu/serve.py` exposes one OpenAI-compatible `/v1/chat/completions` endpoint with an internal TRINITY loop running over a litellm pool, so the pool is hidden behind a standard interface. litellm is what makes the pool wide: any endpoint it can talk to can serve as a worker, which is why one dependency in the requirements file does so much of the work.

What you get is a router wearing an API. That framing matters for expectations, because a caller cannot tell from the response which model answered, and a system built this way inherits every property of its weakest worker.

Routing in your own code is the alternative. If your workload has a recognisable shape, a rule that sends billing questions to one model and code questions to another is cheaper, inspectable, and fails predictably. What a learned router buys is the absence of those rules: the policy is fitted from behaviour instead of written by hand, so it can change when the workload does. What it costs is that the policy is now a set of numbers you have to trust, evaluate and retrain, and that the failure mode becomes a bad route rather than a bug you can read. Its own evaluation is framed against exactly that alternative, comparing a trained router to the best single worker rather than to a hand-written dispatcher.

TRINITY scores workers from one hidden state with a bias-free head

Small enough to reconstruct from a description, the router is specific about what it does. A backbone of about 0.6B parameters, Qwen3-0.6B, never answers the user. It produces a single hidden state at the penultimate token, and a bias-free linear head scores each candidate worker from that state. Dispatch follows, and what comes back is the top-scoring worker's reply.

Tiny by design, the trainable surface is roughly 19.5 thousand numbers, comprising the linear head plus singular-value-fine-tuning offsets on 9 matrices. That is small enough to optimize with a gradient-free method, and the project uses sep-CMA-ES, a separable covariance matrix adaptation scheme, so training needs no gradients through the 0.6B backbone at all. That is the part of the design that would be hardest to arrive at independently, and it is what makes the whole thing runnable on ordinary hardware.

`train/train_trinity.py` trains that coordinator from scratch and does so without Sakana weights, which is the independence claim stated as a code path rather than as a promise. Two scripts under `train/` cover the two architectures: this one for TRINITY, and `train_conductor.py`, which uses GRPO to fit the Conductor on `nvidia/ToolScale`.

Verification is a single command against real weights:

bash
python openfugu/mini.py --self-test       # -> 95% / 100%

Two numbers come back, 95 percent on agent selection and 100 percent on role, over a 37-case fixture. Those are the numbers to check first, because they are the only headline figures in the project produced against actual weights rather than a stand-in.

Ultra names a 7B Conductor, and the published checkpoint is 3B

The larger variant, Fugu-Ultra, replaces the per-turn picker with a Conductor that emits a whole workflow DAG, so one query can fan out across several workers in a planned order rather than choosing a single target. In the mechanism description, that Conductor is 7B.

The weights section describes something smaller in size. Trained on ToolScale, a fine-tune of Llama-3.2-3B-Instruct, the Conductor is published on HuggingFace as `huggingface.co/di-zhang-fdu/openfugu-conductor-3b`. So the architecture narrative points at a 7B model and the artefact you can actually download is a 3B one. That may be a deliberate simplification for a runnable release rather than an error, and the model card is named as the place to check, but the visible text does not reconcile the two numbers.

Where those weights sit matters as much as their size. They are published on HuggingFace, not in this repository, because the Llama 3.2 Community License applies, and the licence section says so directly: Apache-2.0 covers the OpenFugu code, third-party material is fetched rather than redistributed, and trained weights carry the Llama 3.2 licence.

So the project's most interesting trained artefact is the one artefact you cannot get from the repository, under a licence that is not the repository's.

The +107% routing number carries its own caveat

`eval/eval_orchestration.py` asks a narrow question: does per-question routing beat the best single model? That is the right question to ask of a router, and it is the question the project puts to itself rather than asking whether the whole system is good.

The answer it reports is a trained router at plus 107 percent over the best single worker. The same sentence carries the qualifier that matters most: this is query-level routing, not per-step coordination, and the project links to a results caveat. Both halves belong in the same breath. A router that picks a different model per question and a coordinator that plans a multi-step DAG are different claims with different evidence behind them, and only the first is being asserted here.

Evidence is graded per claim, which is the habit worth copying whether or not you use the code. Every statement in the documentation carries one of three grades, EXEC for something executed, CODE for something read in code, and DATA for something measured. A reader can tell at a glance which sentences rest on a run and which rest on an inference.

`docs/HOW_FUGU_IS_IMPLEMENTED.md` holds the full mathematics and `docs/ARCHITECTURE.md` is the investigation log, also graded. Criteria for each grade are not spelled out, so a grade is only as useful as the discipline behind it, which means reading the documents rather than trusting the label.

Two of the three headline numbers come from mocks

Sorting the reported results by what they were produced against is the fastest way to understand this project. Described as running on real weights, the self-test figure of 95 percent agent and 100 percent role comes from the 37-case fixture. That is a measurement against the artefact.

Not so the TRINITY training result. Chance to optimal routing in about 5 generations, and the parenthetical says mock, runs anywhere. That is a demonstration that the optimizer and the head can learn the routing function at all, on synthetic data, without a GPU. It is a real result about the training loop and no result at all about the checkpoint.

Between the two sits the Conductor figure: reward rising from 1.21 to 1.64 over 100 steps, on `nvidia/ToolScale`, with a curve in `results/`. A hundred steps is short, and reward scale is not comparable across tasks, so the number establishes that GRPO training moves the objective rather than that the Conductor is good.

None of this is hidden, and that is to the project's credit, but a reader who takes the three numbers at face value will overstate the result. Reproduce the self-test first.

Artifacts are fetched by a script, and the quickstart stops mid-command

Nothing third-party is vendored. A fetch script pulls the backbone, a vector and the fixture from their licensed sources, and the environment is then pointed at wherever they landed:

bash
export FUGU_MODEL=$(...Qwen3-0.6B path...)
export FUGU_VECTOR=$PWD/artifacts/model_iter_60.npy
export FUGU_FIXTURE=$PWD/artifacts/qwen_router_prompt_eval_cases.json

Three variables, and the middle one carries the most uncertainty. Naming matters here: `model_iter_60.npy` reads as a numbered iteration of something released, and neither the visible text nor the fetch script's description says what iteration 60 corresponds to or which release it came from. A reconstruction that depends on a specific released vector depends on that vector continuing to be available.

After that the quickstart reads the documentation, runs the self-test, and stops. A comment fragment is the last line in the visible block, `# RUN: route one query (of`, cut off mid-sentence, so the command that routes a single query is not shown. Installing is two commands:

bash
pip install -r requirements.txt           # torch, transformers, trl, litellm, ...
python scripts/fetch_artifacts.py         # pull Qwen3-0.6B + model_iter_60.npy + fixture (not redistributed)

and then the documentation is the guide. `docs/handoff.md` is named as the place for version notes, which is the file to read first.

Apache-2.0 covers the code, the weights carry Llama's licence

Licensing is split three ways and stated plainly. Apache-2.0 covers all OpenFugu code, under `LICENSE`. Third-party material is fetched rather than redistributed, documented in `NOTICE`. Trained weights carry the Llama 3.2 licence, which is why the Conductor lives on HuggingFace rather than in the tree.

There is also the affiliation question, answered in a blockquote at the top: this is an independent reimplementation, not affiliated with Sakana AI, and no third-party code or weights are redistributed. Sakana's product and trained weights are closed, so nothing here substitutes for the commercial thing, and the reconstruction is built from the two papers plus released artefacts.

Distribution has no release history. There are no GitHub releases and no project homepage, so version tracking means commit hashes, and the last push was on 2026-06-22. A research reconstruction with a moving default branch and no tags is normal, and it is still a factor in whether you can cite a specific state of the code.

Also at the repository root: `.claude/`, which suggests agent instructions, and four directories whose purpose the visible text does not explain: `verify/`, `pipeline/`, `openspec/` and `assets/`. Two links point at `results/` as a bare directory, which is where the training curves and the routing caveat live.

The requirements file calls its own dependency pinning version hell

Better than the feature list, the dependency list explains the shape of the project, because every entry maps to a mechanism:

text
torch>=2.4
transformers>=4.52,<5
trl>=0.19,<0.20
datasets>=3.6
peft
accelerate
numpy
litellm
hydra-core>=1.3
omegaconf
math_verify
huggingface_hub
cma

`cma` is the optimizer behind sep-CMA-ES, `math_verify` is the reward, `hydra-core` and `omegaconf` are the training configuration system, `peft` and `accelerate` are there for the Conductor fine-tune, and `litellm` is the worker pool. Reading the file top to bottom tells you that running the server and training a Conductor are one installation.

Most interesting are the ceilings, and the file's own comment above them names the problem: the core pins are set to what the GRPO stack needs, with a pointer to `docs/handoff.md` for version notes. `trl` is capped below 0.20 and `transformers` below 5, so a transformers 5 release or a trl 0.20 will not install cleanly against this file until someone moves it.

That is the honest cost of the project. Reproducing it means a GPU-adjacent dependency stack with narrow pins, not a single library, and the training side pulls in more than the serving side needs. If you only want the endpoint, the litellm route is a much smaller dependency than the one the file requires.

Editorial conclusion

Adopt OpenFugu if you are studying learned routing as a technique, if you want a runnable reference for a linear-head router trained without gradients, or if you need the evidence grading as a model for your own reverse-engineering notes. Do not adopt it expecting Sakana's product or its weights, since neither is here and the project states it is not affiliated, and do not adopt it as a router for production traffic on the strength of the routing number, which the project itself scopes to query-level routing rather than per-step coordination. Verify three things first: that `model_iter_60.npy`, fetched rather than vendored, is the artefact your reconstruction needs, since the visible text does not say what iteration 60 corresponds to; which of the three headline numbers you are relying on, because the self-test uses real weights while the training curve does not; and the licence of the Conductor checkpoint, which is the one artefact you would want and the one that is not Apache-2.0.

Frequently asked questions

What is OpenFugu and is it affiliated with Sakana AI?

OpenFugu is an independent reimplementation of the mechanism behind Sakana AI's Fugu, which the project describes as a policy over models rather than a model: a small coordinator that routes each query to a pool of frontier LLMs and returns one answer. It states plainly that it is not affiliated with Sakana AI, and that Sakana's product and trained weights are closed.

How does the TRINITY router in OpenFugu choose a model?

A Qwen3-0.6B backbone never answers the user. It produces one hidden state at the penultimate token, a bias-free linear head scores each worker from that state, and the top-scoring worker is dispatched and its reply returned. The trainable surface is about 19.5 thousand numbers, optimized gradient-free with sep-CMA-ES.

Can I get OpenFugu's trained Conductor weights from the repository?

No. They are published on HuggingFace as huggingface.co/di-zhang-fdu/openfugu-conductor-3b rather than in the repository, because the Llama 3.2 Community License applies. The repository is Apache-2.0 for its own code, third-party material is fetched rather than redistributed, and trained weights carry the Llama 3.2 licence.

What does the +107% figure in OpenFugu measure?

It compares a trained router against the best single worker, and the project scopes it to query-level routing rather than per-step coordination, linking to a results caveat. The evaluation question it poses is whether per-question routing beats the best single model, which is a narrower claim than the Conductor emitting a whole workflow DAG.

How do I install and run OpenFugu?

Install with pip install -r requirements.txt, then run python scripts/fetch_artifacts.py to pull Qwen3-0.6B, model_iter_60.npy and the fixture, since nothing third-party is redistributed. Point FUGU_MODEL, FUGU_VECTOR and FUGU_FIXTURE at the results, then run python openfugu/mini.py --self-test. The quickstart in the README is cut off before the command that routes a single query.

What is the difference between TRINITY and Fugu-Ultra in OpenFugu?

TRINITY, in openfugu/mini.py, picks one worker per query using a linear head over a hidden state. Fugu-Ultra, in openfugu/ultra.py, swaps the per-turn picker for a Conductor that emits a whole workflow DAG. The mechanism text calls that Conductor 7B, while the published checkpoint is a 3B fine-tune of Llama-3.2-3B-Instruct.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. trotsky1997/OpenFugu on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/trotsky1997-openfugu.svg)](https://hysenlabs.com/projects/trotsky1997-openfugu)