# fablize ships four procedures and names the four ideas it refused to ship

> A Claude Code plugin built from a controlled comparison of two models, which concluded that working procedure transfers and capability does not. The interesting part is not the four things it ships, it is the four ideas it deliberately left out and the deterministic hook that blocks a polite offer.

**fivetaku/fablize** — A Claude Code plugin that makes Opus behave like Fable — completion, evidence, and verification enforced as procedure. Ships only what a Fable-vs-Opus comparison proved transferable.

- Repository: https://github.com/fivetaku/fablize
- Stars: 894 · Forks: 127
- Language: Python
- License: MIT
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/fivetaku-fablize

## The comparison behind the plugin, with the sample size attached

The plugin exists because of a specific experiment, and the experiment is described with its size attached rather than as a general impression.

The author ran a controlled comparison of two models when the second one shipped: an A/B set of nineteen runs, plus twenty-six real working sessions, and roughly fifteen hundred tool calls in total.

The results split cleanly. On closed, answer-bearing work, meaning code, logic and builds, the two were effectively tied. The gap appeared only on open-ended work, and its character was narrow: following an implication one step further.

That last sentence is the whole thesis of the project. Depth of that kind is treated as model capability, and the claim is that it cannot be handed over by instructions or by a harness around the model.

The verification for that claim is a specific experiment rather than an argument. An injection was tried, and it failed: the second model could not reproduce the defects the first one found on its own. So the plugin is built on the residue, the part that did transfer, which is described as the procedure of good work: actually running what you build, seeing it through, and investigating systematically.

## Four ideas are named and excluded, which is rarer than shipping them

Most plugin write-ups list what is included. This one also lists what was tried and left out, and the list is specific: style mimicry, broad reasoning injection, a silent recovery guard, and a review recall scan.

They are described as negligible or unverified, and the stated reason for their absence is that they stay in personal development until a controlled experiment confirms their effect.

That is a stricter standard than most projects apply to their own successful features, and it is worth noticing that the same person who could not ship four plausible ideas also shipped four others. The distinguishing question is not whether an idea sounds good, it is whether running it changed the behaviour being targeted.

The shipped set is the residue of the first experiment, expressed as a per-task router that injects only the matching verified discipline. Verification grounding means renderable or executable output, such as HTML, SVG, games and charts, is actually run and observed before the task is called done. The multi-story gate decomposes work and refuses a groundless completion. The investigation protocol is reproduce, compete hypotheses, trace the causal chain. The early-stop hook is deterministic rather than advisory.

## The early-stop hook blocks a polite offer, and the fix is your typing

Three of the four shipped behaviours are procedures the model would follow if reminded. The fourth is different, and it is the one with a documented failure mode.

The early-stop hook is described as deterministic, and its job is to catch a promise without the work behind it: an announcement of an action that never happens. Unlike the other three, it does not instruct the model to do something, it refuses to let a turn end on a promise.

The cost is stated plainly. It can misfire on a declarative offer, and the example given is an offer to write a report if the user wants one. A courteous sentence of the form I will do X if you would like is, structurally, the same shape as I will do X. There is no way for a deterministic check to tell the difference between a promise of work and a promise of availability.

The suggested workaround is to phrase offers as questions. That moves the fix onto the user, and it is worth being clear that this is a permanent behavioural requirement rather than a version that will improve: the hook has to be right every time, and politeness is the case that loses.

The other three procedures fail differently, which is the better failure. A verification step that is skipped means a build that was never run, which you find out about at the next command.

## goals.py is the gate, and the router decides which discipline loads

The multi-story mechanism is a file, not a prompt. `goals.py` decomposes the work and refuses a completion that has no proof attached.

That is the load-bearing part of the design, because a decomposition with a checkpoint and an evidence requirement is a procedure the model has to follow across many turns, which is exactly the case where a one-line instruction decays.

The routing is described as per task rather than always on. Two or more stories means decompose plus the verification gate. Debugging means the investigation protocol. A renderable artifact means verification grounding. A hard task means adaptive thinking plus a suggestion to raise the effort level.

Two details are worth noting. The router registers itself as a prompt submission hook, so it runs on the user's turn rather than being invoked by the model, which means the discipline is applied before the model has decided what kind of task this is. And the effort suggestion is a suggestion, phrased as one, which keeps it inside the plugin's stated boundary of not deciding things for you.

The last route is escalation. At the ceiling, the plugin tells you to go to a stronger model or a person, which is the honest answer given the transfer table, and is also the one a plugin with an incentive to be used would have an easy time omitting.

## Always-on mode is a second install, with a scope choice

Installation has two halves, and only the first is a plugin install.

```
/plugin marketplace add fivetaku/fablize
```

```
/plugin install fablize
```

After that the per-task router registers automatically. That is the whole of the first install: the hook is in place and it will match tasks as they are submitted.

Always-on operation is separate. The rules have to be resident in the context rather than injected per task, and that is done by running a setup script once, which offers a choice between local and global scope, with local recommended:

```
bash ${CLAUDE_PLUGIN_ROOT}/setup/setup.sh
```

Uninstall is its own script:

```
bash ${CLAUDE_PLUGIN_ROOT}/setup/uninstall.sh
```

So there are three states a user can be in, and only the middle one is the default: per-task routing, always-on with rules in context, or not installed. The second is the one that changes how the model sees every request, and it is the one a user is most likely to adopt and least likely to read about.

The two scripts are the only shell entry points named in the documentation, and they sit under a plugin root variable, which means they run from wherever the plugin was installed rather than from a path in the repository.

## The setup script asks once whether you want to star the repository

One instruction in the documentation is about the repository rather than the plugin, and it is worth reading because of what kind of software is asking.

The setup script asks whether you want to give the project a star, and the documentation specifies the terms: it asks once, as a single opt-in question, and it never stars without an explicit yes.

That specificity is the point. A plugin that installs a hook on every prompt submission and can put rules in your context is not a good place for a repeated request, and the documentation is aware of that, so it states the constraint rather than leaving it to the behaviour of the script.

The trigger vocabulary is worth knowing too, because it determines when the plugin activates without you asking. The documented phrases are the slash command, and a phrase like see it through, and in always-on mode it matches automatically.

The honest limits section, which is the most useful part of the file, also covers the one that is hardest to fix. The effect numbers come from a small, single-family self-measurement on the nineteen-run set, and the direction of the result is described as solid while the decimals are not asserted. Note the gap between that and the earlier description of the comparison, which included twenty-six real working sessions; the numbers are attributed to the smaller set.

## One release, and the default branch is weeks ahead of it

The release history is a single entry, version 2.1.0, published on 18 June 2026. The last commit to the repository is 6 July 2026, so the head of the default branch contains roughly two and a half weeks of work that no tag covers.

With one release there is no upgrade path to describe, and no way to tell from the tag list whether the difference is documentation, a new route, or a changed hook. The changelog at the root is where that would be recorded.

The tree is broader than the four commands in the documentation suggest. Alongside the plugin manifest directory, the licence, the changelog and two versions of the readme, there are separate directories for commands, documentation, hooks, packs, scripts, setup, skills and tests. The language is reported as Python, and the two named Python artefacts in the documentation are the completion gate and the setup scripts around it, so the Python sits in the script and test directories rather than in a package of its own.

The second readme is a Korean translation linked from the first line of the English one, so the project is maintained in two languages with no stated rule about which is authoritative when they disagree.

## Conclusion

Reach for fablize if your work with Claude Code fails in a way that is about follow-through rather than intelligence: the model plans, promises, stops early, or claims a change works without running it. That is the failure mode it was built for, and the evidence for it, while small and from one model family, is at least published with the number of runs attached. Three things to know before installing. That it is a prompt hook, so it runs on your submissions and its rules sit in your context in the always-on mode, which is a bigger footprint than a settings change. That the early-stop gate can misfire on a courteous offer, and the documented workaround is to phrase offers as questions, which is a habit you have to maintain rather than a bug that gets fixed. And that it will not make a weaker model reason better; the README says the ceiling is a model-choice decision, and the plugin's own instruction at that point is to escalate.

## FAQ

### What is the fablize Claude Code plugin?

A plugin that injects verified working procedures into Claude Code, selected per task by a router hook. The shipped procedures are verification grounding, where renderable or executable output is run and observed before completion, a multi-story evidence gate in goals.py, an investigation protocol, and a deterministic early-stop hook.

### How do I install and remove fablize?

Add the marketplace and install the plugin with two slash commands, after which the per-task router registers as a prompt submission hook. Always-on operation, with the rules resident in context, is a separate one-time setup script that offers local or global scope. Uninstall has its own script under the plugin root.

### Does fablize make a weaker model reason as well as a stronger one?

No, and the documentation says so directly. An injection experiment refuted capability transfer: the model either finds an out-of-spec defect or it does not. Three capabilities are marked not possible, and the plugin's instruction at the ceiling is to escalate to a stronger model or a person.

### What does fablize deliberately not ship?

Four ideas are named as excluded because they are negligible or unverified: style mimicry, broad reasoning injection, a silent recovery guard, and a review recall scan. They stay in personal development until a controlled experiment confirms their effect.

### How strong is the evidence behind fablize?

It comes from a controlled comparison of two Claude models: an A/B set of 19 runs plus 26 real working sessions and about 1,500 tool calls. The effect numbers are attributed to the 19-run set, and the documentation states that the direction of the result is solid while the decimals are not asserted.

## Sources

- [fivetaku/fablize on GitHub](https://github.com/fivetaku/fablize)
- [Issues](https://github.com/fivetaku/fablize/issues)
- [License: MIT](https://github.com/fivetaku/fablize/blob/main/LICENSE)
- [README](https://github.com/fivetaku/fablize/blob/main/README.md)
- [Releases](https://github.com/fivetaku/fablize/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fivetaku-fablize
