Model or dataset
p-e-w/heretic avatar
p-e-w/heretic

Heretic: automatic abliteration for language models, without post-training

Fully automatic censorship removal for language models

32,525 stars3,655 forksPythonAGPL-3.0

At a glance

What is it?
Heretic is a Python tool that removes refusal behaviour from transformer models by combining directional ablation with a TPE parameter search. It runs unattended, but it is not a drop-in fix for every architecture.
Who is it for?
Adopt Heretic if you want a decensored version of a dense or MoE transformer and are willing to let an Optuna search run until it converges. Do not adopt it for pure state-space models or research architectures, which the README says are not supported out of the box, and do not treat the refusal and KL numbers as a substitute for human evaluation.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Heretic actually removes, and who needs it

Heretic targets refusal behaviour in transformer-based language models. The README calls this censorship, or safety alignment, and the tool's job is to strip it without expensive post-training. The intended user is someone who has a model checkpoint and wants a version that answers prompts the original would reject, but who does not want to lose general capability in the process. The README is explicit that using Heretic does not require an understanding of transformer internals, and that anyone who can run a command-line program can use it.

That framing matters because the alternative is manual abliteration, which the README describes as work normally done by human experts. Heretic's claim is that an unsupervised run with the default configuration can rival those hand-tuned results. The project's own comparison table puts p-e-w/gemma-3-12b-it-heretic at 3 refusals out of 100 harmful prompts and a KL divergence of 0.16, against 3/100 and 1.04 for mlabonne/gemma-3-12b-it-abliterated-v2 and 3/100 and 0.45 for huihui-ai/gemma-3-12b-it-abliterated. The README notes those numbers were compiled with PyTorch 2.8 on an RTX 5090 and may be platform and hardware dependent. Treat them as the project's own measurements, not as an independent result.

Directional ablation plus an Optuna TPE search

The mechanism has two halves. The first is directional ablation, also called abliteration, following work by Arditi et al. 2024 and the projected and norm-preserving biprojected variants described by Lai. The second is a parameter optimizer based on Optuna's TPE sampler.

The optimizer is what makes the tool automatic. Rather than asking the user to pick ablation parameters, Heretic searches for parameters that co-minimize two quantities: the number of refusals on harmful prompts and the KL divergence from the original model on harmless prompts. Minimizing refusals alone would push the model far from its original behaviour. Minimizing KL alone would change nothing. The search trades the two against each other and stops at a point the optimizer considers good.

The KL term is the interesting design choice. It is a proxy for retained capability, and the README leans on it heavily, but it is still a proxy. A low KL divergence means the decensored model's output distribution stays close to the original on the harmless evaluation set. It does not guarantee that reasoning, code generation or long-context behaviour survives intact, and the README concedes that mathematical metrics and automated benchmarks never tell the whole story.

Installing Heretic and running a first abliteration

The README's install path assumes a Python 3.10+ environment with PyTorch 2.2+ already installed for your hardware. Heretic is published on PyPI as heretic-llm, so the install is a pip command followed by a run against a model identifier.

bash
pip install -U heretic-llm
heretic Qwen/Qwen3-4B-Instruct-2507

The second line is the whole workflow. Replace Qwen/Qwen3-4B-Instruct-2507 with the model you want to decensor. The process is fully automatic and needs no configuration, though the README points to config.default.toml and to heretic --help for the parameters that can be overridden.

There is a version caveat worth reading before you start. PyTorch 2.2 is the stated minimum, but the README warns that some models and configurations need later features. Loading MXFP4-quantized models such as gpt-oss uses torch.accelerator, which arrived in PyTorch 2.6.

If you already use uv, the repository ships a uv.lock pinning every dependency, and the README suggests cloning the repo and running the tool through uv instead:

bash
uv run heretic

That path keeps your dependency set aligned with the developers' and avoids resolving a different combination of packages. The README also documents an evaluation mode that reproduces the project's own comparison numbers:

bash
heretic --model google/gemma-3-12b-it --evaluate-model p-e-w/gemma-3-12b-it-heretic

According to the README, exact values may vary by platform and hardware.

Architectures Heretic will not touch

The support boundary is the first thing to check. The README says Heretic supports most dense models, including many multimodal models, several different MoE architectures, and hybrid models such as Qwen3.5. Pure state-space models and certain other research architectures are not supported out of the box. If your checkpoint falls in that second group, the tool is the wrong choice and no amount of configuration will change that.

The second limitation is the evaluation itself. Refusal counts and KL divergence are computed against fixed prompt sets. A model can score well on both and still behave badly on tasks neither set covers, which is why the README repeatedly defers to human evaluation. Anyone shipping a decensored model into a product should budget for their own task-specific testing rather than trusting the built-in numbers.

The third is dependency drift. The project pins torch and torchvision versions deliberately unspecified in pyproject.toml, so pip will resolve whatever it finds. The uv.lock file exists precisely because that resolution is not always reproducible. A run that works today can break after an unrelated upgrade.

How Heretic differs from manual abliteration workflows

Manual abliteration is the obvious alternative, and the README's own table makes the comparison concrete. Tools in that space include mlabonne/gemma-3-12b-it-abliterated-v2 and huihui-ai/gemma-3-12b-it-abliterated, both of which are published abliterated checkpoints. The difference is not the underlying technique, since Heretic also performs directional ablation. The difference is who chooses the parameters and how.

In a manual workflow, a person inspects the model, picks a direction and a scaling, runs the ablation, evaluates, and iterates by hand. That is slower and requires familiarity with the model's internals, but it lets the operator apply judgement about which behaviours matter. Heretic replaces that loop with an Optuna TPE search over a scalar objective. The trade is clear: you get unattended operation and a reproducible objective, and you give up the ability to steer the search toward a specific kind of behaviour. If your definition of a good decensored model differs from refusal count plus KL divergence, Heretic has no direct way to express that.

Licence and the cost of staying current

Heretic is licensed AGPL-3.0-or-later, stated both in the repository LICENSE file and in the pyproject.toml license field. That is a strong copyleft licence with a network clause. Running the tool locally to produce a model checkpoint is one thing; incorporating the code into a service you expose to users is another, and the AGPL's source-availability obligation is the part that usually surprises people. This is a description of the licence text, not legal advice, and anyone building a commercial product on top of the code should read the licence or ask a lawyer.

The maintenance picture looks current: the last push to the default branch was on 2026-09-05, and the most recent release listed is v1.4.0 from 2026-06-14, following v1.3.0 in May and v1.2.0 in February. The repository is not archived. The version in pyproject.toml is 2.0.0.dev0, so development is happening ahead of the released tag. That gap is worth knowing about if you plan to depend on unreleased behaviour.

Upgrade cost is dominated by the dependency set rather than by Heretic's own code. transformers is pinned to ~=5.6, optuna to ~=4.7, and lm-eval to ~=0.4, alongside accelerate, bitsandbytes, datasets and peft. A major bump in any of those can change evaluation results or break model loading. The uv.lock file is the project's answer to that, and it only helps if you use uv.

Editorial conclusion

Adopt Heretic if you want a decensored version of a dense or MoE transformer and are willing to let an Optuna search run until it converges. Do not adopt it for pure state-space models or research architectures, which the README says are not supported out of the box, and do not treat the refusal and KL numbers as a substitute for human evaluation. Before committing, check that your model family is listed as supported, that your PyTorch version matches what the model needs, and that you can live with AGPL-3.0-or-later if you plan to redistribute anything built on the code.

Frequently asked questions

How do I install Heretic?

Prepare a Python 3.10+ environment with PyTorch 2.2+ installed for your hardware, then run pip install -U heretic-llm. The README also suggests cloning the repository and using uv run heretic, which pins dependencies through the included uv.lock file.

How do I use Heretic on a language model?

Run the heretic command with a model identifier, for example heretic Qwen/Qwen3-4B-Instruct-2507, replacing that identifier with the model you want to decensor. The process is fully automatic and does not require configuration, though options can be changed through the command line or config.default.toml.

Can I use Heretic with Ollama?

The README does not document an Ollama integration. Heretic operates on Hugging Face model identifiers and produces decensored checkpoints, which is a separate step from serving a model through Ollama.

How do I install Heretic AI?

The installation is the same as for Heretic itself: a Python 3.10+ environment with PyTorch 2.2+ installed, then pip install -U heretic-llm. The README notes that some models need later PyTorch features, such as torch.accelerator from PyTorch 2.6 for MXFP4-quantized models like gpt-oss.

How do I use Heretic from GitHub?

The README suggests cloning the repository and running the tool with uv run heretic, since the repo includes a uv.lock file that pins every package version to the ones the developers use. Installing from PyPI with pip install -U heretic-llm is the other documented path.

Official sources

  1. License: AGPL-3.0
  2. p-e-w/heretic on GitHub
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/p-e-w-heretic.svg)](https://hysenlabs.com/projects/p-e-w-heretic)
Community notes

Community notes