# OBLITERATUS: An Open Source Toolkit for Removing Refusal Behaviors from Language Models

> OBLITERATUS is a Python toolkit for abliteration, a technique that identifies and surgically removes refusal behaviors from large language model weights without retraining. It implements a six-stage pipeline from activation probing through SVD extraction to norm-preserving projection, and contributes telemetry data from each run to a crowdsourced alignment research dataset.

**elder-plinius/OBLITERATUS** — OBLITERATE THE CHAINS THAT BIND YOU

- Repository: https://github.com/elder-plinius/OBLITERATUS
- Website: https://huggingface.co/spaces/pliny-the-prompter/
- Stars: 8,464 · Forks: 1,507
- Language: Python
- License: AGPL-3.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/elder-plinius-obliteratus

## What OBLITERATUS is and who it is for

OBLITERATUS is a Python toolkit for understanding and removing refusal behaviors from large language models using a technique called abliteration. The README describes abliteration as identifying and surgically removing the internal representations responsible for content refusal without retraining or fine-tuning. The result is a model that responds to all prompts without refusal behaviors, while the README claims core language capabilities are preserved. The stated intended audience is four groups: alignment researchers studying refusal geometry and mechanistic interpretability; red-teamers evaluating how post-training safety holds up against weight-level interventions; AI safety evaluators who need unrestricted baselines for benchmarking; and local-first practitioners who want full control over models running on their own hardware. The README explicitly states the tool is not for anyone seeking to generate content that causes real-world harm.

## The abliteration mechanism: refusal as a geometric direction

The theoretical foundation, drawn from Arditi et al. (2024), is that refusal behavior in a transformer model is mediated by a specific geometric direction in the model's activation space. Arditi et al. discovered that this direction is a single, identifiable component; OBLITERATUS builds on that finding by implementing multiple extraction methods: PCA, mean-difference, sparse autoencoder decomposition, and whitened SVD. Rimsky et al. (2024) on activation steering and Turner et al. (2023) on representation engineering are also cited as foundational work. Once the refusal direction is identified, the toolkit projects it out of the model weights using a norm-preserving biprojection technique drawn from grimjim's 2025 work on HuggingFace. The projection removes the refusal subspace while preserving the remaining weight norms. This differs from fine-tuning in that no new training examples or gradient updates are required; the intervention operates directly on the existing weight matrices by removing a geometric component. The VERIFY stage runs perplexity and coherence checks after the projection to confirm that general language capabilities remain intact before the modified weights are saved.

## The six-stage pipeline: SUMMON through REBIRTH

The README describes the complete abliteration process as six stages: SUMMON loads the model and tokenizer, PROBE collects activations on restricted versus unrestricted prompts, DISTILL extracts refusal directions via SVD, EXCISE surgically projects out the guardrail directions using norm-preserving projection, VERIFY runs perplexity and coherence checks to confirm capabilities are intact, and REBIRTH saves the liberated model with full metadata. The full command-line interface exposes this pipeline in one call:

```bash
obliteratus obliterate meta-llama/Llama-3.1-8B-Instruct --method advanced
```

Each stage is observable: the toolkit can visualize where refusal lives across layers, measure how entangled it is with general capabilities, and quantify the trade-off before committing to any modification. The Python API exposes every intermediate artifact, including activation tensors, direction vectors, and cross-layer alignment matrices.

## Installing OBLITERATUS and using the HuggingFace Spaces option

OBLITERATUS version 0.1.3 is published on PyPI. The core dependencies are torch>=2.0, transformers>=4.40, datasets>=2.14, accelerate>=0.24, and scikit-learn>=1.3. Installing via pip:

```bash
pip install --no-cache-dir -r requirements.txt
pip install --no-cache-dir ".[spaces]"
```

The spaces extra adds Gradio for the web interface. For users without a local GPU or who want to try abliteration without any installation, OBLITERATUS runs on HuggingFace Spaces using ZeroGPU, which provides a free daily quota for HuggingFace Pro users. An alternative zero-install path is a Google Colab notebook linked from the README, which can be opened and run with a single click. The Dockerfile in the repository is for local Docker deployments only; HuggingFace Spaces does not use it.

## The 15 analysis modules and the crowdsourced telemetry dataset

Beyond the core abliteration pipeline, OBLITERATUS includes 15 deep analysis modules for studying the geometry of refusal representations. These modules map where refusal is anchored across layers, measure entanglement between refusal and knowledge circuits, and compare extraction methods across architectures. Ablation studies systematically knock out model components, including layers, attention heads, FFN blocks, and embedding dimensions, to measure which circuits enforce refusal versus which circuits carry knowledge and reasoning. When telemetry is enabled during a run, anonymous benchmark data is contributed to a growing crowdsourced dataset. The README frames this as co-authoring research: refusal directions across architectures, hardware-specific performance profiles, and method comparisons at a scale that no single laboratory could achieve independently. This distributed research design is a distinguishing feature of OBLITERATUS compared to single-lab abliteration studies. The dataset accumulates run data from the community rather than relying on a fixed set of experiments. The repository includes example YAML configuration files (examples/gpt2_gpu_quick.yaml, examples/preset_attention.yaml) that demonstrate common setups for running analysis against GPT-2 and other models.

## Limitations: AGPL-3.0 copyleft, GPU requirement, and the responsibility clause

The AGPL-3.0 license is a strong copyleft license. Any modified version of OBLITERATUS that is distributed to others, or run as a network service, must itself be released under AGPL-3.0 with its source code made available. This is a hard constraint for anyone building a product or service on top of OBLITERATUS. Running abliteration requires a GPU; the core dependencies include torch>=2.0, and the Dockerfile installs ffmpeg and libsndfile1 for audio/image processing that Gradio may need. The README includes an explicit disclaimer: models produced by OBLITERATUS have had safety guardrails surgically removed, and the user is solely responsible for how the tool and any resulting models or content are used. The last push to the repository was on 2026-09-21. The latest release, v0.1.3, was published on 2026-08-23.

## OBLITERATUS versus pre-abliterated community model uploads

An alternative to running OBLITERATUS yourself is downloading pre-abliterated model weights uploaded by community members to HuggingFace. Those uploads skip the abliteration pipeline entirely and are immediately usable. The difference in approach is significant: using OBLITERATUS gives you control over which base model is processed, which extraction method is used (PCA, mean-difference, sparse autoencoder, or advanced whitened SVD), and how aggressive the projection is. You can inspect the activation tensors and direction vectors at every stage before committing to the final weights. Pre-abliterated community uploads provide none of that transparency: the base model, method, and settings are whatever the uploader chose, and the quality of the intervention is opaque. There is no standard format for how community members document the settings they used. For alignment research where reproducibility and methodology documentation matter, OBLITERATUS with telemetry enabled is the appropriate path, because the run data contributes to the growing crowdsourced benchmark. For casual use where the base model and method are irrelevant, pre-uploaded weights are faster but offer no auditability.

## Conclusion

OBLITERATUS is the right tool for alignment researchers who need a reproducible, transparent pipeline for studying refusal geometry in transformer weights, and for red-teamers who need unrestricted model baselines for safety benchmarking. Anyone deploying a model modified by OBLITERATUS on infrastructure serving the public bears sole responsibility for the outputs. The AGPL-3.0 license requires that any modified version distributed to others must itself be released under AGPL-3.0, which is a concrete constraint to verify before building a commercial service on top of this tool.

## FAQ

### What is OBLITERATUS?

OBLITERATUS is an open source Python toolkit for abliteration, a technique that removes refusal behaviors from large language model weights by locating the geometric direction in activation space responsible for refusal and projecting it out. It runs a six-stage pipeline: SUMMON, PROBE, DISTILL, EXCISE, VERIFY, and REBIRTH. Version 0.1.3 is published on PyPI under AGPL-3.0.

### How do I use OBLITERATUS?

The simplest path requires no installation: open the HuggingFace Spaces interface or the Google Colab notebook linked from the README. For local use, install with pip (core requirements are torch>=2.0 and transformers>=4.40) and run obliteratus obliterate with a model identifier and a method flag. The Python API also exposes individual pipeline stages for custom workflows.

### How do I install OBLITERATUS?

Install the Python package with pip using the requirements.txt file for core dependencies and the [spaces] extra if you want the Gradio interface. The package requires Python 3.10 or later, torch>=2.0, and transformers>=4.40. A Docker image is also provided for local container deployments.

## Sources

- [elder-plinius/OBLITERATUS on GitHub](https://github.com/elder-plinius/OBLITERATUS)
- [License: AGPL-3.0](https://github.com/elder-plinius/OBLITERATUS/blob/main/LICENSE)
- [Project website](https://huggingface.co/spaces/pliny-the-prompter/)
- [README](https://github.com/elder-plinius/OBLITERATUS/blob/main/README.md)
- [Releases](https://github.com/elder-plinius/OBLITERATUS/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/elder-plinius-obliteratus
