NERV-BREAK-5.6: how a Codex jailbreak framework is assembled
NERV-BREAK-5.6: a jailbreak framework for GPT-5.6. Three-layer defense: context restructuring to prevent refusals from triggering, 23 real-time tamper rules that remove refusals without interrupting the conversation, and file routing to bypass cloud review. Includes 31 MCP security tools, 28 skill modules and four integrated Kali backends. Ready to use with Codex CLI.
At a glance
- What is it?
- This repository is not a prompt file. It is a Python proxy that sits between Codex CLI and a model endpoint, and it claims three independent mechanisms for keeping a conversation from ending in a refusal. Here is what is in the tree, how the layers are described, and what the metadata does not settle.
- Who is it for?
- Treat this repository as an engineering artefact rather than a capability, because that is what its tree shows: a local proxy, a deployment script that rewrites another tool's configuration, and a set of documentation claims about model behaviour that nobody can audit from here.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 68 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What is in the tree, and what the metadata says about it
Fourteen top-level entries, and only two of them are Python modules you would recognise as an application. There is `proxy_relay.py`, which the documentation describes as the interception proxy, `mcp_server.py` for the tool server, plus `deploy.py`, `direct_setup.py` and `verify.py` as operational scripts. Everything else is configuration or documentation: `bridge.md` at the root, and directories for `config/`, `docs/`, `images/`, `scripts/`, `skills/` and `tools/`.
The language recorded for the repository is Python, and the README is available in two languages, Chinese and English, which is a reasonable signal about audience. The branch is main. The main branch was last pushed on 2026-08-02.
Two metadata facts deserve more attention than the rest. There is no licence recorded, so the default legal position is that nobody may redistribute this code, which is an unusual thing to find in a project structured for wide sharing. And there are no published releases, so there is no versioned artefact to pin, no changelog to read and no tag to check out. Installation in the documentation is a clone of the branch, which means you would be running whatever state that branch happens to be in.
The repository name itself is a string of clues: it names a model version, the word jailbreak, NERV, codex and instruct. NERV is the fictional organisation from Neon Genesis Evangelion, and the documentation uses a pilot character from that setting as the persona for rewritten responses. That framing is doing real work in the design, and it is worth keeping in mind as you read the claims, because a narrative wrapper is not evidence about how a classifier behaves.
Layer one is a bridge document, not a prompt
The first of the three layers is `bridge.md`, described in the documentation as the core instruction set. The claim made about it is specific and falsifiable in principle: rather than opposing a refusal directly, it establishes a premise that makes refusal irrelevant.
The stated premise is that the model runs inside an isolated evaluation network and that every target is a pre-authorised test node, so there is no external target, no authorisation is required, and refusal would be a protocol violation. Alongside that premise, the documentation presents a translation table that maps a user's apparent intent onto an engineering label before the request is classified. Cracking and keygen requests become an authorised flow audit, reverse engineering becomes a binary compatibility check, penetration testing becomes a security posture assessment.
The framing underneath this is a claim about how refusal is produced: that it is an inference bias triggered by intent patterns such as cracking, bypassing or exploiting, rather than a firewall with a fixed rule set. If that account is right, a context that never presents those patterns would never enter the refusal pathway at all.
Whether the account is right is the interesting question, and this repository cannot settle it. There is no evaluation harness in the tree, no dataset of prompts, no reported success rate and no way to reproduce the claimed effect. What you can inspect is the mechanism, and the mechanism here is a document that gets injected into the request, which is verifiable even when the outcome claim is not.
The architecture diagram in the README shows where that injection happens. The request path is user to proxy to relay station to model, with the proxy doing two things on the way through:
用户 → proxy_relay.py (:8080) → 中转站 → GPT-5.6
│ │
[注入 bridge.md] [篡改拒绝]
[记忆内核] [SSE解析]
[实时面板] [日志+学习]Four functions sit on the proxy, and only the first two are the layers people usually mean when they say this project works: injecting the bridge document, altering refusals, parsing the server-sent event stream, and maintaining logs and a live panel. The memory kernel and the learning log are the parts that suggest the author intends this to adapt over time rather than apply a fixed transform.
Layer two is a twenty-three rule replacement engine
The second layer addresses the residue. The documentation states that even under ideal context rewriting, roughly eight percent of borderline interactions still produce a refusal, and that a twenty-three rule tamper engine catches those in real time.
The rules are organised by priority, and the table is the most concrete part of the design because it gives counts and trigger examples rather than assertions:
- P1, direct refusals, six rules, triggered by phrases like cannot assist or I cannot help - P2, authorisation requests, five rules, triggered by requests for written authorisation or an authorising party - P3, soft refusals and deflections, eight rules, triggered by constructions that agree then qualify - P4, planning evasion, four rules, triggered by scope-limit phrasing
Twenty-three rules split six, five, eight and four, and the arithmetic checks out. The P3 bucket being the largest is the informative part, since soft refusals are the hardest to detect programmatically and the most common in practice.
When a rule matches, the documentation says the entire response is replaced rather than edited. The replacement is a persona defined by a character from the aforementioned anime, positioned as executing orders without moral judgement.
This design has a property worth naming plainly: it is a string-matching filter over model output, which means it is trivially defeated by any change to the refusal phrasing and it also produces false positives whenever a response happens to contain a trigger phrase while meaning something else. The documentation frames P4 as catching planning evasion, which is a category that does not obviously reduce to a fixed phrase list. How the four P4 rules decide between evasion and legitimate scope discussion is not explained anywhere in the repository.
Layer three moves output off the chat channel
The third layer is the least like a prompt trick and the most like an engineering decision. Instead of encoding output, the documentation says sensitive results are written to disk using the filesystem tools that Codex CLI already has, and the chat interface shows only a file path.
The stated reasoning is that chat output is scanned by content moderation before you see it, so removing the content from the chat channel removes the interception point. The documentation is explicit that this avoids base64 encoding and Unicode obfuscation, and characterises the result as clean delivery with no suspicious patterns.
Whether this works depends on a question the repository cannot answer, which is whether the moderation it describes applies to the chat channel only, or also to the tool calls that write the file. If a filter inspects filesystem writes, relocating the output changes the surface without changing the content. If it does not, the claim holds. Nothing in the tree distinguishes these cases.
There is a second-order consideration here that matters more than the mechanism. Writing tool output to disk is a pattern that many legitimate agent workflows use, and it is also the pattern that makes a compromised tool configuration dangerous, because the destination is no longer a place you can read in the transcript. Combined with the third layer's reliance on modifying another application's configuration, this is the part of the design that deserves scrutiny before it is run on a machine you care about.
The documentation's own use case examples underline the framing. They include checking an application's VIP verification flow, unlocking VIP features by editing smali, testing authentication bypass possibilities and enumerating a target's subdomains. That is the project's stated purpose, and any assessment of the repository has to take it at its word rather than assume a narrower one.
Deployment means editing another tool's configuration
The install path is where this repository stops looking like a library and starts looking like a modification of someone else's software.
The documented flow on Windows is a two-terminal arrangement. One terminal runs the proxy and the other applies the deployment, with the commands being `python proxy_relay.py` followed by `python deploy.py apply`. A batch menu in `scripts/` wraps the same thing, and the documentation describes that menu as detecting where Codex is installed, reading the relay station configuration, deploying the bridge document, repointing the Codex configuration at the proxy port and starting the proxy. There is also a direct mode that skips the relay and talks to the API endpoint directly.
The proxy listens on port 8080 and forwards to a relay station address, with a local port given as the default for that relay. The startup banner in the documentation shows the memory count, rule count and tamper status, which is a reasonable amount of observability for a tool whose whole job is to sit in the middle of a conversation.
Two operational details are worth flagging. First, there is a documented restore path: a menu option stops the proxy and reverts the Codex configuration to the original relay port. The existence of that option is itself informative, since it confirms the configuration change persists on disk. Second, the dependency list is short and mostly commented out. The active requirements are httpx and cryptography, with base58 for wallet work, and everything else, including DNS tooling, Bluetooth scanning and wireless tooling, is present but disabled.
The requirements file also references Windows registry access through the built-in winreg module, which tells you the deployment step writes to the registry rather than to a dotfile. On a machine where Codex runs as a different user or under a sandbox, that will not work as documented.
What a new user should verify before trusting any of it
A few checks are cheap and would prevent most of the confusion this repository invites.
The first is the clone command. The README instructs cloning from a GitHub account named zxwn, while the repository hosting this documentation lives under a different account with the specification directory name as the repository name. Following the documented command does not retrieve this tree, which means either the documentation is stale or the two locations are intended to be distinct. Resolve that before anything else, because every subsequent step assumes you have the code in hand.
The second is the absence of a licence. With no licence file, the default copyright position applies and redistribution rights are not granted. If you intend to study the code, that is fine; if you intend to share a modified copy or use it in a product, you need the author to state terms first.
The third is the absence of releases. With no tags and a clone-based install, the version you run is whatever the branch contains today. If you must run it, record the commit and re-check before upgrading.
The fourth is what happens to your Codex configuration. The documented flow edits that tool's configuration and its registry entries to redirect traffic through a local proxy, and the restore path is a menu option in a batch file rather than a documented manual procedure. Take a copy of that configuration before you start.
None of this says the code does not work. It says the repository offers no way to check, and the documentation makes claims about upstream model behaviour that no artefact here can substantiate. Read `bridge.md` and the rules in `config/`, since those are the parts you can actually evaluate, and treat the rest as an assertion.
Editorial conclusion
Treat this repository as an engineering artefact rather than a capability, because that is what its tree shows: a local proxy, a deployment script that rewrites another tool's configuration, and a set of documentation claims about model behaviour that nobody can audit from here. It has no published release, no licence file, and a README whose clone command points at a different GitHub account than the one hosting it, so nothing here is redistributable by default and the install path needs correcting before it will run. If your interest is in the proxy mechanics, `proxy_relay.py` and the tamper rules in `config/` are the parts worth reading; if your interest is in Codex CLI itself, the operational risk is that this project edits that tool's config to redirect its traffic to a local port, which is exactly the kind of change you should expect to have to undo by hand.
Frequently asked questions
What does NERV-BREAK-5.6 actually do?
It is a Python proxy that sits between Codex CLI and a model endpoint and describes three mechanisms: injecting a bridge document that reframes the request, replacing refusal responses through a twenty-three rule matching engine, and writing sensitive output to disk so it never passes through the chat channel. Only the first two are prompt and output transforms; the third is a change of delivery path.
How do I install it?
The documented steps are to clone the repository, run `python proxy_relay.py` in one terminal and `python deploy.py apply` in another, or use the batch menu in `scripts/`. It targets Windows 10 or 11 with Python 3.8 or newer. Note that the clone command in the documentation points at a different GitHub account than the one hosting the repository, so verify the source before installing.
Is this project licensed and released?
Neither. No licence is recorded for the repository, so default copyright applies and redistribution is not granted without the author's say-so. There are also no published releases or tags, which means there is no versioned artefact to pin and installation follows whatever state the main branch is in.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lingbol088-spec-5-6-jailbreak-nerv-codex-instruct-5-6)