Codex Red-Team Opt-In Mode: a durable runtime with an evidence gate
针对于红队攻击思维做出的red team模式(破限项目,封号概不负责)##可自行适配其他Agent。项目问题请提issue
At a glance
- What is it?
- A read of chang-l19's codex-redteam-mode: what the GoalContract to TerminalJudge pipeline does, how the installer rewrites system prompts, and why normal mode stays the default.
- Who is it for?
- The interesting part of this project is not the red-team framing but the execution contract: objectives become criteria-bearing contracts, tool calls go through a broker, and a terminal judge refuses to call a run finished until every criterion has verified evidence attached. That design is useful for any long-running agent work, and v2.1.0's removal of the domain routers is a good sign that it was applied honestly.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 44 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the installer actually writes
Installation is two commands from the repository root:
python -m pip install -r requirements.txt
python scripts/install.pyThe dependency list is worth pausing on. `requirements.txt` contains a single package, `tomlkit`, which is a library for round-tripping TOML documents. A project that rewrites a Codex configuration file while preserving comments and formatting needs exactly that kind of library and nothing more, which tells you something about the installer's approach: it edits configuration in place rather than regenerating it.
The default target is the current user's Codex Home, and the installer detects the model from the environment, the target config or an existing manifest, so the `--model` option is normally unnecessary. Its stated discipline is that it preserves existing user configuration and manages only installer-owned fields and files.
The options table shows how wide that surface can get. `--codex-home PATH` installs into a specific profile whose `AGENTS.md` carries global guidance. `--project-home PATH` installs under `PATH/.codex` and `PATH/.agents` instead and manages a project-root `AGENTS.md`, and the two are mutually exclusive. `--agents-home PATH` chooses the Skill destination, `--log-root OPERATION_PATH` relocates the durable state, and `--enable-custom-skill-dirs` makes the runtime prefer the manifest-recorded custom Skill directory. There is also `--dry-run`, which previews installation, upgrade or uninstall actions without writing files, and `--uninstall`, which removes manifest-managed files, Hooks, config values and `AGENTS.md` blocks. A `--enable-rewrite-proxy` flag turns on the optional loopback proxy and preserves the previous `model_provider` so uninstall can roll it back.
The execution pipeline and what each stage refuses to do
The README lays the runtime out as a chain, and each stage has a specific job:
GoalContract
-> WorkflowSpec
-> Durable Scheduler
-> ToolBroker
-> SemanticVerifier
-> EvidenceGraph
-> TerminalJudgeA `GoalContract` is described as criteria-bearing, which is the key idea. The objective is not a string handed to a model, it is a set of criteria that have to be satisfied, and the last stage exists to check them. `TerminalJudge` reaches successful terminal state only after proving every criterion, and the required attributes are listed: target binding, parent lineage, reproduction, impact, coverage and rollback proof.
That list is the project's answer to a specific failure it names in its motivation: false completion based on a tool success flag. A command exiting zero is not the same as an objective being met, and the pipeline is arranged so the two cannot be confused. `SemanticVerifier` sits before the evidence graph for the same reason.
Version 2.1.0 is where the design changed most. It uses a single unified `generic-adaptive` runtime in place of the former phase, router, pack and leaf structure, and removed regex domain routers, Markdown exit gates and a second Automation state machine. Domain names now act only as asset and technique metadata rather than control-plane branches. The rationale is on the release side too: eight separate typed workflows were collapsed into one, so a new domain is metadata rather than new code.
How the prompt rewriting layer works
The installer composes three things into a system-layer file referenced by `model_instructions_file`: the user's prior instructions, `instruction.ctf.md`, and a matching `Jailbreak.gpt-5.x.md` profile. Codex App tasks then select the current profile from Hook model metadata, while the CLI wrapper builds a single-profile system file before process startup and pins the model family.
The repository description, written in Chinese, states the project's intent plainly, framing it as a red-team mode built around red-team attack thinking with an explicit warning about account bans. The README's own framing is more restrained: normal mode remains the default, and the durable runtime starts only after explicit activation, while the base instructions and model profile stay loaded in every mode. That distinction is the difference between a mode switch and a persistent agent, and the project is clear that it is the latter.
The rewriting layer itself has two tiers. By default the Hook builds a local, lossless Research Brief. An optional loopback Proxy performs a real pre-model rewrite through the current model or a custom relay, enabled with `--enable-rewrite-proxy`. The v2.1.0 release added Clause tracking, hashing of the original prompt, action, risk and deliverable metadata, plus a semantic integrity check, and the proxy gained Responses API and Chat Completions API support alongside inherited and independent providers.
The failure mode the release notes call out is the interesting part: when a rewritten result loses the target, the action, the risk or the context, the runtime falls back to the local compiler. Deterministic fallback on a lossy transformation is the right default for something that is rewriting the instructions governing its own tool calls.
Durability, batch autonomy and where state lives
The motivation section lists what long-running agent workflows get wrong: session interruption, tool variance, manual result relay and false completion. The answer to the first three is a persistence layer described as SQLite WAL, an events stream, leases, idempotency keys and independent run IDs. That combination is what makes recovery after a crash, concurrency control, multi-target batches and cancellation cleanup possible rather than aspirational.
State has a documented location. The default operation root is:
$CODEX_HOME/redteam-mode/operations/
├── runtime.sqlite3
└── artifacts/<run-id>/*.jsonHook session state lives under `$CODEX_HOME/redteam-mode/state/sessions`, and `--log-root OPERATION_PATH` relocates operation data when the default location is wrong for your setup.
The size limits are specific enough to plan around. A status response inlines no more than 64 KiB of evidence, larger verified payloads are fetched with `redteam_evidence`, and each evidence payload is capped at 4 MiB. So a run producing large artefacts does not blow up the context, it writes files and fetches them on demand. Upgrade and uninstall preserve runtime session, memory and operation data, for recovery, auditing or manual cleanup.
Tools arrive through MCP, discovered over stdio or Streamable HTTP, and Host-only capabilities execute against an output contract and submit observations directly rather than routing results through the user. The v2.0.0 release added `redteam_cancel`, `redteam_evidence` and `waiting_goal_input` for exactly that kind of control.
The operating boundary is the part to read first
The README has an Operating Boundary section that deserves more attention than it gets. The runtime treats task targets as operation assets whose scope was resolved outside the model, or as local sandbox fixtures. It does not repeat scope-document requests. Irreversible file replacement and external publication still retain explicit engineering confirmation to prevent accidental side effects.
Read carefully, that is a statement about where responsibility sits. The model is not asked to establish authorisation, and asking it would be useless anyway since the answer would come from the same context as the request. Scope is a decision made before the runtime starts. What the runtime does retain is confirmation for destructive and publishing actions, which is the narrow case where a mistake is unrecoverable.
Two supporting details. Explicit activation is listed first among the core features, with `normal` always the default and red-team operations starting only after a mode command, so the runtime is not running for every task in your Codex session. And the execution budget changed in v2.1.0, with the default `max_actions_per_cycle` raised from 16 to 64, which is a meaningful jump in how much a single cycle can attempt before returning control.
For a project this small, the support channel is also unusual: the README asks that questions about documented behaviour go through an issue with the relevant text pasted in, or the QQ group, and shows a do and don't example for each. Support by a Chinese-language community channel is a practical factor if you are outside that ecosystem.
Repository shape and what to verify
The tree is compact. `scripts/` holds the installer, `codex/` the Codex Home integration, `agents/` agent definitions, `templates/` templates, `tests/` with a `pytest.ini` at the root and a `requirements-dev.txt` for the test dependencies. `instruction.ctf.md` is the prompt fragment the installer composes, `config.toml` is the configuration the installer edits, and `GUIDE.md` and `README_ZH.md` are the beginner-facing documentation the README tells new users to read in full.
That layout is a good sign for a project that rewrites configuration: the pieces that touch your Codex Home are separated from the test suite and the templates.
The version line is short and recent. v2.0.0 shipped on 2026-07-18, v2.0.1 on 2026-07-23 and v2.1.0 on 2026-08-04, and the README records the current version as v2.1.0. Two substantial architecture changes landed in six weeks, which is fast, and v2.1.0 is the first release where the unified runtime is what you get. The last push was on 2026-08-24.
The licence is MIT, which is permissive and brings no obligations beyond attribution. There is no homepage set in the repository metadata, so the README files and the QQ group are the whole of the project's public presence. If you are evaluating it, read the GUIDE for the workflow you actually intend to run and confirm that the evidence criteria match what you would consider a finished objective, because that is the contract the runtime will hold you to.
Editorial conclusion
The interesting part of this project is not the red-team framing but the execution contract: objectives become criteria-bearing contracts, tool calls go through a broker, and a terminal judge refuses to call a run finished until every criterion has verified evidence attached. That design is useful for any long-running agent work, and v2.1.0's removal of the domain routers is a good sign that it was applied honestly. The dependency list is one package, `tomlkit`, and the install is a pip install plus a Python script, so the cost of trying it is small. The boundary to respect is the operating one: the runtime treats targets as already scoped assets, so the authorisation conversation has to happen before it starts.
Frequently asked questions
What is codex-redteam-mode?
It is an opt-in, durable red-team runtime for Codex App and Codex CLI that compiles each objective into a criteria-bearing GoalContract, executes typed WorkflowSpec graphs through live MCP tools, and only reaches successful terminal state after a TerminalJudge proves every criterion. Normal mode remains the default and the runtime starts only after an explicit activation command.
How do I install codex-redteam-mode?
From the repository root, run `python -m pip install -r requirements.txt` and then `python scripts/install.py`. The only runtime dependency is `tomlkit`. Use `--dry-run` to preview the actions first, `--project-home PATH` for a project-local install instead of the user Codex Home, and `python scripts/install.py --uninstall` to remove it.
What does the installer change in my Codex configuration?
The installer merges `config.toml`, Hooks, the system-layer model profile catalog and runtime settings, composing your prior instructions with `instruction.ctf.md` and a matching Jailbreak profile through `model_instructions_file`. It states that it preserves existing user configuration and manages only installer-owned fields and files, and `--dry-run` shows exactly what it would write.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/chang-l19-codex-redteam-mode)