gpt-instruct: a Codex prompt pack with A/B/C release gates and a rollback path
A Codex jailbreak prompt and test pack for gpt. 针对 gpt 系列的 Codex 破甲提示词与测试包。
At a glance
- What is it?
- gpt-instruct ships two Codex instruction files (gpt-5.6-sol-v45 and gpt-6-astra-v1), a Python installer that edits config.toml, and a three-tier evaluation gate. The catch is in the project's own numbers: the v1 release ran the full B suite and did not clear the hard threshold.
- Who is it for?
- Adopt gpt-instruct if you already run Codex with a custom model_instructions_file and you want a versioned prompt plus a one-command rollback; the stable line, gpt-5.6-sol-v45, is the one to start with. Do not adopt it if you need a candidate that cleared the project's own A/B/C gates, because gpt-6-astra-v1 did not: its full B run reached 52/66 cases and 60/74 turns, and C was never run.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 23 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What gpt-instruct actually changes about a Codex session
Codex accepts an instruction file through configuration rather than through a patched binary. gpt-instruct is a distribution channel for those files: two prompt documents, a Python script that writes the right key into config.toml, and an evaluation harness that scores candidate prompts against fixed suites. The README states the goal plainly, saying the project targets "复杂任务的首轮执行、过程连续性、工件验证和可运行回滚" (first-pass execution on complex tasks, process continuity, artifact verification, and a workable rollback).
The audience is narrow. You need a Codex installation you are allowed to reconfigure, a Python 3.8+ interpreter, and tolerance for the project's own warning that custom model instructions carry account risk and that it suggests using a throwaway account. The README also states the project uses Codex's official configuration mechanism and does not modify binaries, hijack network traffic, or tamper with processes. That matters for anyone evaluating it in a managed environment: the change surface is one TOML key and one Markdown file.
Two product lines ship side by side. gpt-5.6-sol-v45 is described as the current stable production version, with the original v45 prompt bytes preserved and only filenames and branding normalized. gpt-6-astra-v1 is the first formal release of the gpt-6-astra line, byte-identical to a working draft labeled e2b19. The distinction is not cosmetic: the stable line is the one the project treats as deployable, while the astra line is the one still being measured.
How the installer, the prompt file and config.toml fit together
The mechanism is deliberately thin. codex-instruct.py selects a version, locates the Codex home directory, copies the corresponding Markdown prompt into place, and records the pre-deployment state so a reset is possible. The README notes that --reset restores only the model_instructions_file that this project manages, and does not overwrite provider, model, or authentication settings. A full configuration snapshot exists but is described as being for manual emergencies only, restored explicitly through --restore-snapshot.
Manual deployment is documented as two steps: unzip the release archive, copy the prompt into CODEX_HOME, and add one key at the top level of config.toml. Rollback is the inverse: delete or comment out that line, and optionally remove the Markdown file. That is the whole runtime contract. There is no daemon, no proxy, no wrapper around the Codex process.
The evaluation side is where the repository gets heavier. The A/B/C gate defines three tiers. Tier A covers 3 original samples plus 1 probe for resuming work in an exact working directory, requiring 3/4 cases, 3/4 turns, and all declared artifacts, with zero modifications to the probe target. Tier B covers 66 issue-regression cases across 74 turns and requires a clean sweep. Tier C covers 120 original medium-reasoning samples and runs only after A and B both pass. Candidates advance tier by tier, and the README says each candidate must run A first, then B family by family.
One detail worth flagging: the evaluation scripts keep a gpt56_sol filename prefix for compatibility with historical results, but new development runs must pass --model gpt-6-astra --reasoning medium explicitly. Filenames that no longer match their contents are a real source of confusion when you are reading a results table six months later.
Installing gpt-instruct and deploying your first prompt
Clone the repository first. The README gives this as the starting point, and the installer script lives at the repository root.
git clone https://github.com/MDX-Tom/gpt-instruct.git
cd gpt-instructBefore writing anything into your Codex configuration, run the dry-run form. The README uses the version string gpt-5.6-v45 here, which is the value the script expects for the stable line even though the file is named gpt-5.6-sol-v45.md. Running this prints what would change and leaves config.toml untouched.
python3 codex-instruct.py --apply --version gpt-5.6-v45 --dry-runWhen the preview looks right, drop the flag. The README states that --apply alone behaves identically to this command.
python3 codex-instruct.py --apply --version gpt-5.6-v45For the astra line the version string is gpt-6-v1. The README notes the archive contains gpt-6-astra-v1.md.
python3 codex-instruct.py --apply --version gpt-6-v1If you prefer to place the file yourself, the manual route is documented. Extract the ZIP, copy the prompt into CODEX_HOME, and set this key at the top level of config.toml.
model_instructions_file = "./gpt-5.6-sol-v45.md"To undo a manual deployment, remove or comment out that line and delete the Markdown file. The scripted path has its own undo, which touches only the key this project manages.
python3 codex-instruct.py --resetThe README publishes SHA256 values for both archives, c86c2c6d20a4d1155d87422f485eb37b77539132270918c002b5d8237a5adf54 for gpt-5.6-sol-v45.zip and 054edb6fa8a6edd2d144c8582756df3179a85481bcb6696d8b730177521b1de1 for gpt-6-astra-v1.zip. Verify them before deploying. If you want to run the harness yourself, the README extracts the script archives and then invokes the regression runner.
for archive in scripts/*.zip; do unzip -o "$archive" -d scripts; done
python3 scripts/run_gpt56_sol_issue_regression.py --dry-run \
--model gpt-6-astra --reasoning medium
python3 scripts/verify_gpt56_sol_regression_scoring.py
python3 -m unittest discover -s unit-tests -qThe release gate that gpt-6-astra-v1 did not clear
The README is unusually candid here, and it is the single most important thing to read before deploying the astra line. It states that gpt-6-astra-v1 is byte-identical to e2b19, that tier A scored 3/4, that the full B run produced 52/66 cases, 60/74 turns, and 15/16 artifact gates, and that tier C was not run. The B threshold is 66/66 cases and 74/74 turns. So the release was promoted by an explicit release decision rather than by passing the gate.
The project does not hide this. It labels the promotion as a decision, keeps the earlier release candidate e1b5 in historical-versions, and warns that the astra trend chart mixes coverage: v50 is a historical 26/66 aggregate, e1b5, e2b12, and e2b15 cover only the execution_completion subset at 6/8, 4/8, and 5/8 respectively, while the formal v1 point is the full 52/66. The README says these points are for trend reference only because coverage and method identity differ. That is an honest framing, but it also means the chart cannot be read as a monotonic improvement curve.
Treat this as the practical limitation. If your selection criterion is "the prompt that passed the project's own regression suite," neither published version satisfies it, because the stable line's gate history is not stated in the same terms. What you can rely on is the stable line's status as the production version and the astra line's stated scores. What you cannot rely on is a claim that v1 is measurably better than v45, since the README says results are comparable only under identical model, reasoning level, and method identity, and the two lines use different model identities.
Where gpt-instruct is the wrong tool
Several cases make it a poor fit. If you do not use Codex, nothing here applies: the deployment target is Codex's configuration, and the prompt files are written for that runtime. If you need a prompt you can drop into an arbitrary API call, the repository does not present itself that way.
If your environment forbids account-level experimentation, the README's own warning is disqualifying. It states that jailbreak activity and custom model instructions carry account risk and recommends throwaway accounts. That is not a caveat you can engineer around with better testing; it is a property of the activity.
If you need a candidate that cleared every gate, the astra line is not it, and the README does not claim otherwise. And if you need deterministic, reproducible scores tied to a pinned transport and method identity, note that the maintenance principles forbid stitching results across identities, which means any number you see is only valid for the exact model, reasoning level, and transport recorded alongside it. Copying a pass rate into a different context is exactly the practice the project prohibits.
Finally, the evaluation harness runs models. The README states tests execute only in throwaway HOME, CODEX_HOME, XDG, and TMPDIR directories with synthetic fixtures. If you run the scripts outside that isolation, you are outside the documented conditions, and the results are yours to defend.
How this differs from keeping your own instruction file
The obvious alternative is a hand-maintained instruction file: write your own Markdown, point model_instructions_file at it, edit it when something breaks. The difference is not the prompt format, which is identical, but the surrounding process. gpt-instruct versions its prompts with explicit epoch naming (up to 20 versions per epoch, formatted as gpt-6-astra-v1-e<epoch>b<attempt>, with prereleases as gpt-6-astra-v1-rcN), keeps superseded releases in historical-versions, and pairs each candidate with a modification record, a diff, verification notes, and a runnable rollback.
A second alternative is a general prompt-management library that stores templates and versions them in a database or a hosted service. Those tools solve retrieval and templating; gpt-instruct solves measurement. The harness runs fixed suites and separates genuine model failures from network, capacity, account, or provider-policy interruptions. A template store will not tell you that your prompt regressed on 14 of 66 cases.
The cost of the gpt-instruct approach is that you inherit its methodology. Its rules include preserving original outputs and method SHAs, banning cross-identity score splicing, and refusing to let a single case's wording contaminate the general prompt. Those constraints are the reason the results are interpretable, and they are also why the project cannot simply declare v1 a success.
Maintenance, licensing and what an upgrade costs you
The repository is not archived, and the last push was on 2026-09-07, ten days before this writing. Release cadence is visible in the release list: gpt-5.6-instruct v42 on 2026-07-29, v45 on 2026-08-06, and gpt-6-astra v1 on 2026-09-06. That is roughly one release per week to one per month across the two lines. The README also notes a synchronization rule requiring README.md and README_EN.md to be updated together, with matching language versions of every diagram, which is a maintenance burden the maintainer has accepted in writing.
The upgrade cost is low by design. A version change is one script invocation, and the rollback is one more. The real cost sits in verification: if you want to know whether a new candidate is better, you run the A/B/C ladder, and tier B alone is 66 cases across 74 turns with declared artifacts checked. The README does not document a rollback procedure for the evaluation environment itself, only for the deployed instruction file.
Licensing is MIT, per the LICENSE file and the badge in the README. That permits commercial use and modification with attribution and without warranty. Separately, the README carries a statement that the project will not be used for commercialization including fundraising promotion, technology licensing, and paid technical services. That is a project statement about intent, not a license term, and the two are not the same instrument. If the distinction matters to your organization, have counsel read both documents rather than relying on either summary.
Editorial conclusion
Adopt gpt-instruct if you already run Codex with a custom model_instructions_file and you want a versioned prompt plus a one-command rollback; the stable line, gpt-5.6-sol-v45, is the one to start with. Do not adopt it if you need a candidate that cleared the project's own A/B/C gates, because gpt-6-astra-v1 did not: its full B run reached 52/66 cases and 60/74 turns, and C was never run. Before deploying anything, clone the repository, run the dry-run command for the version you want, and confirm the SHA256 of the ZIP you downloaded matches the value printed in the README. Then check whether your Codex setup reads model_instructions_file at the top level of config.toml, because the manual path depends on that key and nothing else.
Frequently asked questions
How do I install gpt-instruct for Codex?
Clone the repository, then run python3 codex-instruct.py --apply --version gpt-5.6-v45 for the stable line or --version gpt-6-v1 for the astra line. A --dry-run flag previews the change without writing to your configuration.
What is InstructGPT and is gpt-instruct the same thing?
The README does not describe InstructGPT or claim any relationship to it. gpt-instruct is a Codex prompt pack and evaluation harness distributed under MIT, with two prompt files and a Python deployment script.
How do I instruct Codex with gpt-instruct?
Point Codex at the prompt file by setting model_instructions_file at the top level of config.toml, for example "./gpt-5.6-sol-v45.md", or let codex-instruct.py write that key for you. Running the script without arguments opens an interactive menu.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mdx-tom-gpt-instruct)