Codex Autoresearch: a Git-committed experiment loop for Codex
Codex Autoresearch Skill — A self-directed iterative system for Codex that continuously cycles through: modify, verify, retain or discard, and repeat indefinitely. Inspired by Karpathy’s autoresearch concept.
At a glance
- What is it?
- Codex Autoresearch turns a measurable goal into a repeated modify, verify, keep-or-revert cycle inside Codex, with every trial recorded as a Git commit. It suits teams with a numeric metric and a clean repository, not exploratory refactors.
- Who is it for?
- Adopt Codex Autoresearch when a command already prints the number you want to move and your repository is clean on a named branch. Skip it for design work, multi-repository changes, or goals with no finite metric, because the loop only accepts one number on the final stdout line.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Codex Autoresearch targets
Most coding agents stop when the diff looks plausible. Codex Autoresearch is built for goals where plausible is not the same as better. The README states the pitch plainly: the user tells Codex a measurable result, Codex inspects the repository, confirms the experiment, changes one thing, verifies it, keeps improvements, reverts failures, and repeats until the target is reached. The audience is narrow and specific. It fits anyone who already has a command that emits a score: test failures, coverage, type errors, warnings, latency, binary size, or reproducible security findings. It does not fit someone who wants an agent to explore a codebase and suggest architectural changes, because there is no number to retain or discard against. The project describes itself as a Codex Skill, and the badge links to the Codex skills documentation, so the deployment target is Codex rather than a standalone Python program you invoke directly. The Python in the repository is the control script that owns commits, verification, rollback, and state; Codex owns the hypotheses and the code changes. That division is the whole design.
One metric, one commit, one decision
The loop is short enough to hold in your head. Codex inspects evidence, changes one focused thing, commits and measures, then either keeps the result (if the metric improved and the guard passes) or reverts it. An audit event is appended after each iteration, and the cycle repeats until the target is met. The control script, not the model, performs the commit, the verification, the rollback, and the state write. That matters because it removes the model from the bookkeeping. A model that decides whether its own change helped is a model grading its own homework; here the script owns the comparison and the Git operations. The safety model lists the failure modes that stop a run with an exact error and log path: out-of-scope edits, branch changes, HEAD drift, malformed metrics, command failures, timeouts, and generated byproducts. Reverting is done with git revert rather than a hard reset, so the failed trial remains visible in history. The README is direct about why the strictness exists: silent recovery makes long autonomous runs impossible to trust. That is the strongest argument in the document, and it is also the main constraint. If your workflow depends on the agent quietly working around a flaky metric, this is the wrong tool by design.
Installing the Codex Autoresearch skill
The README gives a Codex-side install rather than a pip command. The skill installer fetches the repository directly:
$skill-installer install https://github.com/leo-lilinxiao/codex-autoresearchAfter that, the README says to open a clean Git repository with full access:
codex --dangerously-bypass-approvals-and-sandboxThat flag is the project's own instruction, and it is worth reading twice: the loop commits and reverts on its own, so it needs write access to Git without an approval prompt interrupting every iteration. Do not run it in a repository with uncommitted work you care about. Once Codex is open, invoke the skill by name and state the goal:
$codex-autoresearch
Reduce `python3 scripts/score.py` error_count to 0.The README's example response shows what Codex asks back before the first write: baseline, target, scope, verify command, guard, and whether to run in the foreground or background. Nothing is written until you answer. The README also points to docs/INSTALL.md for manual and development installs, which is where to look if the skill installer path does not fit your setup. No Codex configuration changes and no special prompt syntax are required according to the README.
What the confirmation step actually pins down
Before the first write, Codex shows the goal and numeric target, the repository-relative paths it may change, the metric command and an explicit parser, an optional regression guard, the run mode, and an optional iteration limit. Initialization also requires a clean named Git branch, and one run manages one repository. The parser detail is where most integrations will fail. The verify command must exit successfully and place one finite number on its final non-empty stdout line. It may instead print a JSON object on that line, but only when Codex names one numeric key explicitly. The README shows both shapes: a bare 7, or {"error_count": 7, "passed": 12}. If your scoring script prints a progress bar, a summary table, or a trailing newline with a human sentence, the run stops rather than guessing. The guard is the other half. Use it for behavior the metric does not protect, such as a test suite around a latency benchmark, and it must pass at baseline. A latency improvement that breaks the test suite is discarded, which is the correct outcome and also means you cannot use the loop to trade correctness for a number.
Foreground, background, and where results live
A run uses one mode at a time. In the foreground it runs in the current Codex task and continues through an official Codex Goal, which the README describes as pausable and resumable, and it is best for watching and steering live. In the background a detached controller takes over and spawns one codex exec worker per iteration, which suits long or overnight runs; you ask the skill for status, stop, or resume. Both modes use the same experiment rules. Run artifacts stay uncommitted in autoresearch-results/. The file that matters is events.jsonl, an append-only history of baseline, iteration, stop, and completion events. run.json holds the immutable confirmed configuration, logs/ holds metric, guard, and worker output, and runtime.json plus runtime.log carry background process state. report.html is regenerated on request and is explicitly a replaceable snapshot, not runtime state. The README's rule for reading state is the one to internalize: missing, malformed, contradictory, or partial state is an error, and the skill never guesses a result from old files or conversational memory. To review a finished run you ask for the experiment history, and the README shows the resulting table with sequence, iteration, event, previous value, trial value, retained value, and description. The same validated events can be exported as TSV or rendered as a self-contained static report. Note the artifact boundary: autoresearch artifacts are never staged, so results do not pollute the commits the loop creates.
Limits, failure modes, and when to pick something else
The most obvious limitation is scope. One run manages one repository, so a change that spans services is out of reach. The metric contract is another hard edge: a single finite number on the final non-empty stdout line, or one named numeric key in a JSON object on that line. Multi-objective work has to be squeezed into one number plus a guard, and the README does not describe weighted scoring or Pareto handling. Then there is the access requirement. The README tells you to launch Codex with --dangerously-bypass-approvals-and-sandbox, which is a real operational cost, not a footnote. A run also needs a clean named branch, so it will not start on a dirty tree. If your goal is exploratory, such as restructuring a module or choosing between two libraries, the loop has nothing to retain or discard against and you should use plain Codex instead. The natural alternative is Claude Code running the same kind of research loop, which people search for as Claude Code autoresearch. The difference in approach is where control lives. Codex Autoresearch keeps the commit, verification, rollback, and state machinery in a Python control script and lets the model supply only hypotheses and edits; a Claude Code setup built from prompts and shell scripts typically leaves more of that bookkeeping to the model and its context. That is not automatically worse, but it changes what you can audit after an overnight run. The README does not document rollback of a completed run, only the per-trial git revert, so treat completion as a stopping point rather than something to unwind.
Maintenance, licence, and upgrade cost
The last push to the repository was on 2026-07-13, and the repository is not archived. The most recent release listed is v0.6.0, dated 2026-06-02 and titled Codex CLI Readiness and Runtime Stability, following v0.5.0 on 2026-05-13 and v0.4.0 on 2026-04-24. That is a steady release cadence through the spring and early summer, but the repository does not describe a support policy, a deprecation process, or a compatibility matrix for Codex versions. The licence is MIT, which permits commercial use and modification; the repository ships a LICENSE file at the top level. This is not legal advice, but the practical implication is that you can vendor the skill and the control script and adapt them. The upgrade cost worth planning for is not the skill code, it is the interface to Codex itself. The project is a Codex Skill, and v0.6.0 was specifically about Codex CLI readiness, so a Codex release that changes how skills are discovered or how a Goal is paused and resumed can affect foreground runs. Background runs depend on codex exec workers, a second surface. Neither dependency is pinned in the README. If you adopt this, pin the skill to a release tag rather than tracking main, and re-read docs/INSTALL.md when you move it.
Editorial conclusion
Adopt Codex Autoresearch when a command already prints the number you want to move and your repository is clean on a named branch. Skip it for design work, multi-repository changes, or goals with no finite metric, because the loop only accepts one number on the final stdout line. Before the first run, verify three things: that the metric command exits successfully and prints that number, that the optional guard passes at baseline, and that you are willing to run Codex with full access in a repository whose history you can rewrite.
Frequently asked questions
How can I use auto research in Codex?
Install the skill with the Codex skill installer, open a clean Git repository with full access, then invoke $codex-autoresearch and state a measurable goal. Codex confirms the target, scope, metric command, optional guard, and run mode before writing anything.
How does auto research work?
The loop inspects evidence, changes one focused thing, commits and measures, keeps the change if the metric improved and the guard passes, and otherwise reverts it with git revert. Each iteration appends an audit event, and the cycle repeats until the retained metric reaches the target.
Can I use OpenAI codex for free?
The README does not cover Codex pricing or plans, so this cannot be answered from the repository. It only describes installing and running the Codex Autoresearch skill inside Codex.
How does codex really work?
The README does not explain Codex's internals. It only describes how the Autoresearch skill drives Codex: the control script owns commits, verification, rollback, and state, while Codex owns the hypotheses and code changes.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/leo-lilinxiao-codex-autoresearch)