Library / SDK
adamallcock/codex-chatgpt-control avatar
adamallcock/codex-chatgpt-control

codex-chatgpt-control drives the visible ChatGPT tab, and refuses to do the rest

Unofficial SDK for Codex agents controlling visible ChatGPT web sessions

406 stars32 forksJavaScriptMIT

At a glance

What is it?
codex-chatgpt-control is an unofficial alpha SDK that lets a Codex style agent hand work to a visible, signed in ChatGPT Chat or Work session and read the result back. The interesting design decisions are all about restraint: two language packages over one local backend, a plugin that ships skills but no browser and no credentials, run reports that leave prompt text out, and an explicit list of things the project refuses to become.
Who is it for?
This SDK fits one narrow job: an agent that already runs inside a Codex or browser bridge host, working with a user who is present and watching a signed in ChatGPT tab. It does not fit unattended automation, headless pipelines, or anyone who wants a private API, because the whole surface is user visible controls and the project says so in its own scope list.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two packages over one local backend, and a root that never ships

The install path forks by language, and the split is not cosmetic. The Node package is called the browser control runtime authority, and the Python package is a parity client over the same local backend protocol. In other words the Python side is a second front end onto a backend that already exists, not a reimplementation, and the repository carries parity tooling to keep the two honest.

Both installs are prerelease installs:

bash
npm install codex-chatgpt-control@next
bash
python -m pip install --pre codex-chatgpt-control

The npm detail matters more than it looks. Prereleases are published under the `next` tag, so a plain `npm install codex-chatgpt-control` will intentionally resolve to an older `latest` release rather than to what the repository is currently working on. The Python equivalent is the `--pre` flag. An agent that installs without either marker gets a build that is not the one the maintainer is testing.

At the root, package.json is named `codex-chatgpt-control-repo`, marked private, and versioned 0.5.1-alpha.5, with the description Public source repo for the unofficial visible ChatGPT Chat and Work control SDK. It is the development workspace, not the shipped artifact, with `packages/`, `plugins/` and `skills/` underneath it. The releases follow the same prerelease line: v0.5.1-alpha.3 on 18 August 2026, v0.5.1-alpha.4 on 7 September, and v0.5.1-alpha.5 on 20 September, which is also the date of the last push to main. Three tags, one version line, and no stable release in between is the honest summary of where this sits. The repository declares no homepage, and the badges at the top of the README point at a parity workflow, the npm package and the PyPI package, which are the three things a release has to keep in step.

Chat and Work get detected, inspected, then configured with strict

The premise is that one desktop shell can hold two different execution experiences. Codex is described as the supported home for local repository work, editing, commands, tests, branches and deployment, while visible ChatGPT Chat and Work carry their own conversation and task state, controls, files, progress and artifacts. The SDK's answer is to treat each surface as a capability to be discovered rather than a picker to be assumed, and the API shape shows that immediately.

Detection and inspection come first, and inspection is read only:

ts
const surface = await chatgpt.experience.detect();
const capabilities = await chatgpt.configuration.inspect();

await chatgpt.experience.open({ experience: "work" });
await chatgpt.configuration.apply({
  experience: "work",
  desired: {
    model: "GPT-5.6 Sol",
    effort: "High",
    speed: "Standard"
  },
  strict: true
});

The order is the point. A caller finds out which surface it is on and which controls that surface actually exposes to the signed in user, then asks for a change, and the `strict: true` flag is the verification step that makes the request a postcondition rather than a hope. Configuration is expressed in terms of the controls the product actually renders, which is model, effort and speed in this example, rather than a flat list of every option anyone has ever seen.

The README cuts off mid sentence right after this block, at the words Do not assume the curr, so the guidance that would have followed the example is not visible here.

Submit once, so a timeout cannot create a second task

The second design decision is about what happens when a call takes too long. Work start, status, wait, steer, read and artifact retrieval are separate operations, and the stated reason is that separating them means timeouts do not create duplicate tasks. This is the failure mode of every agent that wraps a long running web action in a single fire and forget call: the request times out, the agent retries, and the user ends up with two of the same task running in a visible tab.

So the run path is deliberately chatty. An agent opens or continues a visible ChatGPT thread, starts a Work task, polls it, steers it, and reads it, and the thread or task identity is preserved across those steps rather than being re established each time. The client is created against a host that provides `globalThis.agent`, a sub agent carries a name and instructions, and the runner takes the target experience and a response format:

ts
import { createChatGPT } from "codex-chatgpt-control";

const chatgpt = createChatGPT({ agent: globalThis.agent });
const reviewer = chatgpt.agent({
  name: "reviewer",
  instructions: "Review carefully and return Markdown."
});

const result = await chatgpt.runner.run(reviewer, {
  input: "Reply with a one-sentence summary of this project.",
  thread: { type: "new" },
  experience: "chat",
  response: { format: "markdown" }
});

console.log(result.output_text);

`thread: { type: "new" }` is where the identity rule shows up, because continuing a thread is the alternative, and the instruction shipped for repositories tells an agent that when the user says a ChatGPT thread is already open, it should reuse that tab through `existingTab` in Node and `existing_tab` in Python instead of opening a replacement. Files follow the same visible path. Approved local files are attached through visible upload controls, generated files and artifacts come back through visible downloads, and image only artifacts get their own case, where the agent waits for a file that has no text response to read.

The visible part matters as much as the sequence. Everything here is a control a person could click, in a tab they can see, and the SDK's stated position is that it keeps the model in front of the product the user is already looking at, rather than reaching around it.

The plugin ships three skills and no browser at all

Codex Desktop users install this as a plugin from the repository itself, at plugins/codex-chatgpt-control, which the README calls the easiest way to make agents use the SDK consistently instead of hand rolling browser commands. The two commands are a marketplace add and a plugin add:

bash
codex plugin marketplace add adamallcock/codex-chatgpt-control --ref main
codex plugin add codex-chatgpt-control@codex-chatgpt-control

Upgrading is two more commands and one instruction: refresh the marketplace snapshot, add the plugin again, then start a new Codex thread so the updated skill metadata is actually loaded. Skipping that last step leaves a thread running with the previous version's instructions in memory, which is the kind of bug that reads as the SDK misbehaving.

Three skills ship inside. `codex-chatgpt-control` is the broad visible Chat and Work workflow with diagnostics, `chatgpt-delegate` is described as the preferred surface neutral delegation workflow, and `chatgpt-pro-consult` is a backward compatible alias for the visible Chat Pro setting. There are also bundled Node runtime files for bridge enabled imports, and a skill only fallback that copies the skill directory into `~/.codex/skills/` with rsync.

What the plugin does not contain is the more useful part. It bundles no browser bridge, no credentials and no ChatGPT account access. A real workflow still needs a compatible Codex or browser bridge and a visible, signed in session, so the plugin is an operating guide with a runtime attached, not a turnkey automation.

Stop reasons and reports that leave the prompt text behind

Two features here exist to keep an agent honest rather than to make it faster.

The first is the stop reason. The stated capability is to tell the agent exactly why it could not continue when ChatGPT needs login, a captcha, permissions, or UI review. That is a short list, and every item on it is a point where a human has to step in. An agent that gets a login wall in front of it has three bad options: retry in a loop, try to solve the captcha, or give up quietly. The SDK names the third as the behaviour, and the skill instruction that goes into a repository says to report the SDK stop reason and not to retry blindly. Login and captcha are also why the project can describe itself as visible session only and still be honest about it.

The second is the run report. Local reports are privacy preserving by default, omitting prompt and response content. For a tool whose entire job is shuttling a user's document into a chat tab and a chat answer back out, a default that keeps both sides out of the report file is the difference between a debugging aid and a leak. The scope list also says the primitives are workflow level, with no private endpoint access, which is the mechanism behind that default rather than a promise layered on top of it.

The scope list is a list of refusals

Most READMEs describe what a project is. This one spends a paragraph on what it is not, and that paragraph is the most useful thing in the file.

The stated refusals: it is not a generic browser automation framework, not a scraping tool, not an OpenAI API wrapper, not an official OpenAI project, and not a replacement for the official Codex SDK. Elsewhere the same point is made again with account automation and hidden ChatGPT access added to the list, and the closing line states that the project is not affiliated with, endorsed by, or sponsored by OpenAI.

Each refusal rules out a plausible use. Without generic automation there is no page-object model to build on. Without scraping there is no data extraction path around the interface. Without an API wrapper there is no unofficial endpoint to depend on, which is also why the visible surface has to be driven through the real UI and why the captcha and login stop reasons cannot be engineered away.

What remains is narrow on purpose: agent, then browser, then chatgpt.com. A reader who needs a background job that nobody watches will find every one of these refusals inconvenient, and a reader who wants a colleague to look over their shoulder while it runs will find the same list reassuring.

The release scripts are the documentation of what a release must prove

The root package.json is a script manifest, and reading it tells you what this project considers finished.

The Node side has separate gates for running tests and for building, and a bundle target that is not one command but five: a main bundle, then `bundle:backend`, `bundle:live-smoke`, `bundle:release-canary` and `bundle:journal`. A live smoke bundle and a canary bundle are what you build when the artifact under test is a browser extension of somebody else's web page, where the thing that breaks is the far side.

The contracts target is four checks in sequence: `contract:validate`, then `docs:drift`, then `parity:fixtures` and `parity:suite`. Docs drift is the one worth pausing on, because it means the documentation is validated against the code rather than trusted. Parity fixtures and suite are the automated form of the claim that the Python client matches the Node runtime.

The Python gate is driven from Node, through `scripts/run-python-gate.mjs`, with test, compile, pyright and ordinary-shell modes, which keeps one entry point for CI. Plugin work has build, check and validate scripts, and the release chain includes version and name checks, a node pack check, a Python sdist and wheel build, a twine check on the built artifacts, and a source smoke install run through `scripts/verify-release-install.mjs --source`, with a matching verify step for what has already been published. A published package that installs cleanly from the registry is a different question from one that installs from a local build, and having both is the point of splitting the script names. The repository also carries AGENTS.md, SECURITY.md, CONTRIBUTING.md and a CHANGELOG.md, and the top level listing shows a docs/ directory, .agents/ and .github/ alongside them, which is the shape of a project that expects agents and humans to work in it at the same time.

Editorial conclusion

This SDK fits one narrow job: an agent that already runs inside a Codex or browser bridge host, working with a user who is present and watching a signed in ChatGPT tab. It does not fit unattended automation, headless pipelines, or anyone who wants a private API, because the whole surface is user visible controls and the project says so in its own scope list. Before depending on it, pin the prerelease tag rather than latest, expect alpha churn across the 0.5.1 line, and read SECURITY.md, since a tool that drives a live logged in session inherits every risk of that session.

Frequently asked questions

What is codex-chatgpt-control and what does it control?

It is an unofficial alpha SDK for agents that need to delegate user directed workflows to the visible ChatGPT Chat and Work experiences in a signed in web session. It detects which surface it is on, inspects the controls that surface actually offers, and drives the user visible UI rather than any private endpoint.

Does codex-chatgpt-control take over your computer or run headless?

No. The stated scope is agent, then browser, then chatgpt.com, and the project says it is not a generic browser automation framework or a scraping tool. A real workflow still requires a compatible Codex or browser bridge and a visible signed in ChatGPT session, and the plugin bundles no browser bridge and no credentials.

How do you install codex-chatgpt-control for Node and for Python?

Node uses npm install codex-chatgpt-control@next and Python uses python -m pip install --pre codex-chatgpt-control. The npm prereleases are published under the next tag, so a plain npm install intentionally resolves to an older latest release instead of the current build.

What happens when ChatGPT shows a login or a captcha during a codex-chatgpt-control run?

The SDK is meant to report the exact stop reason rather than retry, covering the cases where ChatGPT needs login, a captcha, permissions or UI review. The skill instruction shipped with the plugin says to report the stop reason and not to retry blindly.

Does codex-chatgpt-control store my prompts and ChatGPT responses?

Local run reports omit prompt and response content by default, and the project describes them as privacy preserving local reports obtained without private endpoint access. Credentials are not bundled either, since the plugin ships no browser bridge, no credentials and no ChatGPT account access.

Is codex-chatgpt-control an official OpenAI project?

No. The README states that the project is not affiliated with, endorsed by, or sponsored by OpenAI, and it is described as an unofficial SDK that is not an OpenAI API wrapper and not a replacement for the official Codex SDK or CLI.

Official sources

  1. adamallcock/codex-chatgpt-control on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/adamallcock-codex-chatgpt-control.svg)](https://hysenlabs.com/projects/adamallcock-codex-chatgpt-control)