# Codex Security: OpenAI's CLI and TypeScript SDK for Vulnerability Scanning

> Codex Security pairs a CLI with a TypeScript SDK to find, validate and fix vulnerabilities, and adds a preview findings service for deduplication and severity classification. It is early, opinionated about Node and Python versions, and needs an inference provider.

**openai/codex-security** — OpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security

- Repository: https://github.com/openai/codex-security
- Website: https://developers.openai.com/codex/security
- Stars: 10,874 · Forks: 825
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/openai-codex-security

## What Codex Security is for, and who it is aimed at

Codex Security is a CLI and a TypeScript SDK for defining security policy and for finding, validating and fixing security vulnerabilities in a codebase. The README states that purpose in one sentence, and the repository topics (ai-security, devsecops, code-scanning, vulnerability-scanning) point at the same audience: application security engineers and platform teams who want scanning wired into scripts and CI rather than into a hosted dashboard.

The shape of the product matters more than the description. Two entry points exist. `codex-security` is the command line tool; `@openai/codex-security` is the npm package you import in JavaScript or TypeScript. Around them sit a Docker image, several Compose files, and a findings service that the README marks as preview. That is a lot of surface for a package whose latest published version is 0.1.26, released on 2026-09-08.

There is also a gate that has nothing to do with code: the README says some cybersecurity requests and protected findings require approval through Trusted Access for Cyber, with a sign-up page at chatgpt.com/cyber. If your work touches that category, the tool's usefulness depends on an approval process you do not control.

## How the scanner, SDK and findings service fit together

The primary data flow is short. The CLI or SDK takes a directory, runs a scan, and produces a report; the SDK exposes `result.reportPath`, and the README's example prints it. Everything else in the repository is built around that core.

Policy comes first in the documented workflow. `codex-security policy .` drafts repository-wide `SECURITY.md` guidance for future scans, and `codex-security policy . --path services/api --knowledge-base architecture.md` scopes the draft to a component while feeding it supporting documents. The README is explicit that the command saves a draft outside the checkout, does not install it, and does not run a scan; you review the proposed diff and copy the policy yourself. Supporting architecture, threat-model and review documents are expected to stay outside the repository because they may contain sensitive details.

Then there is the findings service. `codex-security serve` starts it without Docker. It stores findings and embeddings in SQLite, lists findings with pagination, and serves a read-only dashboard at `/dashboard` that refreshes every five seconds, showing stored findings and duplicate groups. Duplicate detection works by embedding similarity, scoped either to one repository or, with `--all-repositories`, to everything in the database. `codex-security publish scan --to custom --findings-url http://localhost:3000` uploads completed findings along with a repository ID. The `codex-security dedupe` command and the SDK retrieve candidate duplicates, run independent Codex reviews locally, and persist the groups that survive review.

Severity classification is a separate pass. `codex-security classify-severity --scan SCAN_ID --rubric /path/to/policy.md` assesses selected findings against your own policy before tickets are published. The README notes that classification checkpoints each finding in SQLite and reuses matching assessments on reruns unless `--reprocess` is passed, and that the original scan severity is left unchanged. That separation is sensible: it means a rubric change does not silently rewrite history, but it also means two severity numbers can exist for the same finding and someone has to decide which one triage reads.

## Installing Codex Security and running a first scan

The README states two runtime requirements: Node.js 22.13.0 or later and Python 3.10 or later. The Dockerfile confirms the Python dependency, installing `python3` in both the package and scanner stages. If your build image is Node-only, the scan will not work as documented.

Installation is a single npm package, followed by a login and a scan against a directory. The README gives exactly this sequence:

```bash
npm install @openai/codex-security
codex-security login
codex-security scan /path/to/directory
```

After `login` completes, the scan prints progress and finishes with a report; the SDK example shows the report location being read as `result.reportPath`. For CI, the README says to set `OPENAI_API_KEY` instead of signing in, so an unattended job needs that variable in the environment rather than an interactive session.

Scripting is the second path. The SDK example instantiates the client, runs a default scan, then runs a second scan with explicit limits, and closes the client:

```ts
import { CodexSecurity } from "@openai/codex-security";

const security = new CodexSecurity();
const result = await security.run("/path/to/directory");
await security.run("/path/to/directory", {
  mode: "deep",
  workers: 2,
  subagents: 0,
  stopAfterNoNew: 3,
  maxDiscoveryRuns: 10,
  maxTimeHours: 1.5,
});

console.log(result.reportPath);
await security.close();
```

Those option names are worth reading closely: `stopAfterNoNew`, `maxDiscoveryRuns` and `maxTimeHours` are all stopping conditions. A deep scan is bounded by configuration rather than by the size of the repository, which is the honest way to expose a model-driven process, but it also means a large codebase can return partial coverage when the budget runs out.

For many repositories at once, the repository ships a Compose configuration. `compose.yaml` mounts `./repositories.csv` read-only at `/input/repositories.csv`, `./results` at `/output`, and `./state` at `/state`, then runs `bulk-scan /input/repositories.csv --output-dir /output`. The container runs as user `10001:10001`, drops all capabilities, sets `no-new-privileges:true`, and applies a seccomp profile from `./docker/codex-security-seccomp.json`. The environment block passes `CODEX_API_KEY`, `CODEX_SECURITY_GIT_HOST`, `GH_TOKEN`, `GITHUB_TOKEN` and `OPENAI_API_KEY` through from the host. `examples/findings.csv` shows the expected findings format, and `.env.example` is where you put `OPENAI_API_KEY` for the findings service; it also documents `CODEX_SECURITY_EMBEDDINGS_URL` as an optional full embeddings endpoint URL, unset by default.

## The model provider is a choice, and so is the bill

Codex Security does not lock you to one inference backend. The README documents three alternatives, each selected with `--provider` and `--model`:

```bash
export AWS_BEARER_TOKEN_BEDROCK="<your-bedrock-api-key>"
export AWS_REGION="us-east-2"
codex-security scan . --provider amazon-bedrock --model openai.gpt-5.6-luna

export OPENROUTER_API_KEY="<your-openrouter-api-key>"
codex-security scan . --provider openrouter --model anthropic/claude-sonnet-4.5

export FIREWORKS_API_KEY="<your-fireworks-api-key>"
codex-security scan . --provider fireworks --model accounts/fireworks/models/qwen3-235b-a22b
```

This is a real architectural decision, not a checkbox. Findings, duplicate review and severity classification all depend on a model's judgement, and the model name is passed on the command line, so two people running the same scan against different models can get different results from the same commit. The README does not describe any determinism guarantee, and it does not document a way to pin a model version inside the policy file. The `docs/project-configuration.md` file is referenced for reusable YAML and JSON settings, CLI overrides and editor schema support, which is where a team would look to standardise the provider, but the README itself does not spell out that the provider can be fixed in configuration.

Cost is the other half. The README never states pricing, and the repository contains no quota or budget mechanism beyond the scan's own time and run limits. With a hosted provider you are paying per token for discovery, for duplicate review and for severity classification, and the deep mode with multiple workers multiplies that. Budgeting has to happen at the provider account level.

## Where Codex Security is the wrong tool

The clearest limitation is that this is not a deterministic analyser. A rule-based scanner produces the same findings for the same input every time, which is what makes it usable as a merge gate with a fixed pass or fail condition. Codex Security produces a report from a model-driven discovery process, with stopping conditions like `stopAfterNoNew` and `maxDiscoveryRuns` controlling when it ends. That makes it well suited to triage and to surfacing issues a human then confirms, and poorly suited to a hard gate that blocks a pull request on an exact finding count.

The findings service is labelled preview in the README, and the surrounding details reinforce that reading. It stores findings and embeddings in SQLite, which is fine for a single container with a state volume and awkward for anything that needs concurrent writers or high availability. The dashboard is read-only and refreshes every five seconds; it is an inspection surface, not a workflow tool. There is no mention of authentication on the service, and the publish command points at `http://localhost:3000` in the example, so exposing it beyond a trusted network is something the README does not address.

Policy generation has its own trap. `codex-security policy .` writes a draft outside the checkout and does not install it. A team that assumes the command has configured the repository will find that nothing changed. The README also warns that architecture, threat-model and review documents stay outside the repository and may contain sensitive details, which is a deliberate constraint: you cannot simply commit the context that makes the policy good.

Finally, the Trusted Access for Cyber approval means some requests and protected findings are gated. For an open source maintainer or a small team that cannot join the programme, part of the tool's range may be unavailable, and the README does not describe what the experience looks like when a request falls into that category.

## How it compares with a conventional SAST tool

Semgrep is the natural reference point for anyone evaluating Codex Security, and the difference is in where the intelligence sits. Semgrep executes rules you write or import: the rule is the unit of behaviour, results are reproducible, and adding a language or a pattern means writing or selecting a rule. Codex Security inverts that. The unit of behaviour is a model call, the policy is prose guidance in a `SECURITY.md` draft plus a rubric file passed to `classify-severity`, and the results depend on which provider and model you selected with `--provider` and `--model`.

That inversion has consequences in both directions. Semgrep cannot tell you that a pattern is exploitable in your specific architecture, and it cannot rewrite the code. Codex Security is described as finding, validating and fixing vulnerabilities, and its dedupe step runs independent Codex reviews to decide whether two findings are the same issue. Those are tasks a rule engine does not attempt. But Semgrep will not surprise you on a rerun, and its output can be diffed across commits without asking whether the model changed underneath.

A pragmatic reading: they answer different questions. A rule engine answers "does this code match a known dangerous pattern". Codex Security answers "what looks exploitable here, given this policy". Teams with an existing SAST gate usually want the second as an additional signal, not as the thing that turns the build red.

## Licence, maintenance and upgrade cost

The repository is licensed Apache-2.0. That is a permissive licence with an explicit patent grant, and it imposes no copyleft obligation on code you write around the CLI or SDK. It does not change the fact that scanning sends your source to whichever inference provider you configure, so the licence question and the data-handling question are separate; the README does not describe what is transmitted or retained, and the provider is your choice, so that review belongs with your provider's terms rather than with this repository.

The maintenance signal is strong on recency. The last push was on 2026-09-14, and the release history shows npm-v0.1.24 on 2026-08-29, npm-v0.1.25 on 2026-09-02 and npm-v0.1.26 on 2026-09-08. The repository is not archived. Three releases in eleven days at the 0.1.x line is a fast cadence, and it is the practical upgrade cost: pin the version. Because the CLI, the SDK, the Docker image and the findings service share a version line, a partial upgrade can leave a runner talking to a service from a different release. The `CHANGELOG.md` and `RELEASING.md` files at the repository root are where the project's own notes on that live.

Running the scanner as a container adds a smaller, recurring cost. The Compose setup depends on a seccomp profile at `./docker/codex-security-seccomp.json` and a fixed non-root user, so any host that does not permit custom seccomp profiles needs that configuration changed before the first scan, not after.

## Conclusion

Adopt Codex Security if your team already runs Node 22.13.0 or later and Python 3.10 or later and wants scans driven from a CLI, a TypeScript SDK, or a Docker Compose bulk run over a CSV of repositories. Do not adopt it as a drop-in replacement for a deterministic SAST gate: the scanner's judgement comes from a model, the findings service is explicitly labelled preview, and the policy command writes a draft outside the checkout rather than installing it. Before rolling it out, verify three things: that your inference provider and model name work with `codex-security scan . --provider ... --model ...`, that the seccomp profile at `./docker/codex-security-seccomp.json` runs under your container runtime, and that `codex-security classify-severity --rubric /path/to/policy.md` reproduces the severities your triage process expects.

## FAQ

### What is Codex Security used for?

It is a CLI and TypeScript SDK for defining security policy and for finding, validating and fixing security vulnerabilities in code. The README also documents a policy command that drafts SECURITY.md guidance and a preview findings service that stores findings in SQLite.

### How do I install Codex Security?

Install the npm package, then log in and point the scanner at a directory. The README states the requirements as Node.js 22.13.0 or later and Python 3.10 or later, and notes that CI should set OPENAI_API_KEY instead of signing in.

### How do I use Codex Security in a project?

The README's quick start runs `codex-security login` and then `codex-security scan /path/to/directory`. From TypeScript you import CodexSecurity from @openai/codex-security, call run on a path, read result.reportPath, and call close when finished.

### Is Codex Security free?

The repository does not state pricing. The scanner uses an inference provider, and the README documents OpenAI as well as amazon-bedrock, openrouter and fireworks, each with its own API key, so any cost sits with the provider you configure rather than with the Apache-2.0 package.

### What is the Codex Security plugin?

The repository contains a plugins directory with a codex-security MCP app, whose package.json and lockfiles are copied into the Docker build. The README does not describe the plugin's commands or how to install it, so its behaviour is not documented there.

## Sources

- [License: Apache-2.0](https://github.com/openai/codex-security/blob/main/LICENSE)
- [openai/codex-security on GitHub](https://github.com/openai/codex-security)
- [Project website](https://developers.openai.com/codex/security)
- [README](https://github.com/openai/codex-security/blob/main/README.md)
- [Releases](https://github.com/openai/codex-security/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openai-codex-security
