CLI tool
alibaba/open-code-review avatar
alibaba/open-code-review

Open Code Review moves file selection out of the model, and admits the recall cost

Alibaba Open Code Review combines deterministic checks with a language-model agent to produce line-level findings for issues such as null access, concurrency errors, XSS, and SQL injection.

41,252 stars2,965 forksGoApache-2.0

At a glance

What is it?
Open Code Review is a Go CLI that reviews git diffs with a hybrid design. Deterministic code picks the files and bundles them, and a sub-agent handles each bundle in an isolated context. The README states plainly that this buys precision at the cost of recall, and that is the trade on offer.
Who is it for?
Open Code Review fits a team with a high volume of small and medium pull requests where triage time is the bottleneck, since the design targets fewer false alarms and a fraction of the token spend of a general-purpose agent. Do not adopt it as the only review layer on a security-sensitive change, because the project states its recall is lower than a general-purpose agent's.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

npm installs a launcher that downloads a second binary

Installation is one command, and the prerequisite above it is a specific one: Git 2.41 or higher, because the tool leans on Git for diff generation, code search, and repository operations.

bash
npm install -g @alibaba-group/open-code-review

After that the `ocr` command is available globally. The next step is a model endpoint, and the README is unambiguous that you need one:

bash
ocr config

There is one exception, Delegation Mode, which is exempt from the endpoint requirement. What the README does not say is what Delegation Mode does, and that gap matters before you rely on it.

Now the part worth slowing down on. The npm manifest calls itself version `0.0.0`, its `files` array ships a JavaScript launcher at `bin/ocr.js` plus four scripts and an images directory, and `postinstall` runs `node scripts/install.js`. The real tool is a Go binary, and the install script fetches it from a GitHub release. So a global npm install is a registry download followed by a second network download from GitHub Releases, which is two supply chains for one command, and the second one runs a script.

Every platform package is pinned at 0.0.0

The manifest declares six optional dependencies, one per target, and every one of them is pinned to the same placeholder:

json
"ocrConfig": {
  "urlPattern": "https://github.com/alibaba/open-code-review/releases/download/v{version}/opencodereview-{os}-{arch}",
  "checksumPattern": "https://github.com/alibaba/open-code-review/releases/download/v{version}/sha256sum.txt"
}

`@alibaba-group/ocr-darwin-arm64`, `ocr-darwin-x64`, `ocr-linux-arm64`, `ocr-linux-x64`, `ocr-win32-arm64`, and `ocr-win32-x64` all read `0.0.0`, and so does the manifest itself. Meanwhile the `ocrConfig` block shows where the actual version comes from: a `{version}` placeholder substituted into a GitHub releases download path, alongside a checksum file.

Two consequences. The registry metadata cannot tell you which build you have, because the version it advertises is a placeholder rather than a release number, so your only version signal is what the running binary reports. And an optional dependency stuck at `0.0.0` fails quietly rather than loudly, since npm treats unmet optionals as acceptable, so a platform whose package is missing leaves you with a launcher that has nothing to launch. The `Makefile` closes the loop with a `sha256sum` target that produces the file the install script verifies against.

Bundling is what stops the agent from cutting corners

The design argument is that a purely language-driven architecture lacks hard constraints on the review process, and the project's answer is to move decisions out of the model. Four pieces of deterministic code do the work a general-purpose agent would otherwise improvise.

File selection decides exactly which files need review and which get filtered, so nothing is skipped by an agent losing interest in a large changeset. File bundling groups related files into a single review unit, with the README's own example being `message_en.properties` bundled with `message_zh.properties`, and each bundle runs as a sub-agent with isolated context. That is a divide-and-conquer split, and it is what makes concurrent review and stability on very large changesets possible. Rule matching is template-engine based rather than prompt based, so the rules that apply to a file are chosen by its characteristics. Positioning and reflection are separate modules, one for where a comment lands and one for what it says.

The model keeps only what is genuinely dynamic: scenario-tuned prompts, and a toolset distilled from tool-call traces in production data using call frequency, per-tool repetition rates, and the effect of adding a tool to the chain. The thesis is a deliberately narrower agent.

The tool states that its recall is lower than a general-purpose agent's

The benchmark section makes its case and then undercuts it in the same paragraph. Against Claude Code on the same underlying model, the README claims significantly higher precision and F1, completion faster, and token consumption of roughly a ninth. Then: its recall is lower than general-purpose agents, described as a deliberate trade-off favoring precision over noise.

The metric table explains what that costs. Precision is the proportion of reported issues that are real defects, so higher precision means fewer false alarms to triage. Recall is the proportion of real defects that are found, so lower recall means more issues slip through review. The tool is optimising the number your team reads and accepting a worse number for the one they do not.

That is defensible for a high-traffic repository where a noisy review gets ignored, and wrong if the review is the control on a security-sensitive change. The README offers no mode where you get both. A clean run is evidence that the findings were real, not that nothing is wrong.

The benchmark is the project's own dataset measured against Claude Code

The numbers come from AACR-Bench, published on Hugging Face as Alibaba-Aone/aacr-bench, and the README describes how it was built: 50 popular open-source repositories, 200 real pull requests, and 10 programming languages, cross-validated by more than 80 senior engineers who annotated 1,505 ground-truth issues.

That is serious work, and it is also work the project did. The comparison target is a general-purpose agent doing the same task through a Skill, which is a configuration rather than a fixed system, and the results are published on the project's own site. The ground truth comes from people labelling issues for this dataset, not from the defects your team will ship.

Read the table as a description of the design's intended behaviour, which it does support: a review tool that selects files deterministically and reports fewer speculative findings will score well on precision and poorly on recall, and that is exactly the row it wins. Read it as proof that the tool beats a general-purpose agent on your pull requests and there is a gap between the two claims.

ocr scan audits whole files on the model that misses the most

Beyond diff review there is a second command, `ocr scan`, which reviews entire files. The README frames it for auditing unfamiliar codebases or directories that have no meaningful diff, and that framing is exactly where the recall trade-off bites hardest.

Diff review is a bounded problem. The changeset names the files, the bundler groups them, and the sub-agent works through bundles it was given. A scan has no such bound, so the scope is whatever you point it at, and tokens per review is one of the metrics the benchmark table tracks. Scanning a directory is where that bill lands.

The second problem is that audit is the use case where missing a defect is most expensive. You are reading code you did not write, you have no prior context to catch what the model drops, and the mode is the one where the same precision-over-recall design applies with no diff to keep the review honest. If you scan, plan on treating the output as a first pass and reading the files it did not comment on.

Six CI integrations, and only the GitHub Action has contract tests

Comments have to land somewhere, so the repository ships integration examples for six forges: `bitbucket_pipelines/`, `codeup_ci/`, `gerrit_ci/`, `gitflic_ci/`, `github_actions/`, and `gitlab_ci/` under `examples/`, plus an `action.yml` at the root for the GitHub Action path. The two Codeup and GitFlic directories are Alibaba's own forges, which fits a tool that started as an internal assistant there.

Only the GitHub integration appears to be tested. The manifest has a `test:github-actions` script that runs four files in order: `post-review-comments.test.js`, `check-translation-sync.test.js`, `action-contract.test.js`, and `check-plugin-contract.test.js`. A contract test on the action and a check that translations are in sync are maintenance commitments for that one path.

Nothing in that script name covers the other five. For a Gerrit or GitLab setup the example directory is the whole of the documentation in this repository, so expect to be reading a sample and adapting it rather than installing something tested. The repository also carries four agent-integration roots at the top level, `.claude/`, `.claude-plugin/`, `.kimi-plugin/`, and `.agents/`, plus `skills/` and `plugins/`, so the tool is meant to run both as a standalone CLI and inside a coding agent.

go 1.25.5 and node >=14 guard the same binary

The launcher and the tool it launches have toolchains a decade apart. The manifest asks for `node >= 14`, and `go.mod` opens with:

go
module github.com/alibaba/open-code-review

go 1.25.5

That Go directive names a patch release, not a major version, so building from source needs 1.25.5 or newer and an older toolchain will reach for a download. Add Git 2.41 to the list and a CI image has to satisfy three floors at once.

The `Makefile` shows how the version reaches the binary. `GIT_TAG` comes from `git describe --tags --abbrev=0` with a fallback, `GIT_COMMIT` from `git rev-parse --short HEAD`, and `VERSION` defaults to `v0.0.0-$(GIT_COMMIT)` when no tag is present. Those three values are injected through `-ldflags` into `main.Version`, `main.GitCommit`, and `main.BuildDate`, with `-s -w` added for releases. So an untagged local build reports a `v0.0.0-` version by design, and that is a different string from the `0.0.0` in the npm manifest and from a `v1.12.x` release tag.

Every platform target builds with `CGO_ENABLED=0`, which is why six static binaries can be dropped into `dist/` and fetched by the installer without a matching libc.

Editorial conclusion

Open Code Review fits a team with a high volume of small and medium pull requests where triage time is the bottleneck, since the design targets fewer false alarms and a fraction of the token spend of a general-purpose agent. Do not adopt it as the only review layer on a security-sensitive change, because the project states its recall is lower than a general-purpose agent's. Verify three things before wiring it into CI. Which build you actually installed, since the npm manifest and all six platform packages read 0.0.0 and the real number comes from a GitHub release tag. That your environment permits the postinstall download, which fetches a second artifact and checks it against sha256sum.txt. And that your forge is covered, since GitHub has action.yml and contract tests while Gerrit, GitLab, Bitbucket, Codeup, and GitFlic exist only as examples.

Frequently asked questions

how to use opencode review

Install with `npm install -g @alibaba-group/open-code-review`, which makes the `ocr` command available globally, then run `ocr config` to point it at an LLM endpoint. Git 2.41 or higher is required. The tool reads git diffs and posts structured review comments with line-level precision.

how to use opencode review command

The two commands named in the documentation are `ocr config`, which configures the model endpoint, and `ocr scan`, which reviews entire files rather than a diff for auditing unfamiliar code. A configured LLM is required before reviewing code unless you use Delegation Mode.

Does Open Code Review replace a general-purpose coding agent for review?

The project's own benchmark does not claim that. Against Claude Code on the same underlying model it reports higher precision and F1 with roughly a ninth of the tokens, while stating that its recall is lower, which it calls a deliberate trade-off favoring precision over noise.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/alibaba-open-code-review.svg)](https://hysenlabs.com/projects/alibaba-open-code-review)