PR-AF: Agentic Code Review with Evidence Grounding and Compound Risk Synthesis
#1 open-source code reviewer on Code-Review-Bench
At a glance
- What is it?
- PR-AF is a self-hosted agentic code reviewer built on the AgentField platform that dynamically compiles review plans from pull request topology, spawns specialized reviewer agents, grounds findings in code evidence, and gates them against falsifiability checks before posting inline GitHub comments.
- Who is it for?
- PR-AF is appropriate for CI/CD gates on high-stakes pull requests where architectural depth and compound risk detection matter more than speed. Its 35-to-50 minute execution time makes it poorly suited for interactive inner-loop development or small formatting PRs.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Benchmark Position and What It Measures
The README leads with a specific claim: on 38 runnable pull requests from the Martian Code-Review-Bench evaluation, PR-AF using the GLM-5.2 model achieves a golden recall score of 0.706 across 42 compared tools. Golden recall measures how many of the predefined bugs and issues a reviewer correctly identifies. This places PR-AF ahead of every comparable open-source reviewer and ahead of named commercial products including variants of qodo, CodeRabbit, Greptile, GitHub Copilot, and Devin in that snapshot.
The README also reports 595 independently valid findings on the benchmark set, described as roughly three times more than the leading commercial tools in an adjusted comparison. With Opus-class frontier models instead of GLM-5.2, the README states the margin widens further.
The benchmark results are in the repository at benchmark/martian-code-review-bench, with per-PR judge verdicts and reproduction scripts. This makes the claims checkable, which distinguishes PR-AF from tools that describe their accuracy in general terms without supporting data.
Dynamic Pipeline Architecture: How PR-AF Structures a Review
Most automated code review tools run the same fixed analysis on every pull request. PR-AF does not. When a PR arrives, the system performs what the README calls a topological anatomy mapping: it examines the structure of the diff (which files changed, how they relate, what kinds of changes they are) and uses that structure to compile review dimensions dynamically.
From those dimensions, the system spawns ephemeral reviewer agents tailored to the specific PR. A PR that modifies authentication code gets different agents than one that refactors a data pipeline. The agents run in parallel across semantic, mechanical, and systemic lenses.
After the agents generate findings, three verification steps run before any finding reaches a GitHub comment. Evidence Grounding pulls exact caller snippets and import chains from the repository to confirm that each finding is grounded in actual code, not an assumption about what might exist. Compound Vulnerability Synthesis clusters related findings across files to identify whether isolated issues combine into a larger systemic risk. Falsifiability Gates then attempt to invalidate each finding: if the behavior is safe, intentional, or already mitigated, the finding is discarded. Only findings that survive all three steps become review comments.
Installation and First Review
The quickest start is Docker Compose. Clone the repository, create a .env file from .env.example, fill in at minimum OPENROUTER_API_KEY and GH_TOKEN, then start the services:
docker compose up --buildThis starts two containers: the AgentField control plane on port 8080 and the pr-af agent on port 8004.
To trigger a review using the af CLI (version 0.1.87 or later):
af call pr-af.review --in '{"pr_url": "https://github.com/owner/repo/pull/123"}'Alternatively, use curl directly against the HTTP API:
curl -X POST http://localhost:8080/api/v1/execute/async/pr-af.review \
-H "Content-Type: application/json" \
-d '{"input": {"pr_url": "https://github.com/owner/repo/pull/123"}}'The .env.example lists the key configuration variables. PR_AF_MODEL defaults to deepseek/deepseek-v4-flash-0731 for routine reviews. PR_AF_MAX_COST_USD caps the spend per review. PR_AF_MAX_DURATION_SECONDS sets a timeout. GH_TOKEN is required to clone private repositories and post inline comments.
Model Selection and Cost Profile
PR-AF is model-flexible. The README describes three tiers of use: DeepSeek-class models (such as the default deepseek/deepseek-v4-flash-0731) for routine PRs where cost matters more than depth, GLM-5.2 for open-model reviews where the benchmark score was achieved, and Opus-class frontier models for major PRs where the README states PR-AF tops the benchmark by a wide margin.
The cost comparison in the README puts PR-AF at approximately 10 times cheaper per review than closed-source tools. That figure is for the BYOK (bring your own key) model where you pay the API provider directly, with no per-seat subscription. The trade-off is that Opus-class reviews at high PR volume carry non-trivial API costs, and the PR_AF_MAX_COST_USD limit in .env.example exists precisely to prevent runaway spend on unexpectedly large diffs.
The execution time for a typical review is 35 to 50 minutes according to the README, which contrasts with 2 to 5 minutes for commercial SaaS tools and seconds to minutes for Claude Code CLI reviews. This makes PR-AF better suited to a nightly CI gate or a final merge gate than to a check that runs on every commit push.
Comparing PR-AF to Claude Code CLI and Commercial Reviewers
The README's own comparison table is specific about trade-offs. Claude Code CLI is an interactive inner-loop tool: fast, iterative, useful during development, but limited to a single-threaded approach and focused on the diff context window. Commercial SaaS tools like CodeRabbit offer a cleaner GitHub interface and faster turnaround but use context retrieval plus LLM review rather than parallel cognitive agents.
PR-AF's comparative strengths are compound risk detection (the dedicated Compound Vulnerability Synthesizer clusters related findings across files, which single-pass reviewers miss) and extremely low false positives from the evidence grounding and falsifiability gate steps. The README explicitly recommends using Claude Code for local development and PR-AF as a final GitHub Actions gatekeeper, framing them as complementary rather than competing.
The main practical constraint beyond speed is the self-hosted deployment: PR-AF requires a running AgentField control plane, Docker, and a machine with network access to GitHub and the chosen LLM API. There is no hosted version of PR-AF, so the infrastructure burden falls on the team deploying it.
Maintenance and License
The last push to the repository was on 2026-09-21. The repository is not archived. The project builds on the AgentField framework and requires the agentfield package at version 0.1.130 or later, hax-sdk at 0.2.4 or later, and Python 3.11 or later according to pyproject.toml.
The pyproject.toml declares the license as Apache-2.0. This permits commercial use, modification, and distribution, with the requirement to preserve notices and the Apache license text in distributions. The root directory lists a LICENSE file in the top-level entries, though the README itself does not recite the full license terms.
Railway deployment is supported via a railway.toml configuration file in the root. The README includes a one-click Railway deploy badge, which is useful for teams that want a cloud deployment without managing Docker infrastructure themselves. The docker-compose.yml route remains the recommended path for fully self-hosted setups.
Editorial conclusion
PR-AF is appropriate for CI/CD gates on high-stakes pull requests where architectural depth and compound risk detection matter more than speed. Its 35-to-50 minute execution time makes it poorly suited for interactive inner-loop development or small formatting PRs. The self-hosted model means API costs come from your own keys (OpenRouter or direct provider), and the OPENROUTER_API_KEY and GH_TOKEN values in .env are required before the first review runs. Confirm that the target repository is accessible from the deployed environment and that PR_AF_MAX_COST_USD is set to a value that fits your budget before triggering a large PR.
Frequently asked questions
What is GitHub PR Review Agent?
A GitHub PR review agent is a tool that automatically analyzes a pull request's diff and posts inline review comments on GitHub, identifying bugs, security issues, and design problems without a human reviewer reading the code first. PR-AF is an open-source example built on the AgentField framework that uses parallel specialized agents and evidence grounding to reduce false positives.
What models does PR-AF support for code reviews?
PR-AF supports any model available through OpenRouter or directly configured via OPENROUTER_API_KEY. The .env.example defaults to deepseek/deepseek-v4-flash-0731 for cost-efficient routine reviews. The benchmark result of 0.706 golden recall was achieved with GLM-5.2. The README states Opus-class frontier models (such as Claude Opus) produce the highest review quality.
How long does a PR-AF code review take?
According to the README, PR-AF takes approximately 35 to 50 minutes per review. This is significantly longer than commercial SaaS tools (2 to 5 minutes) because PR-AF runs a multi-phase pipeline with parallel agents, evidence grounding, and falsifiability checks. The PR_AF_MAX_DURATION_SECONDS environment variable in .env.example sets a hard timeout.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/agent-field-pr-af)