bug-hunter's documented version is 3.2.0, its perfect scores come from its own fixture, and droid installs globally
Adversarial AI bug hunter with auto-fix skill for Claude Code, Cursor, Codex CLI, GitHub Copilot CLI, Kiro CLI, Opencode, Pi Coding Agent, and more. Multi-agent pipeline finds security vulnerabilities, logic errors, and runtime bugs — then fixes them autonomously on a safe branch.
At a glance
- What is it?
- codexstar69/bug-hunter is a multi-agent auditing skill for coding agents: a Hunter proposes bugs, a Skeptic challenges every claim, and a Referee decides what the evidence supports, with scan-only single-pass as the default and five separate flags between you and an edit. The pipeline design is careful, and so are the disclosures, because the README tells you its numbers validate its own harness and tells you to install an untagged branch tarball rather than a release.
- Who is it for?
- bug-hunter is worth installing if you want an independent second opinion on a codebase and can accept that its strongest property is refusal rather than accuracy: the default run reports and stops, and every path to editing a file sits behind a flag you have to type.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 50 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The page documents 3.2.0 and there is no 3.2.0 tag
Start with the version, because this project is upfront about the gap and still recommends the unpinned path.
The README describes what it calls the current 3.2.0 source and says in plain terms that the latest published npm release may lag the main branch, and that you should use the current-source command when you need the exact implementation documented on the page. The package manifest also reads 3.2.0.
The release list does not contain it. The three most recent tags are 3.1.0 from 2026-08-03, 3.0.8 from 2026-03-13, and 3.0.7 from 2026-03-12. So the build the README describes exists on the branch and in the manifest but has never been tagged, and the last tagged release predates it by a version number.
That is not hidden, and the project gives you both routes. The recommended one installs the branch tip directly.
npx --yes https://github.com/codexstar69/bug-hunter/archive/refs/heads/main.tar.gz install --agent codex
npx --yes https://github.com/codexstar69/bug-hunter/archive/refs/heads/main.tar.gz doctor --agent codexNote what that command is. It is npx pointed at a tarball URL on the branch head, so it downloads whatever the branch contains at the moment you run it and executes its installer, with no version in the string at all. The alternative pins a published package instead.
npm exec --yes --package=@codexstar/bug-hunter@latest -- bug-hunter install --agent codex
npm exec --yes --package=@codexstar/bug-hunter@latest -- bug-hunter doctor --agent codexSo the choice is an unpinned branch tip or a published version that the documentation does not describe. A separate doctor command follows installation in both cases, and you are told to restart the agent if it was open while you installed.
A perfect score on a fixture that ships with the repository
This is the section to read before you quote any of this project's accuracy claims, including to yourself.
There is a bundled deterministic regression fixture, and its numbers are recorded in the README. Precision of 1.00, recall of 1.00, an F1 of 1.00, repeat stability of 1.00, zero false positives, a median of 12,090 tokens per true positive, a 95th-percentile duration of 61.3 seconds, and an expected calibration error of about 0.048.
Every one of those is a perfect or near-perfect score. On a benchmark. That ships with the tool, in a fixture directory at the repository root.
The README then says the thing that matters, in its own words: these figures validate the bundled harness and fixture, and they are not an independent benchmark of every repository or model. That sentence is the whole caveat, and it is correctly placed.
So there is no claim here that the tool is one hundred percent accurate, and the author is not making one. What the numbers do demonstrate is that the harness works: given known defects, the three-role pipeline finds them, the Skeptic does not discard them, and the Referee upholds them, and it does so repeatably and within a stated token and time budget.
What the numbers cannot tell you is the false-positive rate on your code, and false positives are the failure mode that makes people stop running tools like this. The project does publish the metric that would matter here, false positives per thousand lines, and it lists calibration and repeat stability as part of the quality gate. It just does not have an independent number for them.
For that, the machine-readable results file and the protocol document the README links are where to look.
Twelve stages, three roles, and seven schemas that name them
The pipeline is printed in full, which makes it easy to see exactly where a human can stop it.
your code
-> risk triage
-> optional adaptive plan
-> architecture recon
-> hypothesis-driven retrieval
-> Hunter findings
-> documentation checks
-> Skeptic challenges
-> Referee verdicts
-> optional hybrid verification
-> report
-> optional approved fix plan
-> optional approved fixes and verificationThree of those stages are optional by construction: the plan, the verification, and both fix stages. Everything up to the report runs unconditionally. So the floor is: triage your code, look at it, find bugs, check the docs, have someone argue with the findings, have someone else rule on the argument, write a report.
The separation of roles is the design. A Hunter proposes, a Skeptic is required to challenge each claim rather than confirm it, and a Referee decides what the evidence supports. Two agents looking at the same finding is not a consensus mechanism; it is a debate with an adjudicator, which is the cheapest structure that resists a single confident hallucination.
The schemas line up with the stages one for one. The published package ships schema files for coverage, experiments, findings, fix plans, fix reports, fix strategies, and fixer scope. Coverage maps to the loop, findings to the Hunter's output, and the three fix schemas plus the fixer-scope schema map onto the three fix stages. The README asks for canonical JSON artifacts throughout, which is what makes a scan auditable after the fact rather than only at the time.
One stage is worth naming for what it rules out. Hybrid verification runs tests, type checks, static checks, fuzz checks, and security static checks, and the constraint is argv-only invocation, so no shell string is ever assembled. It also fails closed: if a required check cannot run, the run fails rather than reporting success.
Evidence is reused only on an exact match, which is the anti-staleness rule
If there is one engineering decision here worth copying into your own tooling, it is this one.
The cache reuses evidence only when four things match exactly: the source content, the protocol identity, the role, and the relevant configuration. The stated consequence is that changed source cannot inherit stale conclusions.
That matters more for this class of tool than for a linter. A linter is a pure function of the file. An auditing agent produces a conclusion, and the natural implementation is to store it so the next run is cheaper. The natural bug is that the conclusion outlives the code it was about, and a human reads a confident finding that was true six commits ago.
The source-integrity work is the same idea pushed further. Scan scope, source hashes, resume identity, coverage state, and fixer authorization are all kept explicit, and the project says it rejects source drift. Resume identity is the subtle one: it is what stops a resumed run from silently auditing a different tree than the one it was started on.
The retrieval side is bounded rather than open-ended. Hypothesis-driven retrieval prioritises direct files, symbols, dependencies, dependents, cross-references, and trust boundaries under hard context budgets, and the pipeline is described as keeping adaptive and retrieval context inside explicit file and token budgets. Combined with triage ranking high-risk files before low-risk ones, that is an admission that a full-repository audit is not affordable in context and the design optimises for where it is spent.
Three execution profiles exist, selected automatically from triage risk, security scope, benchmark evidence, stability, calibration, and token efficiency: fast, balanced, and assurance. The same token budget in the quality gate, 12,090 median tokens per true positive, is what a profile is trying to manage.
The droid target installs into a directory that loads everywhere
Nine agent targets are supported, and one of them has a default you will want to change.
The mapping is explicit: Claude Code, Codex, Cursor, GitHub Copilot, Kiro, Windsurf, OpenCode, Factory Droid CLI, and a generic file-based target for everything else. The advice is to always pass the agent flag when more than one coding agent is installed, because auto-detection is available but an explicit target is what prevents the skill being written into the wrong directory. That is a real failure mode when several agents share a machine.
Factory Droid is the one that needs reading twice. Its default target writes the skill into a directory under the user's home and loads it for every repository on the machine. A path option restricts the install to a single repository instead, and that is the form to use unless you genuinely want the skill available everywhere.
There is also a legacy overlap. Droid reads an older shared skills location as well, which means an install made with the generic file-based target already works there, so a user who installed once with the generic target and later added Droid has a working setup without knowing it.
The installation documentation is delegated to a separate page covering paths, source installs, updates, removal, and custom targets, and the README says nothing further about how to uninstall.
Once installed, the interface is natural language, which is the right choice for something meant to work inside nine different agents. There is also a slash command form for agents that expose one, and the recommended first run is explicit about what it will not do.
A CommonJS script package with twenty-five scripts enumerated by name
The package manifest describes something closer to a shell program than a library, and the file list is where that shows.
The package is marked CommonJS, its entry point and its command are the same file with no extension, and it requires Node 22 or newer. The permanent continuous-integration gates verify Node 22 and Node 24, with the 24 lane also running the benchmark quality gate and a package inventory check. So Node 22 is the floor and 23 is untested.
The files array is the interesting part. Instead of shipping a directory, it enumerates the bin directory and then names roughly twenty-five individual scripts by filename, followed by seven or more schema files, also named individually. Everything the tool runs is CommonJS with an explicit extension.
That is a deliberate choice with a real cost. Because the list is maintained by hand, a script added to the source tree and not added to the list is missing from the published package, and nothing fails until the moment a code path reaches it. That is exactly what the package inventory check in CI appears to exist to catch.
The script names also encode the safety model more clearly than the prose does. There is a payload guard, a fix lock, a process runner, a worktree harvester, a state store, a triage module, and an experiment loop. Worktree harvesting implies fixes land somewhere isolated rather than in your working tree, and the fix lock implies two runs cannot fix at once.
The repository root carries two loop directories, one for current state and one for history, both committed. So a scan leaves state in the repository you audited, and that state is version-controlled alongside your code. The root also holds two machine-readable documentation indexes in short and long form, plus an evaluations directory, a plans directory, a modes directory, and a prompts directory.
Editorial conclusion
bug-hunter is worth installing if you want an independent second opinion on a codebase and can accept that its strongest property is refusal rather than accuracy: the default run reports and stops, and every path to editing a file sits behind a flag you have to type. It is not worth installing if you want measured accuracy, because the only figures it publishes come from a fixture committed in the repository, and the README is explicit that they validate the harness rather than any codebase or model. Three things to settle first. The version documented on the page is 3.2.0, there is no 3.2.0 tag among the releases, and the recommended install fetches an untagged branch tarball through npx, so the artefact the documentation tells you to trust is neither a release nor a pin. The droid target installs into a global skills directory that loads for every repository on the machine, so use the per-repository path form unless you want it everywhere. And the pipeline writes its loop state into the repository, with the loop and loop-history directories committed alongside the source.
Frequently asked questions
What does bug-hunter do by default?
It scans and reports, single-pass, without editing. A Hunter proposes possible bugs, a Skeptic is required to challenge each claim, and a Referee decides what the evidence supports. Loop coverage over the full queue requires the loop flag, and editing, autonomous fixing, and commits each require explicit permission.
How do I install bug-hunter for my coding agent?
Pass the target explicitly with an agent flag. The supported targets are Claude Code, Codex, Cursor, GitHub Copilot, Kiro, Windsurf, OpenCode, Factory Droid CLI, and a generic file-based target. The README advises always passing it when more than one coding agent is installed, since auto-detection can write the skill into the wrong directory.
How do I stop bug-hunter from changing any files?
Use the plan or preview flags, which the README says make source edits impossible. For a reviewed run, the fix flag combined with approve asks the host agent's reviewed permission mode and the host decides when prompts appear. The autonomous and auto-commit flags should not be used unless you intend to grant those permissions.
What accuracy numbers does bug-hunter publish?
Only from its bundled deterministic regression fixture: precision 1.00, recall 1.00, F1 1.00, repeat stability 1.00, zero false positives, a median of 12,090 tokens per true positive, a 95th-percentile duration of 61.3 seconds, and an expected calibration error near 0.048. The README states these validate the harness and fixture rather than any repository or model.
Which Node.js version does bug-hunter require?
Node 22 or newer. The package manifest declares that engine floor, and the permanent continuous-integration quality gates verify Node.js 22 and Node 24, with the 24 lane additionally running the benchmark quality gate and the package inventory check.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/codexstar69-bug-hunter)