Model or dataset
adshao/flounder avatar
adshao/flounder

Flounder is a white-hat auditor whose scope boundary is prose, not an allow-list

Autonomous white-hat security auditor for AI-driven code review, bug bounty research, exploit construction, and execution-grounded verification.

517 stars75 forksTypeScriptAGPL-3.0

At a glance

What is it?
Flounder wraps a coding agent in an audit workflow: seven tracked stages, a containerised sandbox, and a rule that a finding is only real once a local command exercising the vulnerable path has passed. Its bounding mechanisms are network and execution level, which is a real control, but the authorization boundary itself is a sentence in the readme, and the accepted inputs include material from other people's disclosures.
Who is it for?
Separate the two claims before you use this. The engineering claim is strong and checkable: model-written code runs in a copied workspace inside a container with capabilities dropped, a read-only root and network switched off, discovery never has egress at all, and confirmation only opens the network for a narrow class of read and fetch commands.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The boundary is a sentence, and two of the inputs are not permission

Install it once and the agent picks it up from ordinary requests. The single documented command runs through the skills installer, asks for the skill by name, adds it globally, and registers it with two named coding agents at once:

bash
npx skills add adshao/flounder --skill flounder -g -a codex -a claude-code

The readme then opens by saying to hand it a public source or an authorized target boundary, and then lists what that boundary can be: a repository, a source tree, a package, a deployed lead, or a previous run. Two of those five are not authorizations. A deployed lead is something you noticed; a previous run is somebody else's disclosure or somebody else's audit. Treating either as scope is the mistake this tool makes easiest, because the surrounding machinery is careful enough that a user will assume the care extends to the boundary too. What is enforced mechanically is network and process, not permission. There is no host allow-list in the visible surface, no dry-run mode, and no documented abort control, so the guard rails are a sealed network during discovery and a narrow egress grant during confirmation. Those bound what an audit can touch. They do not decide what an audit is allowed to touch. The scenario list widens the same gap: alongside blind capability audits and open-world bounty work, the file names incident investigation started from suspicious on-chain evidence, targeted follow-up on findings somebody else reported, and disclosure preparation. Each is normal security work and each needs a different answer to the question of whose permission covers it.

Framework-agnostic, and then strongest on two frameworks

Two bullets in the reasons list pull against each other, and the provider plumbing is worth a mention: the framework exposes itself through a command line, a React dashboard, a self-describing REST interface, an editor extension and a set of provider profiles, with two example profile files in the tree, one per supported coding agent. Everything is routed through one sandbox and one audit contract regardless of which agent is driving it. The first says the tool encodes no audit strategy for any particular stack, naming EVM contracts, proof systems, Rust, Go, JavaScript, protocols and crypto work as things it deliberately does not special case, so that audit capability improves as models improve without the framework being rewritten per stack. The second then explains why two specific stacks fit well: contract targets, because a local fork plus a test harness can prove a real on-chain effect, and proof-system targets, because a local prover and constraint harness can turn a subtle missing constraint into an executable counterexample. The file calls these high-signal examples rather than hard limits, which is the right framing, but a reader choosing a tool will read the second bullet and not the first. What agnosticism buys you here is coverage breadth; what the two named stacks buy you is verifiability, and those are not the same property.

A finding is real once a command exercising it has passed

The rule that gives this tool its shape is a definition of what counts as evidence. A finding is not real because the model called it plausible; it has to cite a passing local command that exercises the vulnerable path. Stronger findings are held to more: differential confirmation, independent refutation, and an explicit check that the proof of concept still does what the report claims once someone else reads it. That last one is the unusual one, and it is the difference between a finding and a claim in a write-up. The pipeline that produces those commands is seven named stages run as one tracked workflow, where a map stage inventories the audit surface, a dig stage writes and runs the proof tests, a verify stage confirms or refutes candidates by local execution, a confirm stage reproduces findings against the real target, and a report stage packages the reproduced ones. The stages are named here as the project describes them, not as instructions.

Discovery is sealed, and confirmation opens exactly one door

The network policy is the sharpest thing in the file and it is stated in two halves. Discovery runs with the network sealed, which is described as the reason findings come from the target material rather than being copied out of public disclosures, and it is also what makes a blind audit blind. Confirmation is where reproduction has to happen against something real, and there the network opens only for an explicit list of read, fork and fetch commands, while arbitrary model-written code stays sealed. Arbitrary code generated by the model never gets egress. That is a narrower grant than most tools offer and it is the difference between a tool that can look things up and one that can exfiltrate. The rest of the containment story is the container: the default backend refuses to fall back to running on the host, bind-mounts only a copied workspace and a package cache, drops Linux capabilities, sets the no-new-privileges flag, uses a read-only root with temporary directories in memory, and applies process, memory and processor limits when they are configured.

npm test builds the whole project before it runs one line

The test script is two commands: a build, then the runtime's own test runner pointed at the compiled test files. There is no test framework in the dependency surface, and because the build runs first, every invocation of the test command compiles the server and the dashboard interface before executing anything. That is a slower loop than most projects accept and a faster one than they need for a project that checks its own generated artefacts. The second half of the story is the set of consistency checks: one script verifies the API catalogue against the built output, one verifies database upgrades, one verifies the public export surface, and a fourth variant of that last one checks the committed state rather than the working tree. Four scripts whose only job is to catch a generated file drifting from the code that generates it is a governance habit, and it is unusual to see it spelled out this explicitly.

Three names, two homepages, and a one-major runtime range

The repository is called flounder. The published package is called flounders, in the plural, while the command it installs is called flounder in the singular, and the site the metadata points at uses the plural domain. So a reader who sees the domain first will search the index for the wrong name, and a reader who installs the package will find a differently named command in it. The manifest and the repository metadata also disagree about the homepage: one points at the product site, the other at the project's own readme. Underneath that, the runtime requirement is unusually narrow, a minimum of 24.13 and a hard ceiling below 25, so the project admits only the current Node major and refuses the next one, a choice that trades convenience for a guarantee that one runtime version is what the sandbox and the test runner were exercised against, and there are both a node version file and a version manager file in the root to say so twice. The package version is 0.4.0 and the repository has no tagged releases at all.

The maintainer surface is behind a flag and a human gate

One feature in the reasons list is deliberately not available to ordinary users. When the control plane is explicitly started with a maintainer flag, an experiment command helps a maintainer agent mine failures from the evaluation suite, constrains source proposals to an allowlist, and compares paired baseline and candidate runs. The file is explicit that this is not triggered by ordinary projects and is absent from the default user surface, and that promoting anything requires distinct cases and distinct bug families, passing hidden holdouts and controls, and no regressions in the paired runs, with merging and deployment left as human gates. The same separation shows up in deployment: the dashboard is described as a control plane while the audits themselves run on a daemon that can sit on another machine, which is how the file keeps target source and provider credentials on the executor host rather than on whatever machine you opened the browser on. That is the correct shape for a self-improving loop: the machine proposes, the gate decides, and a human still merges. The neighbouring feature is the evaluation group command, which runs validated positive, negative, control, replay and multi-target work items through one queue and survives a control plane restart with its concurrency, attempts and evidence intact.

Editorial conclusion

Separate the two claims before you use this. The engineering claim is strong and checkable: model-written code runs in a copied workspace inside a container with capabilities dropped, a read-only root and network switched off, discovery never has egress at all, and confirmation only opens the network for a narrow class of read and fetch commands. That is a real blast-radius reduction and it is the part worth relying on. The scope claim is weaker, because it is a sentence rather than a mechanism: the accepted target list includes a deployed lead and a previous run, neither of which is an authorization, and the visible surface offers no host allow-list, no dry run and no abort control. Do not point it at anything you have not been engaged to test, do not feed it other people's disclosures as if they were permission, and note before installing that it requires a very narrow runtime range, that the npm package, the command and the repository carry three different names, and that no release is tagged.

Frequently asked questions

What is adshao/flounder?

It is an AGPL-3.0 licensed TypeScript workflow that turns a coding agent into a white-hat security auditor, running seven tracked stages from workspace preparation through surface mapping, proof tests, synthesis, local verification, reproduction on a real target and reporting. The package is published as flounders at version 0.4.0.

How do you install the flounder skill?

With a single command run through the skills installer, adding the skill globally for two named agents at once. A second form installs the copy in a local checkout instead of fetching from GitHub, and then the installed skill is triggered by asking the agent to audit a repository with it.

What counts as a real finding in flounder?

Not a model's judgement. A finding has to cite a passing local command that exercises the vulnerable path, and stronger findings additionally pass differential confirmation, independent refutation, and a check that the proof of concept still does what the report claims.

Does flounder need network access?

Not during discovery, which runs network-sealed so that findings come from the target material rather than from public disclosures. During confirmation the network opens for an explicit class of read, fork and fetch commands only, while arbitrary model-written code stays sealed throughout.

What platforms does flounder suit best?

The file names EVM contract targets and zero-knowledge proof-system targets as the strong fits, because a local fork with a test harness can prove a real on-chain effect and a local prover with a constraint harness can turn a missing constraint into an executable counterexample. It states these are high-signal examples rather than hard-coded limits.

What Node version does flounder require?

A narrow range: 24.13 or newer and strictly below 25, so only the current Node major is admitted. The repository carries both a node version file and a version manager file, and it has no tagged releases despite a package version of 0.4.0.

Official sources

  1. adshao/flounder on GitHub
  2. Issues
  3. License: AGPL-3.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/adshao-flounder.svg)](https://hysenlabs.com/projects/adshao-flounder)