Model or dataset
FrancescoStabile/numasec avatar
FrancescoStabile/numasec

numasec keeps findings in a durable operation instead of a chat transcript, and its root test script fails on purpose

The AI Agent for Cyber Security.

803 stars101 forksTypeScriptAGPL-3.0

At a glance

What is it?
An AGPL-3.0 terminal security agent that borrows the tools already on your machine and wraps runbooks, evidence and reporting around them. The monorepo runs on Bun, pins several dependencies to beta channels, and tells you which local tools are missing before you start.
Who is it for?
Judge it as a workspace wrapper rather than a scanner. If your workflow already runs from a terminal and your findings already need evidence attached to them, the operation model and the durable export are the parts that pay.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 147 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The root test script is a guard that fails on purpose

The manifest for a package distributed on npm with a global install command describes itself differently from the document around it: it calls numasec a terminal native AI cyber operator harness, marks itself private, declares the module type, and pins the package manager to Bun at 1.3.11. The scripts block is where the project states its own rules.

json
"dev": "bun run --cwd packages/numasec --conditions=browser src/index.ts",
"lint": "oxlint",
"typecheck": "bun turbo typecheck",
"postinstall": "bun run --cwd packages/numasec fix-node-pty",
"test": "echo 'do not run tests from root' && exit 1"

Four of those are ordinary. The fifth is a deliberate stop: running the test script at the repository root prints a sentence telling you not to do that and exits non-zero. It exists so that a generic pipeline step cannot report a pass by running nothing. Type checking is a turbo task across workspaces, linting is oxlint with its configuration at the root, and the postinstall step runs a fix for node-pty, which is the native piece a terminal agent needs in order to host a real shell rather than a simulated one.

A global npm install, then a doctor command before any runbook

The install path is one command, followed by launching the binary with no arguments:

bash
npm install -g numasec
numasec

The four commands that follow are the intended first session, and the order matters more than the individual entries. The doctor command checks the environment before any work starts. The mode command sets the posture. The runbook command takes a name and a target, and the example targets a local development server rather than anything on a network. The share command exports the operation. Pointing the first run at a local application is a deliberate choice in an example that otherwise covers pentest, OSINT and lab work, and the surrounding text repeats the constraint that the target scope has to stay explicit and that the tool is for systems you own, labs, CTFs, or targets where testing is permitted.

The runbook identifier in that example is worth reading as a naming convention, since it spells out the domain and the task: a web triage runbook for application security work.

Findings live in the operation, not in the transcript

The central claim is that findings are not dumped into chat. They live in the operation, and each one can carry a state, a severity, evidence, a replay status and a next action. The stated purpose is that the agent can keep working without losing the thread, which is a different thing from keeping a conversation long.

Three rules follow from that structure. Weak signals are allowed to stay weak rather than being promoted into findings. Rejected claims remain visible instead of being deleted, so a hypothesis that did not pan out is still readable later. And reportable findings need proof, which is what the evidence and replay state is for: enough material to understand, verify and reproduce the work. That is also what the replay entry in the capability table refers to, keeping the material needed to reproduce something important rather than only describing it.

Operations are described as durable on top of that. They can be named, renamed, resumed and exported, and the line the project uses for it is that a security workflow should not disappear because the chat ended.

TAB switches between five postures, and only two are called mature

The agent is switched with TAB across five kinds of work: AppSec, Pentest, OSINT, CTF or lab, and research-style work. The argument given is that these do not need the same posture, so you switch the agent when the work changes rather than forcing one generic assistant to behave the same way everywhere. The trade is implied and stated: a shared context does not survive a posture switch as easily as a shared transcript would, so the operation record is what carries the work across.

The project is also explicit about which of those surfaces it claims maturity on. It says numasec is strongest today for authorized AppSec and pentest workflows, and that other cyber surfaces exist or are possible but are not marketed as equally mature yet. That sentence is the most useful calibration in the whole document, because a terminal agent whose runbooks are tuned for web triage will behave differently on an OSINT task than its posture name suggests.

The audience list matches that weighting. AppSec engineers triaging web apps, APIs, dependencies, auth flows and reports, pentesters on scoped work with terminal tools and deliverables, bug bounty hunters, security researchers, and CTF and lab users who want structure while keeping direct control of the tools.

The dependency catalogue leans on beta channels and one dated pseudo-version

The workspace catalogue in the manifest is where the stack becomes visible, and it is unusually willing to sit on pre-release versions. The Effect runtime, its OpenTelemetry package and its Node platform package are all pinned to the same 4.0.0 beta build. The database layer and its migration kit are both on a beta build with a commit hash appended. A diff rendering package is on a beta. One authentication package is pinned to a pseudo-version of the form 0.0.0 dated by timestamp rather than by release.

That is not automatically wrong, and pinning is better than floating, but it is a real maintenance surface: a beta tag can move or be republished, and a timestamp pseudo-version points at one specific commit rather than at a release you could audit. Everything else in the list is more conventional, with Hono and its OpenAPI helper, the AI SDK, Playwright for tests, an octokit rest client, npm's own arborist, ULID identifiers, cross-spawn, Luxon for dates, marked with a syntax highlighter extension, DOMPurify, fuzzysort for ranking, semver and Tailwind through its Vite plugin.

Two of those deserve a second look for a security tool specifically. DOMPurify exists to sanitise HTML, which tells you the project renders untrusted content somewhere. Arborist means the project resolves dependency trees, which is what a tool that inspects installed packages would need.

AGPL-3.0, a NOTICE file, and a patches directory

The licence is the GNU Affero General Public License, and for a tool that people may want to embed in an internal platform that choice has consequences the rest of the packaging makes concrete. The root carries a licence file, a NOTICE file for third party attribution, a SECURITY.md, a CONTRIBUTING file, a changelog and an AGENTS.md, plus an assets directory holding the demo recording.

The repository layout is a monorepo with packages under a single workspace pattern, driven by turbo, with a Bun lockfile and a Bun configuration file at the root. Three entries there are worth noting for anyone vendoring it: a patches directory, which in a Bun workspace means dependencies are being modified in place rather than merely pinned; a script directory; and a docs directory. A patches directory in a project with an AGPL licence and a NOTICE file is the sort of combination that deserves a look before you redistribute anything derived from it, since patched dependencies change what you are actually shipping.

The lint configuration is a single JSON file at the root for oxlint, and the editor configuration file sits beside it, so style and lint rules are both centralised rather than per package.

Local tools are reported as available, missing or degraded

The tool story is a posture rather than a bundle. numasec uses what is installed on your machine and, in the capability table's wording, shows what is available, missing or degraded. That third state is the one that carries information: a tool that is present but not working properly is a different problem from one that is absent, and a workspace that only reported presence would hide it.

The negative positioning is stated just as explicitly. It is not a chatbot, not a scanner wrapper, and not a Burp or Kali replacement. Read together with the doctor command in the first session, the intent is a tool that inspects and coordinates the environment you already built instead of shipping another environment that competes with it.

Cyber knowledge is treated the same way, as something brought into the workflow rather than bundled: vulnerability intelligence, advisories, methodology and tool documentation are inputs the runbooks consume. The reports entry follows the same principle, generating deliverables from the operation state instead of asking you to reconstruct everything at the end.

The workflow diagram stops at its first word

The section describing how the workflow fits together opens by making its point in prose: numasec is not just a prompt with tools, and it keeps the target, operation, posture, runbook, local tools, observations, findings, evidence, replay and report connected. That list is eleven nouns and it is the real architecture summary, covering both the security objects and the workflow objects.

What follows the prose is a diagram, and the diagram is where the section ends. A fenced mermaid block begins and stops at the first characters of a declaration, so the relationship between the eleven elements is described in words and never drawn. The next sections the reader reaches are not visible from what the document provides.

For anyone trying to understand the design, the practical source is AGENTS.md in the repository root, which exists for coding agents working on the project itself, alongside the docs directory. The capability table remains the most complete inventory of what the tool claims to hold at once: agent, local tools, runbooks, postures, operation memory, findings workflow, evidence and replay, cyber knowledge, reports and share bundles.

Editorial conclusion

Judge it as a workspace wrapper rather than a scanner. If your workflow already runs from a terminal and your findings already need evidence attached to them, the operation model and the durable export are the parts that pay. Before installing, note three things. The licence is AGPL-3.0, which is the wrong choice if you intend to embed the agent in a closed source product, and a NOTICE file sits next to the licence for third party attribution. The dependency catalogue leans on beta channels, including the Effect runtime and the database layer, plus one package pinned to a dated pseudo-version. And the root test script exits non-zero deliberately, so a naive test run at the repository root tells you nothing about whether the thing works.

Frequently asked questions

How do I install numasec and start a first session?

Run npm install -g numasec and then launch numasec, then use the doctor command, set a mode, run a named runbook against a local target such as the example appsec-web-triage runbook pointed at localhost, and use the share command to export.

Which security work is numasec said to be strongest at?

Authorized AppSec and pentest workflows. Other surfaces such as OSINT, CTF and lab work and research exist as switchable postures, but the project says they are not marketed as equally mature.

What licence does numasec use and can I embed it in a closed source product?

The project is licensed under AGPL-3.0, with a NOTICE file at the root for third party attribution, which is a constraint to weigh before embedding it in a product you do not publish under the same terms.

Why does running the test script at the numasec repository root fail?

The root test script is set to print a message saying not to run tests from root and exit with a failure code, so a generic pipeline step cannot report a passing test run that never executed anything.

How does numasec handle local tools that are not working properly?

It reports tools as available, missing or degraded, treating a tool that is installed but not functioning as a distinct third state rather than folding it into presence or absence.

Official sources

  1. FrancescoStabile/numasec on GitHub
  2. Issues
  3. License: AGPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/francescostabile-numasec.svg)](https://hysenlabs.com/projects/francescostabile-numasec)