# Kiterunner: content discovery that knows an API needs a method

> A Go scanner built on Swagger specifications rather than a wordlist of paths, so it sends the right HTTP method, headers and parameters to each endpoint instead of hoping a GET at /api/v1/users lands.

**assetnote/kiterunner** — Contextual Content Discovery Tool

- Repository: https://github.com/assetnote/kiterunner
- Stars: 3,273 · Forks: 341
- Language: Go
- License: AGPL-3.0
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/assetnote-kiterunner

## Why a wordlist of paths stopped working

The introduction makes the argument about content discovery in two paragraphs, and the second one is the one that matters for understanding the tool. For the longest of time, content discovery meant finding files and folders. That is effective for legacy web servers hosting static files or responding with 3xx codes on a partial path, and it is no longer effective for modern web applications, specifically APIs.

The reason is the routing paradigm. Frameworks such as Flask, Rails, Express and Django follow the model of explicitly defining routes that expect certain HTTP methods, headers, parameters and values. A traditional scanner sends a GET to a guessed path. If the route exists but requires a POST with a JSON body and a bearer token, the scanner gets a 405 or a 404 and records nothing, even though the endpoint is right there in the application.

The README also notes the industry's response to this, and it is worth reading as criticism rather than as background. A lot of time has been invested in making content discovery tools faster so that larger wordlists can be used, but the art of content discovery has not been innovated upon. Kiterunner is positioned as the innovation: not only traditional content discovery at what it calls lightning fast speeds, but also bruteforcing routes and endpoints in modern applications.

The AGPL-3.0 licence matters here in a way it does not for a library you link into a proprietary product. This is a command line reconnaissance tool, and the licence choice is consistent with that: a hosted service built on it would need its source offered.

## Where the route data comes from

The dataset is the actual product here, so its provenance is documented in detail. Swagger files were collected from a number of sources, including an internet wide scan for the 40 or more most common Swagger paths, GitHub via BigQuery's public dataset, and APIs.guru, which is a public registry of machine readable API descriptions.

The GitHub source is the one with real use. Searching public repositories for Swagger specifications finds API descriptions written by people who never intended to publish them, which means the route list reflects what software actually looks like in the wild rather than what a documentation checklist produces. That is the same reasoning behind Assetnote's other published wordlists, which the README also links to for use as raw bruteforce input.

What the tool does with those files is described precisely: by collating a dataset of Swagger specifications and condensing it into their own schema, Kiterunner can use the dataset to bruteforce API endpoints by sending the correct HTTP method, headers, path, parameters and values for each request it sends. That condensing step is what makes the approach practical. A Swagger specification is verbose XML or JSON with descriptions, schemas and examples, and shipping tens of thousands of them raw would be unwieldy. Collapsing them into a minimal intermediate representation is what makes a compiled wordlist small enough to load and fast enough to scan with.

The dataset ships in two sizes, and the README gives both compressed and decompressed figures, which tells you a lot about the format. The large JSON is 118MB compressed and 2.6GB decompressed. The small JSON is 14MB compressed and 228MB decompressed. The compiled kite versions are far smaller: 40MB compressed and 183MB decompressed for large, 2MB compressed and 35MB decompressed for small. A compiled file roughly 14 times smaller than the JSON source is what tells you the intermediate representation is doing real work rather than being a serialization of the same tree.

## Building it and compiling a wordlist

There are three installation paths. The simplest is a pre-built binary from the GitHub releases page. Arch users have a package in the AUR named `kiterunner-bin`, and the README notes you can install it with an AUR helper like `yay`.

Building from source is four commands, and the README presents them as one block because the order matters:

```bash
# build the binary
make build

# symlink your binary
ln -s $(pwd)/dist/kr /usr/local/bin/kr

# compile the wordlist
# kr kb compile <input.json> <output.kite>
kr kb compile routes.json routes.kite

# scan away
kr scan hosts.txt -w routes.kite -x 20 -j 100 --ignore-length=1053
```

Two details in there are load-bearing. The binary is named `kr`, not `kiterunner`, and it lands in `dist/`, so the symlink is what puts it on your PATH. And compiling is a distinct step with its own subcommand, `kr kb compile`, taking a JSON input and producing a `.kite` output. The commented line `# kr kb compile <input.json> <output.kite>` is the README telling you the shape of the command before showing a real invocation.

The scan line at the end is a complete example rather than a placeholder. It takes a host list, a compiled wordlist, sets 20 connections per host and 100 parallel hosts, and sets an ignore-length filter, which is the option that suppresses responses whose body length matches a known generic page.

The ignore-length flag is a good example of the CLI being shaped by real scanning problems. Because the wordlist is large and the endpoints are real APIs, you will get a lot of responses that are technically 200 and are actually your application's generic error page. Filtering by content length removes those without needing to parse anything.

## The flags that matter, read from the CLI help

The CLI help in the README is more informative than a feature list, because these are the knobs that determine whether a scan is useful or a denial of service against yourself.

`-A` selects a wordlist from wordlist.assetnote.io by type and name, for example `apiroutes-210228`. It also accepts a maxlength suffix with a semicolon, so `apiroutes-210328:20000` uses only the first 20000 values. That is how the README's own examples get a large route set without committing to the full dataset.

`-x` and `-j` are the two concurrency controls and they are separate on purpose. `-x, --max-connection-per-host` caps connections to a single host, default 3. `-j, --max-parallel-hosts` caps how many hosts are scanned concurrently, default 50. Scanning a thousand hosts and scanning one host hard are different failure modes, and the tool separates them. The README's quick start pushes these to 20 and 100, which is aggressive.

Several flags exist specifically to reduce false results. `--fail-status-codes` takes a list of status codes to treat as failure and, as the help text says, overrides the success status code list. The README's examples pass 400,401,404,403,501,502,426,411, a sensible list that covers missing, unauthorized, forbidden and the family of malformed-request codes. `--ignore-length` accepts ranges rather than single values, so `100-105`, `1234` or `123,34-53` all work, and it is inclusive at both ends.

Then there are the flags for shaping what gets sent. `-H` adds headers, with a default of `x-forwarded-for: 127.0.0.1`. `--force-method` ignores the method recorded in the wordlist and uses one you specify, which is what you want when a specification is wrong about the verb. `--filter-api` restricts the scan to APIs matching a given ksuid, and `--blacklist-domain` prevents redirects from being followed to particular domains, which prevents a scan from wandering off onto an unrelated host. `--delay` inserts a pause between requests to a single host, `--disable-precheck` skips host discovery, and `--max-redirects` caps redirect following at 3 by default.

## Bruteforce modes, depth scanning and the kite file

The usage section distinguishes several modes beyond API scanning, and the top-level usage line shows how thin the surface is:

```
kr [scan|brute] <input> [flags]
```

`scan` takes hosts and a compiled kite wordlist, or an Assetnote wordlist via `-A`. `brute` is the traditional mode, and it accepts a dirsearch style wordlist directly. The README's own example shows the mutation syntax working: `-w dirsearch.txt` with `-exml,asp,aspx,ashx` expands a wordlist containing `%EXT%` markers into the requested extensions, and `-D` handles the dirsearch case handling.

The quick start also shows the case where you have no local wordlist at all, scanning hosts with `-A=apiroutes-210328:20000` and nothing else, and the case where you want your own routes file plus the Assetnote ones together.

The technical section covers two things that matter for building a curated list. Depth scanning walks a nested path prefix before probing, so `/api/v2` can be explored as a base rather than treating every full path as flat. And the head syntax lets you filter an Assetnote wordlist to entries matching a path prefix, which is how you keep a scan scoped to one part of a large application.

There is also a converter between file formats and a request replay feature, both of which exist because a scan that finds something interesting should lead to a reproducible single request rather than a re-run. The intermediate data type, called PRoutes in the README table of contents, is the internal representation between parsed Swagger and the compiled kite file, and the kite file format itself is documented as the final on-disk artifact.

The practical workflow, then, is: find the compiled or JSON dataset, compile it if you have your own Swagger files, scope it with depth and head filters, scan with sane per-host concurrency, filter results by status code and body length, and replay anything interesting one request at a time.

## The Go implementation and what the dependency list reveals

The repository layout is conventional Go with a few directories that map directly onto the features. `cmd/` holds the Cobra commands, `internal/` the private implementation, `pkg/` the public API, `routes/` the wordlist definitions, `benchmark/` for performance work, `api-signatures/` for the schema the route data is condensed into, and `hack/` for development scripts. A `doc.go` at the root holds package documentation, `example.kiterunner.yaml` shows the config file format, and `.goreleaser.yaml` plus a `makefile` at the root describe how release binaries are built.

The `go.mod` file is more revealing than the layout, because the dependency list is the architecture. Three entries stand out. `github.com/valyala/fasthttp` is the HTTP client, and choosing fasthttp over the standard library is the single decision that makes this tool fast: it reuses buffers and objects across requests instead of allocating per request, which matters enormously when you are issuing tens of thousands of requests. `github.com/valyala/bytebufferpool` and `github.com/valyala/fasttemplate` support that model.

`github.com/fasthttp/router` sits alongside, which is an unusual pairing and suggests the tool also serves its own small interface rather than only making requests. `github.com/gogo/protobuf` and `github.com/francoispqt/gojay` together cover both protobuf and JSON object mapping, which is what you need to serialize route data compactly in both formats. `github.com/segmentio/ksuid` supplies the identifiers used by the `--filter-api` flag, `github.com/valyala/fasthttp` and `github.com/spf13/cobra` and `viper` handle CLI and config, `github.com/rs/zerolog` the logging, and `github.com/schollz/progressbar/v3` with `github.com/vbauerster/mpb/v6` the two progress bar libraries, one simple and one composable.

One thing to note is the `go 1.15` directive at the top of the module file. That is an old language version floor for a project whose last push is dated 2026, and it lines up with the release dates: the newest tagged releases are v1.0.1 and v1.0.2, both from April 2021, while the repository has been pushed as recently as July 2026. So the source has moved well past what the tags capture, and pinning to an old Go directive means you can build it with a modern toolchain without giving up compatibility. It also means the README's badge set is what tells you the project is alive, since there is no recent release to point at.

## Judging it against the tools it replaces

The comparison set for a tool like this is ffuf, feroxbuster, gobuster and dirsearch, and the honest summary is that Kiterunner is not trying to beat them at what they do well.

Those tools are general purpose content discovery engines with strong support for recursive scanning, response filtering by size, word and regex, match and replace rules, and per-request headers and authentication. They are the right choice when you are looking for hidden files, backup extensions, admin panels and directory listings on a conventional web application. Kiterunner's `brute` mode accepts a dirsearch style wordlist and extension expansion, so it can do a version of that, but that is compatibility rather than its purpose.

What Kiterunner adds is route knowledge. If an application publishes or accidentally exposes a Swagger specification, the routes in it are enumerable with the correct method, path parameters and header requirements, and no amount of wordlist size will find them without that structure. For an assessment where the API surface is the target, that is the whole game.

The tradeoffs follow from the same place. The tool needs data. Out of the box it needs the Assetnote CDN or the AUR or GitHub release, and in a custom environment you need to supply Swagger files yourself, which means writing or finding specifications for the target. The dataset also encodes the world as it looked when it was collated, so a route added to an application since then will not be in it, and a route removed since then will produce false positives. Neither is fatal, but both mean a scan is a supplement to reading the application's own documentation and behaviour rather than a replacement.

On the operational side, a tool this fast is easy to misuse against infrastructure you do not own, which is worth stating plainly. The per-host and per-parallel-host caps exist because the failure mode of getting them wrong is a service that stops responding. Use the delay flag if you are pointing this at anything shared, and keep the scope written down in your client agreement, because a tool that issues a few hundred thousand structured requests against an API is doing something a normal scanner never does.

## Conclusion

Kiterunner's contribution is a change of input rather than a change of speed. Traditional content discovery asks whether a path exists, which works against servers that host static files or answer a partial path with a redirect, and fails against frameworks where a route is defined as a combination of path, method, headers and parameters. Kiterunner takes its routes from real Swagger specifications, collapses them into its own schema, compiles that into a compact kite file, and replays each entry as a correctly shaped request. That is why it can find endpoints no wordlist would guess, and also why it needs a source of Swagger data, which is either downloaded from the Assetnote CDN or collated yourself. The operational picture is what to weigh before adopting it. Scanning is fast enough to be unpleasant to run without care, which is why the concurrency flags are per-host and per-parallel-host rather than one global number, why a delay flag exists, and why redirects can be restricted to a host blacklist. The release history is also worth knowing: the newest tagged release is v1.0.2 from April 2021, while the repository was pushed in July 2026, so the tagged binaries trail the source. Build it with `make build` rather than assuming a release tarball matches the current tree.

## FAQ

### How do I install and build Kiterunner from source?

Clone the repository and run `make build`, which produces the binary as `kr` in `dist/`. You then symlink it onto your PATH with `ln -s $(pwd)/dist/kr /usr/local/bin/kr`. Arch users can skip this and install the `kiterunner-bin` AUR package with a helper such as `yay`. Pre-built binaries are also published on the GitHub releases page, but the newest tagged release there is from 2021, so building from source is the safer option if you want the current tree.

### What file does Kiterunner need to scan, and how do I compile it?

Kiterunner scans a `.kite` wordlist, which you create from a JSON dataset with the `kr kb compile` subcommand, taking an input JSON and an output kite path. The README publishes both formats at two sizes: routes-large.json is 118MB compressed and 2.6GB decompressed, while routes-large.kite is 40MB compressed and 183MB decompressed, so compiling shrinks the data considerably. You can download the JSON and compile it yourself, or download a pre-compiled `.kite` directly and skip that step.

### What do the -x and -j flags do in a Kiterunner scan?

They are two separate concurrency limits rather than one global setting. `-x, --max-connection-per-host` caps how many connections are open to a single host and defaults to 3, while `-j, --max-parallel-hosts` caps how many hosts are scanned at the same time and defaults to 50. Splitting them matters because scanning a large host list and hammering one host hard are different failure modes. There is also a `--delay` flag that pauses between requests to a single host if you need to go slower.

## Sources

- [assetnote/kiterunner on GitHub](https://github.com/assetnote/kiterunner)
- [Issues](https://github.com/assetnote/kiterunner/issues)
- [License: AGPL-3.0](https://github.com/assetnote/kiterunner/blob/main/LICENSE)
- [README](https://github.com/assetnote/kiterunner/blob/main/README.md)
- [Releases](https://github.com/assetnote/kiterunner/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/assetnote-kiterunner
