# miasma: a scraper tarpit that admits its own deployment risk

> A Rust server you point at crawlers that steal your public pages, which answers every request with a page of synthetic slop threaded with links that all point back into itself, so a scraper that follows links is held in a loop until it has consumed its budget. The interesting part is the rate limiting that stops the tarpit from becoming a self-inflicted denial of service.

**austin-weeks/miasma** — Trap AI web scrapers in an endless poison pit.

- Repository: https://github.com/austin-weeks/miasma
- Stars: 1,191 · Forks: 40
- Language: Rust
- License: GPL-3.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/austin-weeks-miasma

## The mechanism: slop pages whose links all point back into the pit

Understand the mechanism before anything else, because it is simple and it is the whole product. You run the server, and you point some scraper's requests at it. Every response is a page of synthetic content, described as poisoned training data, drawn from a third-party generator the readme calls a poison fountain. The page is threaded with multiple links, and every one of those links points back into the same server. A scraper that follows links therefore never leaves, because the only thing on the page is more of the same. The readme's framing is that AI companies scrape the public internet at scale for training data, that if you have a public site they are already taking your work, and that this is your way to fight back by feeding the slop machines an endless buffet of their own kind. The self-reference is the trick. Without it you would just be serving junk once; with it you convert a scraper into a loop that keeps requesting, which is the one thing a training-data pipeline is worst at terminating. The design assumes the scraper follows links, and the readme's own robots.txt advice, disallowing the path for well-behaved crawlers, is the mechanism by which polite bots are kept out of the loop so search engines do not index the garbage.

## The server that admits you might DDoS yourself

This is the most useful thing in the readme and the reason to read the whole thing. A tarpit that holds a connection open and makes the client come back is, by design, a way to tie up resources, and the readme is blunt that deploying it carries inherent risk and carries a caution block telling you to read the configuration and the disclaimer before you use it. The specific danger is not the scraper, it is you. If real visitors, or a legitimate crawler you did not mean to trap, are routed into the pit, the pit will happily exhaust your memory and your bandwidth. The readme addresses this in two ways. The server has an in-flight limit, a maximum number of concurrent connections, and when a client exceeds it the client is told to back off rather than queued. And the proxy configuration in the walkthrough is rate-limited per client user agent, with a request status set so that over-limit requests are rejected outright, specifically so the site owner does not accidentally overwhelm the origin with their own trap. There is also a reserved memory zone sized in the example config, which is the kind of thing you only write down if you have measured what a connection actually costs. Taken together this is a tool that has thought hard about being a liability to the person running it, which is rarer than it should be in this category.

## Bounding the loop: link count, depth, and letting the scraper go

The readme walks through a worked example and the numbers in it are the operational lesson. The setup configures a link prefix so all generated links route through your reverse proxy, sets a maximum in-flight connection count, and then, crucially, uses a link count and a maximum depth to let a trapped scraper go once it has consumed a target number of pages. The example targets roughly a hundred thousand poisoned pages per scraper, which the readme estimates at about 250 megabytes of total data per scraper, and it chooses a link count and depth to reach roughly that. The in-flight figure in the example, fifty, is tied to a memory estimate of about fifty to sixty megabytes at peak. So the tool is not designed to hold a scraper forever. It is designed to hold one for a bounded budget of your choosing, then release it, and the readme says so plainly. That is a deliberate and, to my mind, correct engineering choice on three grounds. It bounds the egress bandwidth you pay for, which matters because a tarpit that never releases is a money leak. It bounds your memory, so the tool cannot exhaust the host. And a finite, self-terminating interaction is far less likely to trip upstream abuse detection and get your address blocked or your host reported, which would cost you the scraper protection you installed it for. The one-line example configuration, which sets a link prefix, a max in flight, a link count and a max depth together, is the shape of a tool you are meant to tune per site rather than run blind.

## Hidden links, a reverse proxy, and why the hidden part is rigged to humans

The deployment walkthrough is the part you will copy, and it is worth reading closely because the hidden-link technique is built to defeat human and assistive navigation specifically. The site embeds links to the trap path using a set of attributes together: the style is set to hide the element, the element is marked as hidden from assistive technology, and it is removed from the tab order with a negative tab index. The readme states plainly that this combination makes the links invisible to human visitors and ignored by screen readers and keyboard navigation, so that they are only visible to a scraper. That is a deliberate accessibility rigging: the trap is only effective if a human never clicks it and a screen reader never announces it, and the technique achieves exactly that. The second half of the setup is the reverse proxy. The walkthrough uses a standard web server, configures the trap path on it to pass requests through to the server, adds a redirect so the path with and without a trailing slash both work, and adds the per-user-agent rate limit discussed above. The walkthrough installs it with the package manager, or runs the container with a port mapping, and the whole invocation is a handful of flags: 

```sh
miasma --link-prefix '/naughty-bots' -p 9855 -c 50 --force-gzip --link-count 5 --max-depth 8
```

The point of routing through a proxy rather than sending scrapers straight to the server is control. The proxy is where you decide what gets trapped, how fast, and how much of your own traffic might be caught, which is exactly the decision the server should not be making on your behalf.

## Forcing compression to cut your own egress bill

One configuration detail in the example is quietly one of the more important operational choices. The server can be forced to compress all of its responses regardless of what the scraper's accept-encoding header asks for. The readme's reasoning is straightforward: gzipped responses are significantly smaller, and smaller responses mean less egress cost. That matters because the entire model is that you are spending bandwidth to waste someone else's compute, so your own bandwidth bill is the price of admission and it is directly proportional to how much you serve. Uncompressed slop pages are much larger than the same pages compressed, so a scraper you hold for a hundred thousand pages costs you several times more if you leave compression off. It is a small flag with a large effect on the one number that decides whether running this is cheap or expensive. Combined with the in-flight and depth limits, the readme has quietly given you the three levers that govern cost: how many connections you hold at once, how long you hold each one, and how many bytes each response is. The readme also mentions a metrics capability for tracking scrapers, which is the piece you need in order to tune those levers against real traffic rather than guessing.

## The contradiction of a scraper trap shipping its own crawl rules

The last section of the walkthrough is a joke with a serious point in it. After telling you to feed robots to an infinite loop, the readme tells you to protect well-behaved bots and search engines from the trap via your crawl-rules file, and shows the exact disallow line for the trap path. This is not a contradiction, it is the difference between a denial of service and a selective trap. The whole design rests on the assumption that the target follows links, and that assumption is true of scrapers and false of search engine crawlers, which respect the rules and which you still need to index your real pages. If you trap your own search engine you do not just lose a crawler, you lose indexing for the site you were trying to protect. The readme is right to insist on the disallow line, and it is worth being explicit about why, because the instinct when you deploy an aggressive anti-scraping tool is to forget that the good guys read the same file the bad guys ignore. It is also a good signal about the tool's overall posture. It is defensive, self-limiting, and honest about the fact that it needs a carve-out for the traffic you actually want.

## Conclusion

Adopt miasma if you run a public site that is being scraped for training data and you are willing to spend a little egress to waste a scraper's crawl budget, because the mechanism is sound: a tarpit that serves slop and links to itself, bounded so it cannot run away from you. Do not adopt it expecting to block anything, because it does not deny a scraper, it wastes its time, and a polite crawler that ignores the links is unaffected. Two things to set before you point it at your own traffic. The in-flight and rate limits, because the readme is explicit that without them you can DDoS yourself, and it even shows a proxy config that rate-limits by user agent precisely to avoid that. And the link count and depth caps, because the readme recommends releasing a scraper after it has eaten a set number of pages rather than holding it forever, which bounds the egress you pay and is the difference between a tarpit and an accidental denial of service against your own origin.

## FAQ

### What does the miasma scraper trap actually do?

It serves synthetic poisoned pages threaded with links that all point back into the same server, so a scraper that follows links is held in a loop. The content is generated from a third-party poison generator the readme calls a poison fountain, and the intended target is scrapers collecting public sites for AI training data.

### How do I run miasma?

Install it with cargo or a prebuilt binary and run it with default configuration, or use the official container image with a port mapping. The same configuration flags apply to the container and to a compose cluster, and the tool can also be embedded as a library in an existing Rust server.

### Why does miasma warn about deployment risk?

Because a tarpit can tie up your own resources if real traffic is routed into it. The readme carries an explicit caution, ties a peak memory figure to the in-flight limit, rejects excess connections outright rather than queueing them, and shows a proxy configuration that rate-limits by user agent specifically so the owner does not overwhelm their own origin.

### Does miasma trap scrapers forever?

No. The readme recommends bounding the interaction with a link count and a maximum depth, releasing a scraper after it has consumed a target number of pages, with an example that targets roughly a hundred thousand pages and about 250 megabytes per scraper, to bound egress cost and memory.

### How are the hidden trap links kept away from human visitors and search engines?

The links are hidden with a display-none style, marked as hidden from assistive technology, and removed from the tab order with a negative tab index, so only scrapers see them. The readme also tells you to disallow the trap path for all agents in your crawl-rules file so search engines and polite bots are not trapped.

## Sources

- [austin-weeks/miasma on GitHub](https://github.com/austin-weeks/miasma)
- [Issues](https://github.com/austin-weeks/miasma/issues)
- [License: GPL-3.0](https://github.com/austin-weeks/miasma/blob/main/LICENSE)
- [README](https://github.com/austin-weeks/miasma/blob/main/README.md)
- [Releases](https://github.com/austin-weeks/miasma/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/austin-weeks-miasma
