Self-hosted service
germondai/trawl avatar
germondai/trawl

trawl's tier ladder climbs to a residential proxy on its own, so the authorization boundary has to be yours

Self-hosted scraping engine — bypasses any JS challenge & captcha: Cloudflare, Turnstile, reCAPTCHA, hCaptcha, GeeTest. FlareSolverr & Byparr alternative and drop-in replacement for your *arr stack.

933 stars54 forksTypeScriptAGPL-3.0

At a glance

What is it?
trawl is a self-hosted scraping engine that presents itself as a FlareSolverr and Byparr replacement for the *arr stack, moving from plain HTTP through cached browser sessions to a fresh challenge solve and an optional residential proxy. The engine internals are documented in detail. The escalation policy is not gated per host, and the two features that widen the blast radius most, the MITM forward proxy and the MCP endpoint, are both switched off by default.
Who is it for?
trawl is well specified as an engine and poorly bounded as a policy, and those are different problems. Use it only for hosts where you hold written permission, on a machine you would be willing to have a CA certificate installed on, and with MCP and the metrics dashboard left off.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Nothing in the tier ladder checks whether you were authorized to hit the target

The execution model has four rungs: plain HTTP fetch, cached browser session, fresh challenge solve, and an optional residential proxy. Climbing them is automatic. A request starts at the cheapest rung and escalates when a wall is detected, so the decision to send traffic through a third party network is made by the engine, not by you, on a per request basis. The single control that touches the ladder is SCRAPE_MIN_TIER, and its own comment in the settings file is unambiguous about scope: lowest scraper tier allowed for every endpoint, where 1 is HTTP, 2 is cached browser, 3 is fresh browser, and 4 is residential proxy, with a note that values above 1 increase browser load. One number, global scope, no host list. No dry run mode and no abort switch appear among the documented settings, so there is no way to watch what the fourth rung would do before letting it happen. The one setting shaped like an allowlist, MCP_ALLOWED_ORIGINS, has nothing to do with scrape targets. It names browser origins allowed to call the MCP endpoint. In practice the authorization boundary is entirely outside the process, which means it has to be enforced somewhere else: a separate proxy, an egress firewall, or simply a discipline about what you paste into the request field.

The forward proxy asks you to install its CA certificate in your trust store

A second feature widens the scope in a different way. The compose file publishes two ports. The first carries the API and the FlareSolverr compatible /v1 endpoint. The second is a MITM forward proxy on a configurable port, and the comment above it says it is only used when MITM_ENABLED is true. It also states where the certificate comes from and where it has to go: the CA cert is served from port 8191 at the path /proxy-ca.crt and must be installed in each client's trust store, with the system keychain and Java cacerts named as the examples. Java cacerts is the entry worth pausing on, because that is the shared trust store for every JVM on the machine, not one application. Once a CA is trusted there, TLS interception stops being a property of one container and becomes a property of everything on that host that speaks TLS, and it stays after the container is removed. The proxy itself is described as challenge aware, forwarding normal traffic directly and escalating tiers for detected walls, with WebSocket, binary body, and Range or 206 support carried through.

The MCP endpoint's origin list is not an authentication control

The MCP surface for agent clients ships switched off. MCP_ENABLED is set to false, and the comment attached to it is a warning rather than a default: keep disabled on public networks. The guard that does exist is MCP_ALLOWED_ORIGINS, a comma separated list of browser origins permitted to call /mcp, and it is empty by default. The behaviour around it is the part to understand before exposing anything. The settings file states that requests without an Origin header, described as normal server to server clients, are accepted. That means the allowlist only constrains requests that present a browser origin, and any client that omits the header is admitted by design rather than by failure. The threat model behind it is therefore browser based cross origin calls, not authenticated API consumers, and it should not be read as a second authentication layer. For anything beyond local experimentation, the choices are to leave the flag off, to place a reverse proxy that does authenticate in front of it, and to bind the service to loopback rather than relying on the origin list.

The dashboard stores request history in SQLite, while the compose file publishes the port on all interfaces

The metrics dashboard is optional and disabled by default, with two conditions attached. METRICS_DASHBOARD_ENABLED is false, and the guidance is to set it true only when its port is bound to 127.0.0.1. A token of at least 32 characters protects the data if the port turns out to be reachable from elsewhere, and METRICS_DASHBOARD_TOKEN is empty in the shipped file. What the dashboard records is not ephemeral: persistent request history, tier outcomes, failure causes, and live updates, written to SQLite at /data/metrics/trawl.sqlite, and the compose file mounts a persistent named volume at /data/metrics so the file survives container replacement. That history is a record of which hosts were requested and by which tier, which is exactly the record you want if something goes wrong and exactly the record you do not want lying around unmanaged. There is a gap between the guidance and the shipped port mapping. The compose file publishes the API port as a plain host to container mapping with a configurable left side and a fixed 8191 on the right, with no host address, so Docker binds it on all interfaces by default rather than on loopback. The dashboard itself is served over plain HTTP at localhost:8191/dashboard on that same port, and the image used in that dashboard is described as illustrative traffic rather than output from an instance.

The default pool size cannot reach DataDome, and the only mention of it is a comment in the settings file

The advertised coverage names Cloudflare, Akamai Bot Manager, and Imperva or Incapsula, with captcha handling for CF Turnstile and Interstitial, reCAPTCHA v2, hCaptcha, GeeTest v4 Slide, ALTCHA, and Friendly Captcha v1 and v2. DataDome does not appear in that list. It appears in a comment beside BROWSER_HEADFUL_POOL_SIZE, whose default is 0, with the note that this is the number of browsers in the headful sub pool, used only for DataDome escalations, warmed on first use, additional to BROWSER_POOL_SIZE at roughly 380 MB each, and the instruction to set it to 1 to scrape DataDome targets. So one named vendor is reachable only through a setting that ships off, and the cost of turning it on is stated in the file as extra memory. The rest of the pool numbers describe what happens under load. One warm browser by default, two content processes each, recycling after eight temporary challenge contexts with 0 disabling it, and a 15 second acquire timeout that is documented as the wait for a free browser before rejecting a request. That last value is the one to watch: concurrency above the pool size turns into rejected requests, and the surrounding numbers, a 180 second grace period before a stalled browser is reclaimed, 10 seconds to close a wedged one, and 90 seconds to launch a fresh one, describe recovery rather than queueing.

reCAPTCHA v2 is answered with speech recognition, and solved sessions are cached for an hour

Two mechanisms are worth naming precisely because they define what the project actually does. The first is captcha handling by audio. reCAPTCHA v2 is listed with the parenthetical free STT, and the corresponding feature line states that no paid solver API is required, since reCAPTCHA audio can use Google's free speech to text endpoint or an optional local Whisper service running locally. The second is session reuse. Solved cookies and browser identity are stored in Redis, and an accepted session can avoid a fresh solve, which is what keeps repeat traffic off the expensive rungs. The cache is configured through SESSION_CACHE_DRIVER defaulting to redis, REDIS_URL, REDIS_SESSION_TTL_SECONDS defaulting to 3600, and a memory driver alternative capped at MEMORY_SESSION_CACHE_MAX_ENTRIES of 1000, alongside connect and retry budgets of 5000 milliseconds each. Underneath the browser is Camoufox, a Firefox build described as fingerprint patched at the C++ and Juggler level to reduce automation signals. The practical consequence for anyone running this is that identity state persists across requests, which is what makes it fast and also what makes an unscoped instance a liability.

The FlareSolverr comparison is hosted on the project's own site and covers selected cases

Compatibility is the reason many people arrive here, and it is concrete. The /v1 endpoint accepts the FlareSolverr command shape, and the stated consumers are Prowlarr, Jackett, Sonarr, and the wider *arr ecosystem. The documented request sample for it is incomplete: the JSON body stops partway inside the example hostname, so the full request shape is not shown there. Getting it running is three commands and a health check:

bash
# Clone and configure
git clone https://github.com/germondai/trawl
cd trawl
cp .env.example .env

# Start scraper + Redis
docker compose up -d

# Verify
curl http://localhost:8191/health

First boot takes 15 to 30 seconds while the browser pool warms, and later starts are quick. There are three compose files at the root, the default one plus minimal and prod variants. The container image is pulled from the ghcr.io registry under the latest tag, and the root package.json is marked private, so the registry image rather than a published package is the distribution channel. The toolchain is Bun, with Biome 2.5.14 for lint, format, and check, and a verify script that chains check, typecheck, test, and build. Two of the alternative install paths are third party: TrueNAS and Unraid community app catalogs, both maintained by outside contributors. The default branch is dev rather than main, and the newest tags are v1.7.0 on 2026-09-28, v1.6.5, and v1.6.4. The speed claim against FlareSolverr and Byparr is backed by selected same-machine benchmarks published on the project's own homepage, with no method described and an explicit note that results vary by site and session state.

The sponsor block is three scraping and proxy vendors with referral links

The top of the README is a sponsor table, and it lines up exactly with the top of the tier ladder. The fourth rung is an optional residential proxy, and the three sponsors are a scraping and unlocker platform, a general proxy provider, and a residential proxy provider for data collection. Each entry is linked through a vendor referral parameter rather than a plain vendor URL, and two of the three attach discount codes specifically for people using this project, so the page a reader reaches from the project carries commercial terms for the reader. That is worth weighing against the tier design, since the engine's own escalation logic moves traffic toward a paid residential network whenever a target resists. Nothing in the settings makes that step deliberate, and nothing in the settings names a vendor. Separately, the settings surface is not fully visible in the two files that configure it: the compose file's environment block ends at SCREENSHOT_JPEG_QUALITY with no value given, and the settings example ends on the bare prefix SCREENSHOT_. Anyone auditing this instance should read those values from the running container rather than assume the shipped file is the whole picture.

Editorial conclusion

trawl is well specified as an engine and poorly bounded as a policy, and those are different problems. Use it only for hosts where you hold written permission, on a machine you would be willing to have a CA certificate installed on, and with MCP and the metrics dashboard left off. Do not use it as a general purpose fetcher pointed at sites you have no relationship with: tier 4 hands the request to a residential proxy, the origin allowlist on /mcp is skipped by any client that omits an Origin header, and the compose file binds the API port on all interfaces while the dashboard guidance assumes loopback. Before you deploy anything, check four things yourself. Whether a per-host allowlist or dry-run switch exists in your version, since the published settings do not include one. What SCRAPE_MIN_TIER is set to on your install. Whether the port is bound to 127.0.0.1. And whether the performance comparison against FlareSolverr, hosted on the project's own site as selected same-machine benchmarks, applies to your targets, since the write-up itself says results vary by site and session state.

Frequently asked questions

What is germondai/trawl and what does it replace?

It is a self-hosted scraping engine in TypeScript licensed under AGPL-3.0 that presents itself as a FlareSolverr and Byparr alternative, exposing a FlareSolverr compatible /v1 endpoint for Prowlarr, Jackett, Sonarr, and other *arr tools. The newest tag is v1.7.0 dated 2026-09-28, and the default branch is dev.

What are the four execution tiers in trawl?

Plain HTTP fetch, cached browser session, fresh challenge solve, and an optional residential proxy, with escalation happening automatically when a wall is detected. SCRAPE_MIN_TIER sets the lowest tier allowed for every endpoint, where 1 is HTTP through 4 being the residential proxy.

Does trawl require a paid captcha solving service?

No. reCAPTCHA v2 audio can go through Google's free speech to text endpoint or an optional local Whisper service, and the listed captcha types include CF Turnstile and Interstitial, hCaptcha, GeeTest v4 Slide, ALTCHA, and Friendly Captcha v1 and v2.

How is trawl's MCP endpoint guarded?

MCP_ENABLED ships as false with a note to keep it disabled on public networks, and MCP_ALLOWED_ORIGINS is a comma separated list of browser origins allowed to call /mcp. Requests that carry no Origin header are accepted, so the list constrains browser origins rather than authenticating clients.

What does trawl's MITM forward proxy ask you to install?

With MITM_ENABLED set to true the proxy port is published and the CA certificate is served from port 8191 at /proxy-ca.crt. It has to be installed in each client's trust store, with the system keychain and Java cacerts named as examples.

How do you install trawl, and where does it keep request history?

Clone the repository, copy .env.example to .env, run docker compose up -d for the scraper and Redis, and check curl http://localhost:8191/health, with first boot taking 15 to 30 seconds. History goes to SQLite at /data/metrics/trawl.sqlite on a persistent named volume, and the dashboard is disabled by default.

Official sources

  1. germondai/trawl on GitHub
  2. License: AGPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/germondai-trawl.svg)](https://hysenlabs.com/projects/germondai-trawl)