AnakinScraper OSS: a self-hosted Go scraping API that falls back from HTTP to an anti-detect browser
Open-source web scraping API. Turn any website into clean markdown or structured JSON. Anti-detect browser, proxy auto-selection, self-hosted. One command: make up
At a glance
- What is it?
- AnakinScraper OSS turns pages into markdown or Gemini-extracted JSON behind a REST API you run yourself. The handler chain, the Camoufox step and the AGPL-3.0 licence are the parts that decide whether it fits your stack.
- Who is it for?
- Adopt AnakinScraper OSS if you want a scraping endpoint you operate yourself, your stack can carry a Go 1.25 server plus a Camoufox container for JavaScript-heavy pages, and AGPL-3.0 fits how you ship. Do not adopt it if you need a managed service, cannot run Docker for the browser step, or intend to resell scraping as a closed product.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 42 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AnakinScraper OSS actually replaces
The project describes itself as an open-source web scraping API that turns a website into markdown or structured data, aimed at RAG pipelines and AI agents. The concrete job it removes is the glue code between fetching a page and getting clean text into a model. Instead of writing a fetch loop, a readability pass and a retry policy per site, you send a URL to a local endpoint and read `.markdown` back.
The audience is narrow in a useful way. It is for engineers who already have somewhere to run a container or a Go binary and who want the scraping step to stay inside their own network. The README states there are no cloud dependencies and no Redis, no AWS and no message queues. That matters if your data cannot leave your infrastructure, or if you want the scraping layer to survive a vendor changing its pricing.
It is not a crawler you point at a whole domain and walk away from. The API takes a URL, or a batch of up to 10 URLs. Discovery, scheduling and deduplication are your problem.
The handler chain: HTTP first, browser second, API last
The mechanism the README leads with is a handler chain with fallback. Each handler tries in order, and if one fails the next picks up. The documented order is HTTP fetch, then the anti-detect browser, then an external API. The README claims most pages resolve on the free local HTTP handler and that paid APIs are only called for roughly 5% that need them. That figure is the project's own estimate, not a measured result, and it will move with your target list.
The browser step is Camoufox, described as an anti-detect Firefox with realistic fingerprints rather than headless Chrome. In the Docker stack it runs as its own container exposing a WebSocket, and the server reaches it through `BROWSER_WS_URL`. That separation is the design choice worth noting: the browser is a service, not a library linked into the Go binary, so you can run the server without it and lose only the JavaScript-heavy path.
Proxy selection is the other moving part. The README states that Thompson Sampling picks the best proxy per domain, learning from success and failure in real time, and the web dashboard exposes a Proxy Scores page for it. This is the part most likely to disappoint in a small deployment, because the algorithm needs traffic per domain before its choices mean anything. With one proxy configured, the scoring has nothing to choose between.
Failure detection is configured per domain rather than globally. You define failure patterns and required patterns, and if scraped content matches a failure pattern (the README gives a CAPTCHA page as the example) or misses a required pattern, the job retries with the next handler. This is a better fit for sites you scrape repeatedly than for one-off URLs, since the patterns have to be written by hand.
Installing AnakinScraper OSS and scraping your first URL
There are two documented paths. The zero-config path needs only Go 1.25 or later and no database. From a clone of the repository, the README gives this command to start the server, which stores jobs in memory and loses them on restart:
cd server && go run cmd/server/main.goWith that running, the README's example request posts a URL to the synchronous endpoint on port 8080 and pipes the response through `jq` to print only the markdown field:
curl -s -X POST http://localhost:8080/v1/scrape \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com"}' | jq .markdownYou should get clean markdown for the page back in the response body, with no polling step. The default timeout is 30 seconds and the request field `timeout` can raise it, up to a maximum of 120 seconds.
The full stack is the Docker path. The Makefile defines `up` as `docker compose up -d` followed by a printed summary of the API and health URLs, so the documented sequence is a clone and one command:
git clone https://github.com/Anakin-Inc/anakinscraper-oss.git
cd anakinscraper-oss
make upThree containers come up: the server on 8080, the browser service on 9222, and PostgreSQL on 5432. The compose file also maps the browser service's health check to port 8090 and sets `WORKER_POOL_SIZE` to 5 and `JOB_TIMEOUT` to 120 on the server. The server waits on both dependencies being healthy before it starts, so a failing browser container will hold the whole stack back.
For structured output rather than markdown, set `GEMINI_API_KEY` and add `generateJson` to the request. The README's example posts the same URL with that flag set to true and reads `.generatedJson` from the response. The README states no API keys are required for the scraper itself, so this step is optional and separate from the scraping path.
If you want the dashboard, it is a separate front end. The README gives `cd webapp && npm install && npm run dev`, then http://localhost:3000, with API calls proxied to the server on 8080. Its pages cover health and quick scrape, sync and async scraping, job history, domain config CRUD, and proxy scores.
Where the defaults are dangerous, and where it is the wrong tool
The configuration file is blunt about the open-instance case. `API_KEY` is empty by default, and the comment states that when it is set, every `/v1` route requires it as `X-API-Key` or an `Authorization: Bearer` header, and that leaving it empty is only safe when the port is not reachable by anything you do not trust. Domain config writes always require a key regardless. If you expose port 8080 before setting `API_KEY`, you have published a scraper that anyone can drive from your IP address.
`CORS_ALLOW_ORIGINS` defaults to the webapp dev server, and the comment warns that `"*"` lets any website a browser visits drive the instance. `ALLOW_PRIVATE_TARGETS` defaults to false, and the comment explains that scrape targets resolving to loopback, private or link-local addresses, including the cloud metadata endpoint at 169.254.169.254, are rejected with a 400. That guard is correct for a shared deployment and inconvenient for an internal one: if your legitimate targets live on a private network, you have to flip the flag and accept the SSRF exposure that comes with it.
State is the other boundary. The Go-only path keeps jobs in memory, and the README says they are lost on restart. Persistence requires setting `DATABASE_URL`. There is no documented queue, so a restart mid-batch is a restart mid-batch.
As a tool choice, this is the wrong fit when you need a managed endpoint with an SLA, when you cannot run Docker and your targets are JavaScript-heavy, or when your product is itself a closed-source scraping service. The AGPL-3.0 licence is the reason for that last case, and it is covered below.
AnakinScraper OSS compared with Scrapy and Crawlee
The README's own comparison table puts AnakinScraper next to Firecrawl, Crawlee and Scrapy, and the meaningful difference is where the complexity sits. Scrapy is a Python framework: you write spiders, and the framework gives you scheduling, concurrency and pipelines. There is no anti-detect browser in its column and no automatic proxy selection, so both are things you assemble. Crawlee is the Node.js option, built on Playwright, with proxy handling described as manual in the same table.
AnakinScraper inverts that. You do not write a spider per site. You run a server and post URLs to it, and the per-site logic lives in domain configs that select handlers, set timeouts and retries, add headers, block domains and define the failure and required patterns. The trade is flexibility for uniformity: a Scrapy spider can express almost any traversal, while a domain config can only express what the documented keys allow. If your extraction needs to follow links and build a graph, this is not the layer for it.
Firecrawl is the closer comparison, since the table lists it as also returning markdown for LLM consumption. The stated differences are the browser (Camoufox against headless Chrome), the proxy strategy (Thompson Sampling against round-robin), the runtime (a Go binary against Node.js) and the fact that AnakinScraper's chain can fall through to an external API. The README's zero-config row is the sharpest distinction it draws: `go run` with no database, against Docker required for Firecrawl. Treat that row as the project's claim rather than a verified comparison.
Licence, telemetry and what an upgrade costs
The repository is AGPL-3.0, and the licence file is at the top level alongside a NOTICE file. The practical implication, without giving legal advice, is that the copyleft terms attach to the network service case: if you modify the software and let users interact with it over a network, the AGPL's source-availability condition is the one your legal team will want to read. Using it internally to feed your own RAG pipeline is a different situation from offering a scraping endpoint to customers. This is the single fact most likely to rule the project out for a commercial team, and it is worth resolving before any engineering time is spent.
Telemetry is on by default and documented in a separate TELEMETRY.md. The `.env.example` comment states that anonymous usage data is collected and that setting `TELEMETRY=off` disables it. The v0.1.1 release is titled "Anonymous Telemetry", so the mechanism arrived in the second release, after the initial open-source drop in v0.1.0. If your environment forbids outbound calls, that is a configuration change rather than a blocker, but it is not something you discover by reading only the API docs.
The upgrade surface looks small. Two releases exist, both from March 2026, and the last push to the repository was on 2026-08-18. The stack is a Go server, a Python browser service and a React webapp, so a full-stack upgrade means tracking three dependency trees rather than one. The Go-only mode is the cheapest thing to keep current, since it needs nothing but a Go toolchain. The CHANGELOG.md at the top level is where release notes live, and it is the file to read before pulling a new tag.
Editorial conclusion
Adopt AnakinScraper OSS if you want a scraping endpoint you operate yourself, your stack can carry a Go 1.25 server plus a Camoufox container for JavaScript-heavy pages, and AGPL-3.0 fits how you ship. Do not adopt it if you need a managed service, cannot run Docker for the browser step, or intend to resell scraping as a closed product. Before committing, verify that the browser-service container actually starts on your host, that your target domains do not trip the SSRF guard, and that the handler chain resolves them without falling through to a paid API key.
Frequently asked questions
How do I install AnakinScraper OSS?
Either run the server directly with Go 1.25 or later from the server directory, or clone the repository and run make up for the full Docker stack with the server, Camoufox browser service and PostgreSQL. The README gives both paths.
Does AnakinScraper OSS need an API key to scrape?
No. The README states that no API keys are required for the scraper itself. Keys are optional for the Gemini structured extraction step and for the anakin.io fallback handler, and the server's own API_KEY setting is what protects your instance.
Where does AnakinScraper OSS store scraped jobs?
In memory by default, and the README says those jobs are lost on restart. Setting DATABASE_URL switches storage to PostgreSQL, which the Docker stack provides on port 5432.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/anakin-inc-anakin)