reader ships two packages with different method names, and one wrong default
Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.
At a glance
- What is it?
- A scraping engine built on Playwright with fingerprint injection and a browser pool, offered both self-hosted and hosted. The hosted client is not the local client with a key added, and the pool example contradicts its own comment.
- Who is it for?
- Use reader if your agent needs the web and you would rather not maintain browser pooling, fingerprint injection and proxy rotation yourself, and pick the self-hosted package if your traffic patterns or compliance requirements rule out a hosted service. Before you write against it, read the hosted client separately, because its method names and result shapes differ from the local ones rather than extending them.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 47 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The hosted client is not the local client with a key
There are two packages and they do not present the same interface. The self-hosted one constructs a client with no arguments and calls a scrape method taking a list of URLs, reading the markdown off an indexed array of results. The hosted one is installed under a different package name, constructs the client with an API key, and calls a read method taking a single URL. Its result is a tagged union: you check the kind field for scrape before reading the markdown off the data. So the differences are four. Different package, different method name, different argument shape, different result shape. None of that is a problem if you pick a path and stay on it. It is a problem if you expect to move between them, because a hosted key does not turn the local call into a hosted call, and the documentation does not present them as the same surface with an optional credential. Installing either is the same two steps:
npm install @vakra-dev/readernpx playwright install chromiumThe second command is a first-run step rather than part of the install, because the automation driver bundles the browser per platform but does not place it on disk by itself. Node 18 or newer is the stated floor. The hosted route installs its own package under a different name and needs no browser on your machine at all, which is the real difference between the two paths rather than the method names.
The TLS fingerprint is generated and injected, not inherited
The problem the tool claims to solve is stated as a table row rather than as prose: real browsers have fingerprints, the automation driver does not, and sites know. The feature list answers with TLS fingerprinting, navigator spoofing, WebRTC masking and a cleared automation flag. Two of the runtime dependencies exist for exactly that, a fingerprint generator and a fingerprint injector, pinned to matching versions. The distinction matters for what you are trusting. Inheriting a fingerprint from a real browser means the browser supplies it. Generating and injecting one means the engine synthesises it, which shifts the maintenance burden onto you: whenever a site changes how it checks, the generator library has to change with it, and nothing in a scraper can guarantee parity with a real browser on every version. That is the trade the engine makes on your behalf, and it is why it describes itself as built from the ground up rather than as a wrapper.
The shipped environment example sets a pool size its own comment denies
The example environment file annotates every value with its default, and one annotation disagrees with the value above it. The pool size line sets five browser instances while its comment states the default is two. Every other line in that block matches: the page-load retirement limit is set to a hundred and documented as a hundred, the time-based limit is set to thirty minutes and documented as thirty minutes, the maximum queue depth is a hundred and documented as a hundred, and the queue timeout is two minutes and documented as two minutes. So the inconsistency is isolated to one number, and it is the one number that determines how much memory the service holds at once. Copying the example as shipped roughly doubles the intended browser count. Worth noting that the daemon reads proxy configuration from the environment at startup, which means these values are also fixed at process start rather than being live-tunable.
Two proxy tiers, and the expensive one is aimed at a named list
The proxy design is a genuine two-tier choice rather than a list of interchangeable endpoints, and the example file explains what each is for. The datacenter tier is fast and cheap and covers most sites. The residential tier is slow and expensive and is reserved for aggressive anti-bot targets, with three named examples in the comment. The API reflects the same split with a proxy type of premium. So the cost difference is something you choose per request rather than something the engine hides, and the second tier exists to cover hosts where the cheap tier is known not to work. This is also the part of the system most likely to need a commercial proxy provider, since a residential pool is not something you run yourself. The README does not say where those addresses come from or how they are billed, so that is a question to answer before relying on the premium tier.
Two independent retirement limits sit in front of a bounded queue
Browser recycling is governed by two limits that trigger independently: a count of page loads and an elapsed time. A browser is retired after either threshold, so a long-lived idle instance is not held forever and a busy one is not left running indefinitely. Both thresholds are configurable and both are documented with their defaults in the example file, which as noted mostly agrees with itself. In front of the pool sits a queue with a maximum depth and a timeout. That pair is the graceful degradation story: once the queue is full, further work waits up to a bounded time and then fails, rather than accumulating without limit or blocking a caller indefinitely. The feature list groups this together with rate limiting, retries and caching, which is a fair description, but the queue is the piece that actually determines what a client experiences when you are overloaded.
Control is a remote endpoint, not a launched browser
The escape hatch from the markdown interface is a browser session, and the shape of it is what makes adoption cheap. Calling the browser method returns a session carrying a debugging-protocol WebSocket address rather than a browser object. You then connect an existing Playwright script over that address, change one line, create a context and a page, and drive it exactly as before, with the stealth settings already applied to the browser you connected to. With the other driver the change is the constructor argument, passing the endpoint where a WebSocket address belongs. So the cost of the stealth layer is that you are no longer launching a local browser; the cost of not using it is that you only get markdown. For an agent that needs to fill a form or click through something, that trade is usually worth one line. The anti-detection settings are described as active for the whole session rather than per call.
The tree is two patch versions ahead of its newest tag
The manifest declares version 0.3.4 while the newest release tag is 0.3.2, so the working tree is ahead of anything you can download. The recent history is compressed: three releases in five days in the last week of June, then nothing since the start of July. That pattern is consistent with a burst of work followed by a quiet period rather than with a monthly cadence. The packaging tells you what kind of project this is. It is an ES module with a built entry point and a published type definition, plus a command line binary wired to the built CLI. The build is a bundler, the type check is separate, tests run on a dedicated runner, and there are separate lint, format and licence-licence-check scripts over the source. Publishing runs a clean and a fresh build first, which is why the tree being ahead of its tag matters: what you get from the registry is not what the repository contains.
Editorial conclusion
Use reader if your agent needs the web and you would rather not maintain browser pooling, fingerprint injection and proxy rotation yourself, and pick the self-hosted package if your traffic patterns or compliance requirements rule out a hosted service. Before you write against it, read the hosted client separately, because its method names and result shapes differ from the local ones rather than extending them. And check the pool size in the shipped environment example, which sets five instances against a documented default of two and will roughly double your memory use if you copy it as-is.
Frequently asked questions
How do I install reader for scraping?
Install the self-hosted package with npm, and then install the bundled browser binary with the Playwright installer command, since the automation driver does not fetch it automatically. Node 18 or newer is required.
What does reader use to bypass bot detection?
TLS fingerprinting, navigator spoofing, WebRTC masking and clearing the automation flag. The fingerprint is generated and injected by two runtime dependencies rather than inherited from a real browser.
Is there a hosted version of reader?
Yes, under a different package name with an API key. Its interface differs from the self-hosted one: it uses a read method taking a single URL and returns a tagged result whose kind you check before reading the markdown.
How many browsers does the reader pool use by default?
The documented default is two. The shipped environment example sets five, which is the one value in that file that contradicts its own comment, and copying it as-is roughly doubles the intended browser count.
How do I control a stealthed browser with Playwright?
Ask the client for a browser session, which returns a debugging-protocol WebSocket address rather than a browser object, then connect over that address instead of launching. Everything else in an existing script stays the same and the stealth settings are already active.
What is the difference between the standard and premium proxy tiers?
The datacenter tier is fast and cheap and covers most sites. The residential tier is slow and expensive and is meant for aggressive anti-bot targets, with the environment example naming a few. The API selects between them per request with a proxy type.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vakra-dev-reader)