PullMD: a self-hosted URL and file to Markdown service with an MCP endpoint
Self-hosted URL- and file-to-Markdown service for humans and AI agents - web pages, documents, images, audio, YouTube. PWA + REST + MCP + Claude Code skill, Reddit-aware, refreshable share links.
At a glance
- What is it?
- PullMD turns web pages, documents, images, audio and YouTube videos into Markdown from a single container, and exposes the same conversion over a REST API, a PWA and an MCP server. The v3 body change is the part that will bite existing self-hosters.
- Who is it for?
- Adopt PullMD if you want one container that answers both a browser and an MCP client with the same Markdown, and if you are willing to run the Playwright and Trafilatura sidecars for JavaScript-heavy pages. Stay away if you need a hosted service with no database to manage, or if you cannot accept AGPL-3.0 obligations in a closed product.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 10 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem PullMD solves, and who it is actually for
Feeding a web page to a language model usually means pasting rendered HTML and paying for navigation, cookie banners and footer links. PullMD's answer is a service that accepts a URL, or an uploaded file, and returns Markdown with the boilerplate stripped. The README frames it as a tool for "humans and AI agents", and the shipping surfaces match that split: a PWA for people, a REST endpoint at GET /api?url=… for scripts, an MCP server at POST /mcp for agent runtimes, and a Claude Code skill distributed as a downloadable zip.
The audience is narrow but real. If you already run a Docker host and want a single internal endpoint that both your browser and your agent tooling can call, PullMD is built for that. If you want a managed API where someone else handles the renderer, this is the wrong shape entirely: it is a container you operate, with a SQLite cache file and optional sidecars you also operate.
How extraction actually flows through the stack
PullMD is not one extractor, it is a ladder. Static HTML goes through Mozilla Readability and Trafilatura. Pages that need JavaScript are rendered by a headless Chromium sidecar (Playwright) before extraction. Cloudflare's native Markdown is used when the source offers it. Reddit and Hacker News threads get purpose-built converters, and the README states Reddit comment trees are captured in full, with subreddit, author, upvotes and published date placed in YAML frontmatter rather than the body.
The sidecars are separate services, not libraries inside the Node process. The bundled docker-compose.yml points the main container at TRAFILATURA_URL=http://trafilatura:8001/extract, PLAYWRIGHT_URL=http://playwright:8002/render and MARKITDOWN_URL=http://markitdown:8003/convert. That is a deliberate trade: the Node image stays small (node:22-alpine with better-sqlite3, cheerio, express, linkedom and friends), while the heavy rendering and document parsing live elsewhere and can be scaled or omitted. The cost is that you are now running a multi-container stack, and a sidecar that dies degrades extraction quietly. Version 3.10 added GET /api/status, which returns 503 when a configured sidecar stops responding, precisely because silent degradation was the failure mode.
Everything beyond plain web extraction is opt-in. Left unconfigured, the README says v3 handles web pages exactly like v2. Documents, images, audio, YouTube transcripts and the OCR PDF tier all require configuration before they do anything.
Installing PullMD with Docker and pulling your first page
The README's quick start assumes Docker and a published multi-arch image on Docker Hub. It needs no .env file, because the README states every variable has a default and the service listens on port 3000.
mkdir pullmd && cd pullmd
curl -O https://raw.githubusercontent.com/AeternaLabsHQ/pullmd/main/docker-compose.yml
docker compose up -d
# → http://localhost:3000After that, the PWA is at http://localhost:3000. Paste a URL into the interface and you should get Markdown back, with a rendered view, a raw view and a frontmatter toggle. The same conversion is available headlessly through the REST API documented as GET /api?url=…, which is the endpoint to script against if you are not using the browser.
The container writes its cache to a SQLite file. The bundled compose files mount ./data:/data and set CACHE_DB=/data/cache.db; for a local npm start the default is ./data/cache.db relative to the working directory. The Dockerfile creates /data and chowns it to the unprivileged app user, and docker-entrypoint.sh runs as root first to fix permissions on bind mounts the daemon may have created as root, then drops privileges with su-exec. That entrypoint detail matters if you bind-mount a directory that already exists on the host with unexpected ownership.
The v3 clean-body change is the one thing that will break your pipeline
The README is unusually direct about this: the Markdown body is now just a # Title plus content, and the source URL, fetch date and metadata moved into YAML frontmatter. Reddit posts follow the same rule. The README calls it "the one breaking change" and points to MIGRATION.md for a one-line opt-out.
If you have a downstream step that scrapes the "Source:" line out of the body, that step is now reading nothing. The escape hatch is PULLMD_SOURCE_HEADER=true, which restores the legacy inline header. PULLMD_FRONTMATTER_FIELDS lets you trim which frontmatter keys are emitted, which is the right lever if your consumer chokes on a field rather than on the format itself. Treat this as a schema migration for your own pipeline, not as a bug: the token savings are the point of the change, and reverting it gives that back.
Share links, cache retention and the retention trade-off
Every conversion gets an 8-hex share id. GET /s/:id returns the cached Markdown and re-fetches from the source if the cached copy is older than one hour. The README suggests using that id as a fixed URL for subreddit feeds and similar recurring sources, which is a genuinely different usage pattern from one-shot conversion: you are treating the service as a refreshable pointer rather than a converter.
Version 3.11 added PULLMD_CACHE_RETENTION_DAYS. The default is 90 days, and the .env.example explains the mechanics: every cache write prunes rows past that age, and a /s/:id link stops resolving once its row is older than the window, with the row deleted on the next cache write. Setting it to 0 keeps rows forever and turns the cache into an archive. The accepted range is 0 to 36500, and malformed values warn once at startup and fall back.
The trade-off is worth naming. An archive-shaped cache means your SQLite file grows without bound, and the growth is driven by whatever URLs people submit, including ones you did not intend to keep. The 90-day default is the safer posture for a shared instance. If you set 0, you have chosen to own that data indefinitely.
SSRF protection and the limits of a self-hosted converter
Version 3.3 added SSRF protection: private, loopback, link-local, CGNAT and cloud-metadata targets are rejected by default, on every fetch path and every redirect hop. That is the correct default for a service that fetches arbitrary URLs on request, and the fact that it covers redirect hops rather than only the initial URL is the part that matters, since a public URL can redirect to 169.254.169.254.
This is also where PullMD is the wrong tool. It is a URL fetcher with an allowlist-shaped safety net, not a sandbox. If your requirement is to convert untrusted input inside a network where the service can reach internal systems, the SSRF filter reduces the blast radius but does not remove the category of risk, and the README does not document a supported way to reach internal hosts. Similarly, if you need deterministic, offline conversion of a document you already hold, running a service that fetches URLs is more machinery than the job requires.
The other limitation is operational: the README does not document rollback or downgrade between versions. The cache is a SQLite file that persists, and MIGRATION.md covers the v2 to v3 body change, but a downgrade path is not described. Back up /data before upgrading if you care about the cached share links.
How PullMD differs from MarkItDown and other converters
MarkItDown is the obvious comparison, and the repository itself makes it: the compose file runs a markitdown-sidecar on port 8003 and the topics list includes markitdown. The difference is architectural rather than a matter of output quality. MarkItDown is a library and CLI you invoke on content you already have. PullMD is a long-running service that fetches URLs, caches results in SQLite, hands out 8-hex share ids, and speaks MCP. Document conversion in PullMD is routed through the MarkItDown sidecar, so the two are not really competitors at the conversion layer; the difference is everything around it.
Against a plain Readability script, the difference is the ladder. Readability alone gives up on JavaScript-rendered pages; PullMD falls back to headless Chromium. Against a hosted extraction API, the difference is that you own the data path and the cache, and you also own the sidecars, the volume and the upgrades.
Licence, upgrade cost and what to verify first
PullMD is AGPL-3.0, stated in the repository metadata, the LICENSE file and the package.json license field (AGPL-3.0-or-later). The Dockerfile labels the image with the same identifier. If you modify PullMD and let users interact with it over a network, the AGPL's network clause is the part to read with your own counsel; this is a description of the licence, not legal advice.
Upgrade cost is low if you track releases. The 3.x line has shipped frequently, and the changelog entries are specific: sidecar health endpoint and recipe append blocks in 3.10.0, the renderer no longer announcing itself as HeadlessChrome in 3.10.1, configurable cache retention in 3.11.0. The last push to main was on 2026-09-03.
Before you commit, verify three things. Read MIGRATION.md and decide whether PULLMD_SOURCE_HEADER belongs in your environment. Confirm your deployment mounts a writable /data volume, since that is where cache.db lives and where share links survive or do not. And decide whether you will run the Playwright and Trafilatura sidecars at all, because GET /api/status only alerts you about sidecars you have configured.
Editorial conclusion
Adopt PullMD if you want one container that answers both a browser and an MCP client with the same Markdown, and if you are willing to run the Playwright and Trafilatura sidecars for JavaScript-heavy pages. Stay away if you need a hosted service with no database to manage, or if you cannot accept AGPL-3.0 obligations in a closed product. Before committing, read MIGRATION.md for the clean-body change, check whether PULLMD_SOURCE_HEADER needs to be set, and confirm that your deployment path handles the /data volume that holds cache.db.
Frequently asked questions
How do I install PullMD?
The README's quick start downloads docker-compose.yml from the repository and runs docker compose up -d, after which the service is reachable at http://localhost:3000. No .env file is required because every variable has a default.
Does PullMD need an API key for YouTube transcripts?
The README states that YouTube transcripts, including title, description and clickable timecodes, are available with no API key required. The compose file does expose a MARKITDOWN_YOUTUBE variable, which is empty by default.
What is the PullMD MCP endpoint?
PullMD ships an MCP server at POST /mcp using Streamable-HTTP transport, and the README describes it as stateless. It sits alongside the REST API at GET /api?url=… and the PWA frontend.
Can I change how long PullMD keeps cached pages and share links?
Yes. PULLMD_CACHE_RETENTION_DAYS sets how long cache rows and share links live; the default is 90 days and 0 keeps them forever. The accepted range is 0 to 36500, and a malformed value warns once at startup and falls back.
Why did the Markdown body change in PullMD v3?
The README describes the clean body as the one breaking change: the body is now a title plus content, with the source URL, fetch date and metadata moved into YAML frontmatter to save tokens. Setting PULLMD_SOURCE_HEADER=true restores the old inline header.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/aeternalabshq-pullmd)