Model or dataset
jina-ai/reader avatar
jina-ai/reader

Jina AI Reader: URL-to-Markdown API for LLM Pipelines

Convert any URL to an LLM-friendly input with a simple prefix https://r.jina.ai/

12,070 stars891 forksTypeScriptApache-2.0

At a glance

What is it?
Jina AI Reader is an open-source TypeScript service that converts any URL to LLM-friendly markdown by prepending r.jina.ai/ to the URL. It handles JavaScript-rendered pages, PDFs, Office documents, and images, and provides a companion search endpoint at s.jina.ai that fetches and converts the top web results for a query.
Who is it for?
Jina AI Reader is a practical fit for engineers building RAG pipelines, web agents, or LLM tools that need to consume arbitrary web content. The free hosted API at r.jina.ai is the lowest-friction starting point: no setup required, just prepend the URL.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 131 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Jina AI Reader Does and Who It Is For

Jina AI Reader converts web content into clean, structured markdown suitable for passing to large language models. The core interface is simple: prepend https://r.jina.ai/ to any URL. For example, r.jina.ai/https://en.wikipedia.org/wiki/Artificial_intelligence returns the Wikipedia article on artificial intelligence as markdown, stripped of navigation, ads, and sidebar content.

The intended users are engineers building RAG (retrieval-augmented generation) systems, LLM agents that need to read web pages, or AI pipelines that process arbitrary URLs as input. The README describes the project as providing "improved output for your agent and RAG systems at no cost."

Reader handles four document types: web pages rendered with headless Chrome or fetched with curl-impersonate, PDFs parsed with PDF.js, Microsoft Office documents (Word, Excel, PowerPoint) converted via LibreOffice, and images captioned by a vision-language model so text-only LLMs get descriptive hints about image content.

The r.jina.ai service is free, and the README describes it as stable and scalable. Rate limits are documented at jina.ai/reader#pricing.

The Search Endpoint: s.jina.ai

Alongside the URL reading API, Jina AI Reader provides a web search endpoint at s.jina.ai. Prepend https://s.jina.ai/ to a search query to get back the content of the top 5 web results, already converted to markdown.

The behavior differs from a standard search API that returns titles, URLs, and snippets. According to the README, s.jina.ai searches the web, fetches the top 5 results, visits each URL, and applies r.jina.ai to it. This means the result is the full converted content of each page rather than search-engine-provided metadata.

For in-site searches, specify a site parameter:

bash
curl 'https://s.jina.ai/When%20was%20Jina%20AI%20founded%3F?site=jina.ai&site=github.com'

This restricts results to pages from jina.ai and github.com. The search queries should be URL-encoded before being appended to the s.jina.ai prefix.

The README describes s.jina.ai as launching in May 2024, adding the search capability to the earlier URL reading service.

Output Formats and Request Headers

By default, r.jina.ai returns markdown. The x-respond-with header selects a different output format:

- markdown: returns markdown without going through Mozilla's Readability.js. - html: returns the page's full documentElement.outerHTML. - text: returns document.body.innerText. - screenshot: returns the URL of the page's viewport screenshot. - pageshot: similar to screenshot but attempts to capture the whole page. - frontmatter: returns Markdown with a YAML frontmatter block containing title, description, and URL. - markdown+frontmatter: like frontmatter but covers the full page without Readability filtering.

The frontmatter format is useful when the downstream LLM benefits from structured metadata:

bash
curl -H 'X-Respond-With: frontmatter' 'https://r.jina.ai/https://example.com'

Other useful headers include x-engine (enforce a specific fetching engine: browser, curl, or auto), x-proxy-url (route through a proxy), x-cache-tolerance (how stale a cached page is acceptable in seconds), and x-no-cache (bypass the cache entirely). The full list of supported headers is at r.jina.ai/docs or in src/dto/crawler-options.ts in the repository.

Architecture: Open Source Branch vs. SaaS

The repository is the open-source branch of the codebase behind the hosted r.jina.ai and s.jina.ai services. The README makes a clear distinction: the full SaaS service uses MongoDB Atlas as a storage layer. The open-source branch does not include this storage layer and runs in stateless mode or with bucket-cached mode using MinIO.

For local development, the docker-compose.yml provides a MinIO instance for S3-compatible caching. MinIO runs on ports 9000 and 9001, with the console accessible on port 9001. This optional caching layer stores previously fetched pages to avoid re-fetching them on repeated requests.

The repository was re-synchronized with the SaaS codebase in April 2026, as noted in the README's update history. Before that re-sync, the open-source branch had diverged significantly from the production code. The README notes that binary file uploads (PDFs and Office documents via the file body field) were added in December 2025, and the major migration from Firebase to Cloud Run with MongoDB happened in March 2025.

Node 22.15 or newer is required. The package.json version is 0.5.0.

Deploying the Service Locally

The Dockerfile builds a Node 24 image with Chrome or Chromium (depending on the target architecture: google-chrome-stable on amd64, chromium on arm64), LibreOffice for Office document conversion, and various font packages for CJK and emoji rendering. The image exposes port 8080 for the main service.

The build uses cargo-chef style layer caching through a multi-stage Dockerfile: a build stage installs system dependencies and installs Node packages, then a final stage copies the built output. The Docker image runs as the jina user rather than root.

For production environments, the README notes that the SaaS runs on Cloud Run. The self-hosted path from this repository runs in stateless mode by default, which means no page caching. Adding the MinIO docker-compose service provides bucket-based caching that is more appropriate for repeated crawling workloads.

Limitations and Cases Where Reader Is Not the Right Tool

The open-source branch does not include the MongoDB-backed storage layer. Any feature of the hosted service that depends on persistent storage, user accounts, or history is not present in a self-hosted deployment from this repository. If your use case requires tracking which URLs have been fetched or maintaining user-specific caches, the self-hosted version needs additional infrastructure.

JavaScript-heavy pages with anti-bot measures may not render correctly even with the browser engine. The auto mode selects between browser and curl-impersonate based on the page's requirements, but some sites specifically block headless browsers or require authentication.

PDF parsing with PDF.js works well for text-based PDFs but may lose structure in scanned PDFs or image-only PDFs, where the output becomes a description from the vision model rather than extracted text. LibreOffice conversion of complex Office documents may not preserve all formatting.

The service is designed for content extraction, not for rendering interactive web applications. Pages that require user interaction beyond initial load (infinite scroll, login gates, CAPTCHA) will return only the initially rendered content.

Jina AI Reader vs. Mozilla Readability

Mozilla Readability.js is a JavaScript library that extracts the main readable content from a web page. It is widely used in browser extensions and read-mode implementations. Jina AI Reader uses Readability internally (it appears as @mozilla/readability in the package.json dependencies) but adds significant capabilities beyond what Readability alone provides.

Readability.js operates on already-fetched HTML. It does not handle JavaScript rendering, PDF parsing, or Office document conversion. When you pass a Readability-processed page to an LLM, images are described only by their alt text. Jina AI Reader handles headless browser rendering for JavaScript pages, parses PDFs and Office documents, and passes images to a vision model for captions.

Readability is suitable for simple content extraction in a browser context where you already have the HTML. Jina AI Reader is suitable for pipelines that need to process arbitrary URLs server-side, including complex JavaScript apps, PDFs, and Office files, and return clean markdown regardless of the original format.

Editorial conclusion

Jina AI Reader is a practical fit for engineers building RAG pipelines, web agents, or LLM tools that need to consume arbitrary web content. The free hosted API at r.jina.ai is the lowest-friction starting point: no setup required, just prepend the URL. Self-hosting is available through the open-source repository, but requires Node 22.15+, Chrome or Chromium, LibreOffice, and optional MinIO for caching. The open-source branch does not include the MongoDB-backed storage layer used in the SaaS version, so session-based or account-based features are not present in self-hosted deployments. The last push to this repository was on 2026-05-22.

Frequently asked questions

How does Jina AI Reader handle JavaScript-rendered pages?

Reader uses headless Chrome or Chromium for pages that require JavaScript execution. The default x-engine mode is auto, which intelligently selects between the browser and curl-impersonate depending on the page's requirements. You can enforce browser-only rendering with x-engine: browser.

Can I self-host Jina AI Reader?

Yes. The repository provides a Dockerfile and docker-compose.yml for self-hosted deployment. The open-source branch runs in stateless mode or with optional MinIO bucket caching. It does not include the MongoDB-backed storage layer used in the SaaS version.

Is there a rate limit on the free r.jina.ai API?

The README states the service is free, stable, and scalable, and directs users to jina.ai/reader#pricing for rate limit details. The rate limit documentation is on the Jina AI pricing page rather than in the repository.

Official sources

  1. Issues
  2. jina-ai/reader on GitHub
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jina-ai-reader.svg)](https://hysenlabs.com/projects/jina-ai-reader)