Open-source project
StarTrail-org/PixelRAG avatar
StarTrail-org/PixelRAG

PixelRAG: Retrieval-Augmented Generation Over Web Screenshots

https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/

10,116 stars867 forksPythonApache-2.0

At a glance

What is it?
PixelRAG renders web pages and documents as screenshot tiles and retrieves over those images directly, preserving tables, charts, and layout that text-based RAG discards. It ships a pre-built FAISS index of 8.28 million Wikipedia pages and a Claude Code plugin for in-editor visual browsing.
Who is it for?
PixelRAG is the right choice when the documents you need to retrieve contain visual structure that text parsers lose: tables, charts, infographics, or complex layouts. For plain text documents with no visual complexity, text-based RAG is simpler and avoids the GPU and storage overhead that the embedding and index stages require.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What PixelRAG Solves and Who It Is For

Text-based RAG parses a web page or document to a string and discards the layout. A table becomes a flat list of values, a chart becomes nothing, and a formatted answer grid becomes unstructured text. The reader model then has to reconstruct what the original page made visually obvious.

PixelRAG takes a different approach: it renders the document to screenshot tiles first and then builds a vector index over those images. The official codebase accompanies the paper "PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation" (arxiv.org/abs/2606.28344) from Berkeley SkyLab, BAIR, and Berkeley NLP.

The primary users are developers building RAG systems over documents where visual structure carries meaning, researchers experimenting with vision-language retrieval, and teams using Claude or opencode who want a plugin that lets the AI see pages the way a human would.

How PixelRAG Works: Rendering and Embedding

Two components make the system work. First, a rendering engine converts a document (web page, PDF, or image) into screenshot tiles using Playwright with Chrome DevTools Protocol (CDP). The `pixelshot` command handles this step.

Second, a vision-language embedding model converts those image tiles into vectors. According to the README, PixelRAG uses a `Qwen3-VL-Embedding` model LoRA-fine-tuned on screenshot data, which embeds page images into a space where visual content is searchable.

The README contrasts this with text-based RAG on a specific example: when a table holds the answer, text parsing discards the table structure and the reader cannot find the answer. PixelRAG retrieves the correct image tile and the reader reads the number directly off the image.

The pipeline stages are: render (pixelshot), chunk and embed, build index (FAISS), and serve. Each stage is a separate install extra, so a deployment that only searches a pre-built index does not need the embedding or training dependencies.

Installing PixelRAG and Running the First Search

The core package installs with pip:

bash
pip install pixelrag

This puts the `pixelshot` command and the `pixelrag` umbrella CLI on the path. The core install covers rendering and the basic CLI; heavy ML stages are opt-in extras.

To render a web page to screenshot tiles:

bash
pixelshot https://en.wikipedia.org/wiki/Python --output ./tiles

The hosted endpoint serves a pre-built index of 8.28 million Wikipedia pages. No setup and no API key are required:

bash
curl -X POST https://api.pixelrag.ai/search \
  -H "Content-Type: application/json" \
  -d '{"queries": [{"text": "What is the capital of France?"}], "n_docs": 5}'

The API also accepts an image as the query for visual search, as documented at pixelrag.ai/docs. To run a local search server over a downloaded index:

bash
pip install 'pixelrag[serve]'

Then download the base FAISS index (~217 GB) from Hugging Face and start the server:

bash
pixelrag serve --index-dir ./index/search_index_normed_v2 --port 30001

The Claude Code Plugin and opencode Integration

PixelRAG ships a plugin called `pixelbrowse` for Claude Code. Instead of fetching raw HTML, Claude screenshots a page with `pixelshot` and reads the image, so it sees charts, diagrams, and tables as a person would.

The README recommends installing `pixelshot` with `uv tool` or `pipx` to keep it on `PATH` across all contexts; a plain `pip install` into a project virtual environment may leave `pixelshot` off `PATH` when Claude runs it:

bash
uv tool install pixelrag
claude plugin marketplace add StarTrail-org/PixelRAG
claude plugin install pixelbrowse@pixelrag-plugins

After installation, you can ask Claude to look at any page:

bash
claude -p "screenshot https://news.ycombinator.com and summarize the top stories"

For opencode users, the same tool is available as the `@startrail/pixelbrowse` npm package. Add it to `opencode.json`:

json
{
  "$schema": "https://opencode.ai/config.json",
  "plugin": ["@startrail/pixelbrowse"]
}

The plugin requires `pixelshot` to be on `PATH`. No MCP server or backend is needed; the skill calls `pixelshot` directly on the local machine.

Pipeline Commands and Install Extras

The `pixelrag` umbrella CLI exposes separate stages. The README documents which install extra each stage requires:

| Command | What it does | Install extra | |---|---|---| | pixelshot | Document to image tiles | pip install pixelrag (core) | | pixelrag chunk / embed / build-index | Tiles to vectors to FAISS index | pixelrag[embed] | | pixelrag index | Full pipeline orchestration | pixelrag[index] | | pixelrag serve | FAISS search API (FastAPI) | pixelrag[serve] |

The `train` stage is a separate `uv` project inside `train/` with its own pinned environment (`torch==2.9.1+cu129`, `transformers==4.57.1`, cuDNN 9.20). The README explicitly states it must be installed from inside `train/`, not from the root. This is a meaningful constraint for anyone planning to fine-tune the embedding model: the training environment is not merged with the main package and has hard version pins.

Limitations and What PixelRAG Does Not Cover

The image-based approach comes with real costs. The base FAISS index for Wikipedia alone is approximately 217 GB, and building a custom index from scratch requires the `[embed]` extra and a GPU that meets the cuDNN version requirements.

Rendering throughput depends on Playwright and Chrome CDP. The `twitter` extra, which adds Twitter as a source in related tooling, requires Playwright browsers and system packages that the current Dockerfile does not install, as the README notes directly.

For documents that are already plain text with no visual complexity, the added rendering and embedding work produces no benefit over text-based RAG. The cost in storage, compute, and pipeline complexity is only justified when visual structure in the source documents is what carries the answer.

The `qdrant` extra lists a Qdrant client as an alternative to FAISS for the vector store, but the README does not document this path in detail. Teams considering Qdrant for production deployments should review the `embed/` and `serve/` directories in the repository to understand the integration scope before committing to it. The project is licensed under Apache-2.0, which allows use in commercial applications without requiring source disclosure.

Editorial conclusion

PixelRAG is the right choice when the documents you need to retrieve contain visual structure that text parsers lose: tables, charts, infographics, or complex layouts. For plain text documents with no visual complexity, text-based RAG is simpler and avoids the GPU and storage overhead that the embedding and index stages require. The pre-built Wikipedia index is free and keyless, which makes evaluation easy. The training environment is pinned to specific torch and cuDNN versions and must be installed separately from the root; anyone planning to fine-tune should verify those hardware requirements before committing to the stack.

Frequently asked questions

What is PixelRAG?

PixelRAG is a Python library and pipeline that performs retrieval-augmented generation over screenshot images of documents rather than parsed text. It renders web pages and PDFs to image tiles, embeds them with a Qwen3-VL-Embedding model, and searches a FAISS index to retrieve visually relevant tiles.

Does PixelRAG require a GPU to use?

The core install (pip install pixelrag) and the hosted Wikipedia endpoint require no GPU. The embedding and index-building stages (pixelrag[embed], pixelrag[index]) require a GPU, and the training environment pins specific torch and cuDNN versions. The pixelrag[serve] extra lists faiss-cpu as a dependency, so local search can run on CPU.

Can PixelRAG search documents other than Wikipedia?

Yes. The pixelshot command renders any web page or PDF to screenshot tiles, and the pixelrag pipeline can build a FAISS index from any set of rendered tiles. The pre-built index covers Wikipedia; building a custom index requires the pixelrag[embed] or pixelrag[index] extras and sufficient storage.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. StarTrail-org/PixelRAG on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/startrail-org-pixelrag.svg)](https://hysenlabs.com/projects/startrail-org-pixelrag)