# PixelRAG: Retrieval-Augmented Generation Over Web Screenshots

> PixelRAG renders web pages and documents as screenshot tiles and retrieves over those images directly, preserving tables, charts, and layout that text-based RAG discards. It ships a pre-built FAISS index of 8.28 million Wikipedia pages and a Claude Code plugin for in-editor visual browsing.

**StarTrail-org/PixelRAG** — https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/

- Repository: https://github.com/StarTrail-org/PixelRAG
- Website: https://arxiv.org/pdf/2606.28344
- Stars: 10,116 · Forks: 867
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/startrail-org-pixelrag

## What PixelRAG Solves and Who It Is For

Text-based RAG parses a web page or document to a string and discards the layout. A table becomes a flat list of values, a chart becomes nothing, and a formatted answer grid becomes unstructured text. The reader model then has to reconstruct what the original page made visually obvious.

PixelRAG takes a different approach: it renders the document to screenshot tiles first and then builds a vector index over those images. The official codebase accompanies the paper "PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation" (arxiv.org/abs/2606.28344) from Berkeley SkyLab, BAIR, and Berkeley NLP.

The primary users are developers building RAG systems over documents where visual structure carries meaning, researchers experimenting with vision-language retrieval, and teams using Claude or opencode who want a plugin that lets the AI see pages the way a human would.

## How PixelRAG Works: Rendering and Embedding

Two components make the system work. First, a rendering engine converts a document (web page, PDF, or image) into screenshot tiles using Playwright with Chrome DevTools Protocol (CDP). The `pixelshot` command handles this step.

Second, a vision-language embedding model converts those image tiles into vectors. According to the README, PixelRAG uses a `Qwen3-VL-Embedding` model LoRA-fine-tuned on screenshot data, which embeds page images into a space where visual content is searchable.

The README contrasts this with text-based RAG on a specific example: when a table holds the answer, text parsing discards the table structure and the reader cannot find the answer. PixelRAG retrieves the correct image tile and the reader reads the number directly off the image.

The pipeline stages are: render (pixelshot), chunk and embed, build index (FAISS), and serve. Each stage is a separate install extra, so a deployment that only searches a pre-built index does not need the embedding or training dependencies.

## Installing PixelRAG and Running the First Search

The core package installs with pip:

```bash
pip install pixelrag
```

This puts the `pixelshot` command and the `pixelrag` umbrella CLI on the path. The core install covers rendering and the basic CLI; heavy ML stages are opt-in extras.

To render a web page to screenshot tiles:

```bash
pixelshot https://en.wikipedia.org/wiki/Python --output ./tiles
```

The hosted endpoint serves a pre-built index of 8.28 million Wikipedia pages. No setup and no API key are required:

```bash
curl -X POST https://api.pixelrag.ai/search \
  -H "Content-Type: application/json" \
  -d '{"queries": [{"text": "What is the capital of France?"}], "n_docs": 5}'
```

The API also accepts an image as the query for visual search, as documented at pixelrag.ai/docs. To run a local search server over a downloaded index:

```bash
pip install 'pixelrag[serve]'
```

Then download the base FAISS index (~217 GB) from Hugging Face and start the server:

```bash
pixelrag serve --index-dir ./index/search_index_normed_v2 --port 30001
```

## The Claude Code Plugin and opencode Integration

PixelRAG ships a plugin called `pixelbrowse` for Claude Code. Instead of fetching raw HTML, Claude screenshots a page with `pixelshot` and reads the image, so it sees charts, diagrams, and tables as a person would.

The README recommends installing `pixelshot` with `uv tool` or `pipx` to keep it on `PATH` across all contexts; a plain `pip install` into a project virtual environment may leave `pixelshot` off `PATH` when Claude runs it:

```bash
uv tool install pixelrag
claude plugin marketplace add StarTrail-org/PixelRAG
claude plugin install pixelbrowse@pixelrag-plugins
```

After installation, you can ask Claude to look at any page:

```bash
claude -p "screenshot https://news.ycombinator.com and summarize the top stories"
```

For opencode users, the same tool is available as the `@startrail/pixelbrowse` npm package. Add it to `opencode.json`:

```json
{
  "$schema": "https://opencode.ai/config.json",
  "plugin": ["@startrail/pixelbrowse"]
}
```

The plugin requires `pixelshot` to be on `PATH`. No MCP server or backend is needed; the skill calls `pixelshot` directly on the local machine.

## Pipeline Commands and Install Extras

The `pixelrag` umbrella CLI exposes separate stages. The README documents which install extra each stage requires:

| Command | What it does | Install extra |
|---|---|---|
| pixelshot | Document to image tiles | pip install pixelrag (core) |
| pixelrag chunk / embed / build-index | Tiles to vectors to FAISS index | pixelrag[embed] |
| pixelrag index | Full pipeline orchestration | pixelrag[index] |
| pixelrag serve | FAISS search API (FastAPI) | pixelrag[serve] |

The `train` stage is a separate `uv` project inside `train/` with its own pinned environment (`torch==2.9.1+cu129`, `transformers==4.57.1`, cuDNN 9.20). The README explicitly states it must be installed from inside `train/`, not from the root. This is a meaningful constraint for anyone planning to fine-tune the embedding model: the training environment is not merged with the main package and has hard version pins.

## Limitations and What PixelRAG Does Not Cover

The image-based approach comes with real costs. The base FAISS index for Wikipedia alone is approximately 217 GB, and building a custom index from scratch requires the `[embed]` extra and a GPU that meets the cuDNN version requirements.

Rendering throughput depends on Playwright and Chrome CDP. The `twitter` extra, which adds Twitter as a source in related tooling, requires Playwright browsers and system packages that the current Dockerfile does not install, as the README notes directly.

For documents that are already plain text with no visual complexity, the added rendering and embedding work produces no benefit over text-based RAG. The cost in storage, compute, and pipeline complexity is only justified when visual structure in the source documents is what carries the answer.

The `qdrant` extra lists a Qdrant client as an alternative to FAISS for the vector store, but the README does not document this path in detail. Teams considering Qdrant for production deployments should review the `embed/` and `serve/` directories in the repository to understand the integration scope before committing to it. The project is licensed under Apache-2.0, which allows use in commercial applications without requiring source disclosure.

## Conclusion

PixelRAG is the right choice when the documents you need to retrieve contain visual structure that text parsers lose: tables, charts, infographics, or complex layouts. For plain text documents with no visual complexity, text-based RAG is simpler and avoids the GPU and storage overhead that the embedding and index stages require. The pre-built Wikipedia index is free and keyless, which makes evaluation easy. The training environment is pinned to specific torch and cuDNN versions and must be installed separately from the root; anyone planning to fine-tune should verify those hardware requirements before committing to the stack.

## FAQ

### What is PixelRAG?

PixelRAG is a Python library and pipeline that performs retrieval-augmented generation over screenshot images of documents rather than parsed text. It renders web pages and PDFs to image tiles, embeds them with a Qwen3-VL-Embedding model, and searches a FAISS index to retrieve visually relevant tiles.

### Does PixelRAG require a GPU to use?

The core install (pip install pixelrag) and the hosted Wikipedia endpoint require no GPU. The embedding and index-building stages (pixelrag[embed], pixelrag[index]) require a GPU, and the training environment pins specific torch and cuDNN versions. The pixelrag[serve] extra lists faiss-cpu as a dependency, so local search can run on CPU.

### Can PixelRAG search documents other than Wikipedia?

Yes. The pixelshot command renders any web page or PDF to screenshot tiles, and the pixelrag pipeline can build a FAISS index from any set of rendered tiles. The pre-built index covers Wikipedia; building a custom index requires the pixelrag[embed] or pixelrag[index] extras and sufficient storage.

## Sources

- [License: Apache-2.0](https://github.com/StarTrail-org/PixelRAG/blob/main/LICENSE)
- [Project website](https://arxiv.org/pdf/2606.28344)
- [README](https://github.com/StarTrail-org/PixelRAG/blob/main/README.md)
- [Releases](https://github.com/StarTrail-org/PixelRAG/releases)
- [StarTrail-org/PixelRAG on GitHub](https://github.com/StarTrail-org/PixelRAG)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/startrail-org-pixelrag
