# Vision is All You Need: a V-RAG demo that skips chunking and embeds PDF pages as images

> Softlandia's repository is a working demo of Vision RAG, wiring pypdfium2, ColPali, Qdrant and GPT-4o behind a Modal-hosted FastAPI service with a React frontend. It is a reference architecture, not a library, and the README is honest about how thin the demo scaffolding is.

**Softlandia-Ltd/vision-is-all-you-need** — Serverless Modal + FastAPI + React + ColPali + Qdrant + GPT4o Vision RAG (V-RAG) Demo

- Repository: https://github.com/Softlandia-Ltd/vision-is-all-you-need
- Website: https://softlandia-ltd-prod--vision-is-all-you-need-web.modal.run/
- Stars: 404 · Forks: 52
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/softlandia-ltd-vision-is-all-you-need

## The chunking problem this demo sets out to remove

Most retrieval pipelines start by cutting documents into text chunks, and the quality of the whole system depends on where those cuts land. Split a table across two chunks and neither half retrieves well. Strip a chart and the surrounding prose loses its referent. The V-RAG approach in this repository avoids the decision entirely: the README states that the architecture uses a vision language model to embed pages of PDF files as vectors directly, "without the tedious chunking process". One page, one vector, no text extraction step in between.

The audience is engineers who already run a RAG stack and want to know whether page-level image retrieval is worth the storage and compute it costs. This is not a tool you install and point at a folder. It is a demonstration of an architecture, published alongside a blog post on the Softlandia site, and the repository is organized accordingly: a single main.py, a vrag/ package, a frontend/ directory, and a requirements.txt that pins the whole dependency set.

## How the V-RAG pipeline moves data from PDF to answer

The README lays out eight steps, and they divide cleanly into an indexing path and a query path.

On the indexing side, pypdfium2 converts PDF pages to images. The README notes that in theory these images can be anything, but the demo uses PDFs because the underlying model was trained on them. ColPali then produces the embeddings, and Qdrant stores them.

On the query side, the user's query goes through the same VLM to get a query embedding, that embedding searches Qdrant for similar page vectors, and the top matches come back as images. Those images plus the original query go to a model that understands images, GPT-4o or GPT-4o-mini in this demo, which generates the answer.

The design consequence is that retrieval operates on rendered pages rather than extracted text. A page containing a diagram, a two-column layout or a scanned table is represented by everything visible on it. The cost is that you store and search image embeddings, and the requirements file shows what that pulls in: colpali_engine, transformers, torch, einops, opencv_python_headless and vidore_benchmark alongside the web stack of fastapi, sse-starlette and pydantic. That is a GPU-oriented dependency tree, which is why the README points at Modal.

## Installing and running the demo against a real PDF

The README requires Python 3.11 or higher, a Hugging Face account, and an OpenAI API key. Authentication to Hugging Face is done with transformers-cli login before anything else, and the two secrets go into a dotenv file at the repository root:

```bash
OPENAI_API_KEY=
HF_TOKEN=
```

With those in place, the documented sequence is to install Modal, authenticate the CLI, and serve the application:

```bash
pip install modal
modal setup
modal serve \main.py
```

Modal prints a URL for the running app. Append /docs to it and the FastAPI-generated interface appears in the browser. From there the README's own walkthrough is: open POST /collections, click Try it out, upload a PDF, and click Execute. The documentation states that this indexes the file into an in-memory vector database and that it takes time depending on PDF size and the GPU allocated in Modal. The demo uses an A10G GPU.

Once indexing finishes, POST /search takes a query, sends the matched page images and the query to the OpenAI API, and returns the generated response. There is also a React frontend in frontend/: install Node.js, cd into the directory, set VITE_BACKEND_URL in .env.development, then npm install and npm run dev, which the README says starts it on http://localhost:5173.

## The in-memory vector store is the demo's real boundary

The most consequential line in the README is easy to skim past: indexing a PDF puts it into an in-memory vector database. That is fine for a demonstration and wrong for almost anything else. Restart the Modal container and the embeddings are gone. Two concurrent users are indexing into the same process memory. There is no described mechanism for collection management beyond the POST /collections endpoint that creates one from an upload.

It is also the wrong tool when your documents are already clean text. If you have well-structured Markdown or HTML, page-image embedding throws away the structure you paid to produce and replaces it with a rendered bitmap that a VLM has to interpret. Text chunking is cheaper, faster and easier to debug in that case. V-RAG earns its cost when layout carries meaning.

Two more gaps worth naming. The README does not document rollback or versioning of a collection, so re-indexing a corrected PDF means starting over. And the documented deployment path, modal deploy, publishes the FastAPI app to a Modal URL with no authentication step described, while every search call spends OpenAI credits. That combination deserves attention before anything is exposed beyond a laptop.

## How this differs from a text-first RAG stack

The obvious alternative is a conventional text RAG pipeline built on a chunker, a text embedding model and the same Qdrant instance. The difference is not the vector database, it is what gets embedded. A text pipeline runs a PDF parser, splits the extracted string on token boundaries, embeds each chunk, and stores the chunk text as payload so the language model receives words. This repository embeds rendered page images with ColPali and stores page images as what the answering model sees.

That changes three things. Retrieval granularity moves from chunk to page, which is coarser and can dilute a match that lives in one paragraph of a dense page. Storage grows because you keep images rather than strings. And the answering model must be multimodal, which is why the demo reaches for GPT-4o rather than a text-only completion endpoint. In exchange, you never write a parser, never tune chunk overlap, and never lose a figure to extraction. If your corpus is born-digital text, the trade is bad. If it is scanned reports, slide decks or manuals where the diagram is the answer, the trade is the point.

## Maintenance, licensing and what upgrading actually costs

The repository is MIT licensed, which permits commercial use and modification provided the copyright notice and permission notice are retained. That is the whole of the licence implication here; the LICENSE file is the authority, not this article.

The maintenance picture is thinner than the licence. There are no releases in the repository, so there is no versioned artifact to pin and no changelog describing breaking changes. Upgrading means tracking main. The last push was on 2026-09-10, which is recent, but a demo repository with no releases gives you no contract about what stays stable between commits.

The dependency surface is where upgrade cost concentrates. requirements.txt pins exact versions for the web and retrieval layers (fastapi[standard]==0.114.0, qdrant-client==1.11.1, colpali_engine==0.3.1, openai==1.44.1, pydantic==2.9.1) but leaves transformers at >=4.45.0 and torch unpinned entirely. ColPali support in transformers has moved quickly, so an unpinned transformers means a fresh install can resolve to a version the colpali_engine pin was not tested against. If you fork this, pin transformers and torch yourself.

## Conclusion

Adopt this if you want to see the V-RAG pipeline end to end before committing to it, or if you need a starting point for a page-level retrieval service where layout and figures matter more than text extraction. Do not adopt it as a production dependency: the README describes an in-memory vector database, there are no releases, and the code is a demo rather than a maintained library. Before building on it, verify two things in the repository: how vrag/ constructs and persists the Qdrant collection, and whether main.py exposes any authentication or rate limiting in front of the OpenAI calls, because the documented deployment path puts the API on a public Modal URL.

## FAQ

### What does vision-is-all-you-need mean as a project name?

It is the title Softlandia gave to a demo of the Vision RAG architecture, where a vision language model embeds PDF pages as images instead of embedding extracted text chunks. The README describes it as a demo of that architecture, published with a background blog post.

### What do I need before running the vision-is-all-you-need demo?

Python 3.11 or higher, a Hugging Face account logged in through transformers-cli login, and an OpenAI API key. Both secrets go into a dotenv file as OPENAI_API_KEY and HF_TOKEN.

### Which models and services does the vision-is-all-you-need pipeline use?

pypdfium2 converts PDF pages to images, ColPali produces the embeddings, Qdrant stores the vectors, and GPT-4o or GPT-4o-mini generates the final answer from the query and the retrieved page images.

### Where are the vectors stored in the vision-is-all-you-need demo?

The README states that indexing a PDF puts it into an in-memory vector database through the POST /collections endpoint. That means the index does not survive a container restart.

## Sources

- [Issues](https://github.com/Softlandia-Ltd/vision-is-all-you-need/issues)
- [License: MIT](https://github.com/Softlandia-Ltd/vision-is-all-you-need/blob/main/LICENSE)
- [Project website](https://softlandia-ltd-prod--vision-is-all-you-need-web.modal.run/)
- [README](https://github.com/Softlandia-Ltd/vision-is-all-you-need/blob/main/README.md)
- [Softlandia-Ltd/vision-is-all-you-need on GitHub](https://github.com/Softlandia-Ltd/vision-is-all-you-need)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/softlandia-ltd-vision-is-all-you-need
