Chunkr: self-hosted document parsing for RAG pipelines
Vision infrastructure to turn complex documents into RAG/LLM-ready data
At a glance
- What is it?
- Chunkr is an AGPL-3.0 Rust service that turns PDFs, slides, Word files and images into layout-aware chunks with OCR and bounding boxes. The open-source build runs community models; the hosted API does not.
- Who is it for?
- Adopt the open-source build if you need document parsing to stay inside your own network and you accept community OCR and VLM quality plus the AGPL-3.0 obligations. Do not adopt it if you need native Excel parsing, or if you want the accuracy of the hosted models, because the README states the open-source release uses community and open-source models while the Cloud API runs proprietary in-house ones.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Chunkr solves for retrieval pipelines
Most RAG systems fail at the ingestion step, not at the vector database. A PDF is not a stream of text; it is positioned glyphs, tables, columns and figures. Naive extraction returns reading order that jumps between columns and drops table structure entirely. Chunkr targets that gap. It is a service, not a library, that accepts a document and returns chunks with layout analysis, OCR, bounding boxes, and structured HTML or Markdown, according to the README. The audience is engineers building retrieval over messy business documents: contracts, slide decks, scanned forms, reports. The README positions the open-source repository for development and testing, and the hosted Cloud API for production workloads. That distinction matters more than any feature table, because it tells you the maintainers consider the self-hosted build a different accuracy tier, not just a different deployment model.
How the pipeline is wired in compose.yaml
The architecture is visible in compose.yaml, and it is a set of specialized HTTP services behind nginx rather than one monolith. A server container exposes port 8000 and depends on postgres, redis and minio. A task service is the same image scaled to 30 replicas by default, which is where processing work is distributed. Two backends do the heavy lifting: segmentation-backend reserves NVIDIA GPUs with MAX_BATCH_SIZE=4, BATCH_WAIT_TIME=0.2, OVERLAP_THRESHOLD=0.025 and SCORE_THRESHOLD=0.2, and ocr-backend is built from docker/doctr/Dockerfile. Each backend sits behind its own nginx proxy on ports 8001 and 8002. A web container serves the UI on port 5173. The task workers reach the backends through WORKER__GENERAL_OCR_URL and WORKER__SEGMENTATION_URL from .env.example, so the data flow is: upload to server, object stored in minio, job queued in redis, task worker calls segmentation then OCR, results assembled and written back. The 30 task replicas against 6 segmentation replicas is an explicit bet that segmentation is the slow stage.
Installing Chunkr with Docker Compose
The README documents Docker Compose as the supported path. Clone the repository, copy the environment template, copy the model configuration, then choose the compose files for your hardware. The GPU path is the plain compose.yaml. CPU-only deployments add compose.cpu.yaml, and Apple Silicon adds compose.mac.yaml on top of that.
git clone https://github.com/lumina-ai-inc/chunkr
cd chunkr
cp .env.example .env
cp models.example.yaml models.yamlAfter that, start the stack. Pick exactly one of these three commands; the README lists them as alternatives, not as steps to run in sequence.
# GPU
docker compose up -d
# CPU only
docker compose -f compose.yaml -f compose.cpu.yaml up -d
# Mac ARM
docker compose -f compose.yaml -f compose.cpu.yaml -f compose.mac.yaml up -dWhen the containers are healthy, the README says the web UI answers on http://localhost:5173 and the API on http://localhost:8000. The server container mounts your models.yaml read-only at /app/models.yaml, so editing that file and restarting is how you change model configuration. Shut down with the matching down command, for example docker compose down for the GPU stack.
Configuring an LLM through models.yaml
Chunkr does not ship a model; it calls one you provide. The README offers two mechanisms and recommends the file. The basic path is three environment variables in .env: LLM__KEY, LLM__MODEL and LLM__URL, which the README describes as working with any OpenAI-compatible endpoint. The recommended path is models.yaml, which supports several providers at once, a default and fallback model, and per-model rate limits. The README gives this example entry for gpt-4o with provider_url pointing at the OpenAI chat completions endpoint and a rate-limit of 200 requests per minute.
models:
- id: gpt-4o
model: gpt-4o
provider_url: https://api.openai.com/v1/chat/completions
api_key: "your_openai_api_key_here"
default: true
rate-limit: 200The id is what you reference in API requests, which is the point of the file: it decouples your callers from the underlying provider. The README points at models.example.yaml for the full option set, and that is the file to read before you assume a key exists. Note that .env.example sets LLM__MODELS_PATH to ./models.yaml, so the compose files expect that filename at the repository root.
What the open-source build does not do
The README is unusually direct about the gap between the repository and the hosted product, and it should shape your decision. Excel support is marked as absent in the open-source column and present in both the Cloud API and Enterprise columns, where it is described as a native parser. If your corpus includes spreadsheets, this build is the wrong tool. The same table lists layout analysis as open-source models versus proprietary in-house models, OCR as community OCR engines versus an optimized stack, and VLM processing as basic open VLMs versus enhanced proprietary ones. That is the maintainers telling you the self-hosted output will be less accurate. There is also an operational constraint in compose.yaml: segmentation-backend reserves NVIDIA devices, so the default stack assumes a GPU. The CPU and Mac variants exist, but the README does not state what throughput or accuracy difference to expect from them. Finally, the repository carries a CONTENT-MIGRATION-GUIDE.md and a COMMERCIAL_LICENSE.md at the top level, which suggests the licensing boundary is a live concern rather than a formality.
Chunkr compared with unstructured and LlamaParse
The closest alternatives are Unstructured and LlamaParse, and the difference is architectural. Unstructured is primarily a Python library you call inside your own process, with a hosted API as an option; you own the orchestration, the queueing and the scaling. Chunkr ships that orchestration as part of the product: postgres for state, redis for the queue, minio for object storage, nginx-fronted segmentation and OCR backends, and a task tier scaled to 30 replicas. LlamaParse sits at the other end, a hosted parsing API where you send a document and receive structured output, with no self-hosted path. So the real question is not which parses better in the abstract but where you want the queue to live. If you already run a document pipeline and want a function call, a library fits. If you want a running service with its own storage and worker pool that you operate, Chunkr's compose stack is that. If you want none of the operations, the hosted APIs are the answer, and Chunkr's own README recommends its Cloud API for production.
Licence and maintenance cost
Chunkr is AGPL-3.0. For internal use this is usually unremarkable. For a product where you modify Chunkr and expose it to users over a network, the AGPL's source-availability obligation is the thing your legal team will want to read, and the presence of COMMERCIAL_LICENSE.md and THIRD-PARTY-NOTICES.md in the repository indicates the maintainers offer an alternative track. This is not legal advice; read LICENSE and COMMERCIAL_LICENSE.md yourself. On maintenance, the last push to main was on 2026-09-09, and the most recent tagged release in the repository is v2.2.1 from 2025-07-31, alongside a chunkr-services-v0.1.6 tag published the same day. The gap between a September push and a July release means you should expect to track main if you want recent changes, not just tagged versions. Upgrading also has a cost specific to this project: models.yaml is mounted into the server and task containers, so a schema change to that file is an upgrade step, not a background detail.
Editorial conclusion
Adopt the open-source build if you need document parsing to stay inside your own network and you accept community OCR and VLM quality plus the AGPL-3.0 obligations. Do not adopt it if you need native Excel parsing, or if you want the accuracy of the hosted models, because the README states the open-source release uses community and open-source models while the Cloud API runs proprietary in-house ones. Before committing, verify three things on your own hardware: how the segmentation and OCR containers behave on your document mix, whether your chosen LLM endpoint satisfies the models.yaml configuration, and whether your legal team accepts AGPL-3.0 for your distribution model.
Frequently asked questions
What does "chunk" mean in AI?
In retrieval systems, a chunk is a passage of a source document small enough to embed and retrieve. Chunkr's contribution is producing those passages with layout analysis, OCR and bounding boxes rather than splitting raw extracted text.
What is a Chunkr alternative for self-hosted document parsing?
Unstructured is the closest alternative, but it is primarily a Python library you call in-process, so you own the queueing and scaling that Chunkr ships in its compose stack. LlamaParse is the opposite trade: a hosted API with no self-hosted path.
Is Chunkr open source?
Yes. The repository lumina-ai-inc/chunkr is licensed AGPL-3.0, and the README states the open-source release uses community and open-source models rather than the proprietary in-house models behind the hosted Cloud API.
Does Chunkr support Excel files?
Not in the open-source build. The README's comparison table marks Excel support as absent for the repository and present as a native parser in both the Cloud API and Enterprise editions.
Does Chunkr need a GPU to run?
The default compose.yaml reserves NVIDIA devices for the segmentation backend, so the GPU path is the primary one. The README also documents CPU-only and Mac ARM variants through compose.cpu.yaml and compose.mac.yaml, but it does not state the throughput difference.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lumina-ai-inc-chunkr)