Morphik Core: a multimodal retrieval engine for visually rich documents
Open-source multimodal retrieval engine (Morphik Core). By Morphik — AI back office for skilled nursing & senior living (morphik.ai).
At a glance
- What is it?
- Morphik Core is the open-source retrieval layer behind Morphik's back-office AI, offered to developers as a standalone platform. It targets PDFs, images and video with ColPali-style visual search, and ships under the Business Source License 1.1 rather than an OSI-approved licence.
- Who is it for?
- Adopt Morphik Core if your retrieval problem is genuinely visual: engineering drawings, scanned tables, chart-heavy reports where text extraction loses the answer. Stay away if you need predictable per-query cost, a fully supported self-hosted deployment, or an OSI-approved licence, because the README states self-hosted support cannot be guaranteed and the code ships under the Business Source License 1.1.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Morphik Core actually solves for visually rich documents
The README frames the problem bluntly: pipelines assembled from separate text extraction, OCR, embedding, vector store and retrieval components work in a proof of concept and then break under load. The specific failure it names is visual. Charts become meaningless text fragments, diagrams lose spatial relationships, tables turn into unreadable strings. If your documents are mostly prose, that argument does not apply to you and a conventional text-chunking pipeline will be cheaper and easier to reason about.
The target user is a developer building an AI application over documents where the answer lives in layout: assembly instructions, technical specifications with mixed text and visuals, scanned forms, video. Morphik Core exposes ingest, shallow and deep search, transform and management as one system rather than five, with a single endpoint for images, PDFs and videos. The README also points out a cost angle: re-processing the same 500-page manual for every query is what pushes inference bills up, which is the argument for retrieval happening before the LLM sees anything.
ColPali search, rules-based metadata and the ingestion worker
Two mechanisms are visible in the repository. The first is multimodal search built on ColPali, linked from the README under docs/concepts/colpali and demonstrated in examples/colpali.py. ColPali-style retrieval embeds page images rather than extracted text, which is why the project claims search that understands visual content. The practical consequence is that ingestion must render pages to images, and the dependency list reflects that: pymupdf, pdf2image, docling, weasyprint and torch all appear in pyproject.toml.
The second is rules-based ingestion, documented at morphik.ai/docs/concepts/rules-processing, for metadata extraction including bounding boxes, labeling and classification. Metadata is extracted during ingestion rather than inferred at query time.
The architecture is a FastAPI service plus a separate arq worker. docker-compose.yml defines two application containers: morphik, which builds from the repository root and maps port 8000, and worker, which runs arq core.workers.ingestion_worker.WorkerSettings. Ingestion is therefore asynchronous and queue-backed, with Redis as the broker and Postgres with pgvector for storage. The compose file also carries a config-check service that greps morphik.toml for provider = "ollama" and writes a flag file, so the Ollama container is only started when the config asks for it. That is a small but telling detail: local model serving is opt-in through morphik.toml, not a default.
Installing Morphik Core and running a first query
The README's recommended path is not self-hosting. It points to dev.morphik.ai/signup for a free tier, and the Python SDK example assumes you already have a Morphik URI from that account. The client package is morphik, pinned at 1.0.3 in pyproject.toml.
from morphik import Morphik
morphik = Morphik("<your-morphik-uri>")
morphik.ingest_file("path/to/your/super/complex/file.pdf")After ingest_file returns, the document is registered with the service. The README does not state whether ingestion is synchronous at this call site or whether the SDK polls the worker queue, so expect the answer to arrive asynchronously if you query immediately.
Querying is a single call, and the README's own example is deliberately visual: a question about the height of a screw in chair assembly instructions.
morphik.query("What's the height of screw 14-A in the chair assembly instructions?")For self-hosting, the README defers to dev.morphik.ai/docs/getting-started rather than reproducing steps, and the repository ships install_and_start.sh, install_and_start.ps1, install_docker.sh, install_docker.ps1 and start-dev.sh alongside docker-compose.yml. Before any of that, .env.example requires three values: JWT_SECRET_KEY, SESSION_SECRET_KEY and POSTGRES_URI.
JWT_SECRET_KEY="your-super-secret-key-change-in-production"
SESSION_SECRET_KEY="your-session-secret-key-change-in-production"
POSTGRES_URI="postgresql+asyncpg://morphik:morphik@localhost:5432/morphik"The same file notes that COMPOSE_PROJECT_NAME should be set once before the first Docker start to keep volume names stable, and warns against changing it on an existing deployment without migrating named volumes. Provider keys for OpenAI, Anthropic or Gemini are set separately, and only the ones you use. The README also mentions a web console for uploading files and chatting with the data, and an MCP integration documented at dev.morphik.ai/docs/using-morphik/mcp.
Where Morphik Core is the wrong choice
The README is unusually direct about support: due to limited resources, full support for self-hosted deployments is not provided. There is an installation guide and a Discord community, but no guarantee. For a team without the capacity to debug a FastAPI service, an arq worker, Postgres with pgvector and Redis, that sentence should decide the question. The hosted tier is the path the project itself recommends.
The second boundary is operational. ColPali-style visual retrieval means page images and torch in the dependency set, so the resource profile is not that of a text embedding service. The README does not publish memory or GPU requirements, and pyproject.toml does not constrain the compute target, so sizing has to come from your own deployment attempt rather than from documentation.
Third, the licence. The README describes Morphik Core as source-available under the Business Source License 1.1, with personal and indie use free and commercial production use free only if the deployment generates something the truncated README does not finish stating. The LICENSE file is the authority here, and it is the first thing to read rather than the marketing page.
Morphik Core against a conventional text RAG stack
The real alternative is the stack the README argues against: a text extractor such as Docling or PyMuPDF, a chunker, a text embedding model, and pgvector or another vector store, wired together by hand. The difference is not quality in the abstract, it is what gets embedded. A conventional stack embeds extracted text, so a chart's meaning is whatever the extractor produced, often a caption and a jumble of axis labels. Morphik Core embeds page images through ColPali, so the visual content is itself retrievable.
That choice has a cost the README does not quantify. Image embeddings are larger and slower to produce than text embeddings, and the pipeline carries torch and a document rendering step. A text-only corpus pays that cost for nothing. Conversely, if your PDFs are born-digital prose with clean text layers, the conventional stack is simpler to operate and easier to swap components in and out of, which matters when one vendor's embedding model changes under you.
A second comparison point is scope. Morphik Core bundles ingestion, storage, search, metadata extraction and integrations with Google Suite, Slack and Confluence into one service. That reduces integration work and increases lock-in to one project's abstractions, its config file and its worker model. The trade is real in both directions.
Licence terms and the cost of staying current
Morphik Core is licensed under the Business Source License 1.1, which the README explicitly calls source-available rather than open source. Personal and indie use is free. Commercial production use is free only under a condition the README begins to state and does not finish in the available text, so the LICENSE file is the only reliable source for what your deployment is allowed to do. This is not legal advice; if your company has a policy on source-available licences, that policy applies before you write any code.
The upgrade picture is also constrained by packaging. The repository pins morphik==1.0.3 as a dependency of morphik-core itself, and pins torch==2.8.0, torchaudio==2.8.0, torchvision==0.23.0 and pymupdf==1.25.5. Those exact pins mean an upgrade is not a version bump in one file; it is a coordinated change across the SDK, the ML runtime and the PDF renderer, and it has to be tested against your own documents because the README publishes no compatibility matrix. The last push to the default branch was on 2026-09-08, and there are no retrieved releases, so tracking changes means following the main branch rather than waiting for tagged versions. Running the compose stack also means maintaining Postgres, Redis and, if you enable it, Ollama, plus a Hugging Face model cache volume that docker-compose.yml mounts at /root/.cache/huggingface.
Editorial conclusion
Adopt Morphik Core if your retrieval problem is genuinely visual: engineering drawings, scanned tables, chart-heavy reports where text extraction loses the answer. Stay away if you need predictable per-query cost, a fully supported self-hosted deployment, or an OSI-approved licence, because the README states self-hosted support cannot be guaranteed and the code ships under the Business Source License 1.1. Before committing, verify two things yourself: whether your deployment's commercial use falls inside the free tier described in LICENSE, and whether the self-hosting guide at dev.morphik.ai/docs/getting-started matches your environment, including the Postgres, Redis and optional Ollama services that docker-compose.yml expects.
Frequently asked questions
Is Morphik Core free to use?
Personal and indie use is free under the Business Source License 1.1, and the README also points to a free tier on the hosted platform at dev.morphik.ai. Commercial production use is free only if your deployment meets a condition the README states but does not complete in the available text, so the LICENSE file is the place to confirm it.
Is Morphik an AI product?
Yes. Morphik Core is described as an AI-native toolset for visually rich documents and multimodal data, using techniques such as ColPali for multimodal search and LiteLLM for model access. The company behind it builds AI workers for back-office operations in skilled nursing and senior living.
How do I install Morphik Core?
The README recommends signing up at dev.morphik.ai rather than self-hosting. For self-hosting it links to dev.morphik.ai/docs/getting-started, and the repository ships install_and_start.sh, install_and_start.ps1, install_docker.sh and install_docker.ps1 alongside docker-compose.yml.
What does Morphik Core use ColPali for?
ColPali is the technique behind multimodal search, documented at dev.morphik.ai/docs/concepts/colpali and demonstrated in examples/colpali.py. It lets the engine search over the visual content of images, PDFs and videos through a single endpoint instead of relying only on extracted text.
Does Morphik Core require Ollama?
No. docker-compose.yml runs a config-check service that greps morphik.toml for provider = "ollama" and writes a flag file, and the Ollama service is marked required: false. Ollama starts only when the config selects it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/morphik-org-morphik-core)