# QMedia: Self-Hosted Multimodal AI Search for Content Creators

> QMedia is an open-source AI content search engine that extracts, indexes, and answers questions over text, images, and short videos using a fully local RAG pipeline. It targets content creators who want to mine multimedia material without sending data to external APIs.

**QmiAI/Qmedia** — An open-source AI content search engine designed specifically for content creators. Supports extraction of text, images, and short videos. Allows full local deployment (web app, RAG server, LLM server). Supports multi-modal RAG content Q&A. 

- Repository: https://github.com/QmiAI/Qmedia
- Stars: 635 · Forks: 76
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/qmiai-qmedia

## What QMedia Solves for Content Creators

Most AI search tools treat content as plain text. QMedia is built specifically for the mixed media a content creator actually handles: screenshots, short-form videos, long articles, and image-heavy posts. It indexes all three formats in a unified retrieval layer, so a single query can surface a relevant frame from a video, a paragraph from a blog post, and a captioned image in one result set.

The target user is a creator who accumulates material from research, social media, and their own archives and needs to answer questions like "show me everything I have about product pricing" across formats. The system is designed for full local deployment, meaning no material leaves the host machine.

## Three-Service Architecture and Data Flow

QMedia splits its work across three independent services that communicate over HTTP.

`mm_server` (port 50110) is the model layer. It hosts the Ollama LLM, the CLIP image embedding model, the BGE text embedding model, the Qanything OCR engine, and the Faster Whisper video transcription engine. It exposes a unified API so the rest of the system never calls model libraries directly.

`mmrag_server` (port 50111) is the retrieval and Q&A layer, built on LlamaIndex and Python. When content is ingested, this service calls `mm_server` to extract transcripts, captions, and embeddings, then stores vectors and metadata. At query time it retrieves matching chunks and calls the LLM to generate an answer, returning both the answer and the source content cards.

`qmedia_web` (port 50112) is the front end, built with TypeScript, Next.js, TailwindCSS, and Shadcn/UI. The README notes it was inspired by the XHS web layout: content surfaces as cards rather than a list of links. The web service depends on `mmrag_server`, which in turn depends on `mm_server`, so the startup order matters.

All three services can also be deployed independently and embedded in other systems for content extraction alone.

## Installing with Docker Compose

The repository provides a `docker-compose.yml` at the root that builds and wires all three services. Clone the repository and start everything with:

```bash
git clone https://github.com/QmiAI/Qmedia.git
cd Qmedia
docker-compose up --build
```

The compose file maps the `./assets` directory into `mmrag_server` as `/assets`, which is where ingested media is stored. Once running, the web interface is available on port 50112, the RAG API on 50111, and the model API on 50110. The `SERVICE_ENDPOINT` environment variable on the web container points it at the RAG server:

```yaml
environment:
  - SERVICE_ENDPOINT=http://qmedia_mmarg_server:50111
```

For users who want to swap the LLM, `mm_server` supports switching between Ollama models. The README lists `llama3:8b-instruct-q4_0` as the lightweight option and `llama3:70b-instruct` for higher-quality output. Changing the model means reconfiguring the Ollama instance inside `mm_server`; the RAG layer is designed to be model-agnostic.

Separate installation instructions for each service live in `mm_server/README.md`, `mmrag_server/README.md`, and `qmedia_web/README.md`.

## Multimodal Extraction: What Actually Happens to Each Format

Text and image content goes through two paths. Text is chunked and embedded with the BGE multilingual encoder, which the README describes as aligned to GPT Encoder for cross-lingual retrieval. Images are passed to CLIP for visual embedding and to Qanything for OCR, so both the visual style and any text inside the image become searchable. The README notes that image style and text layout are part of the breakdown stored per content card.

Short video processing depends on Faster Whisper. The model runs on CPU, which is important because GPU availability is not assumed. Whisper produces a transcript; the LLM then generates a summary. At the time of the last push, several video features were still marked with empty checkboxes in the README: highlight detection, style classification, and detailed content breakdown. The transcription and summarization paths are listed as complete; the more granular video analysis is listed as future work.

Google content search is listed as a supported source alongside local data, meaning the RAG pipeline can retrieve from both private archives and live search results.

## Real Limitations to Weigh Before Adopting

Running three separately built services locally is a meaningful operational commitment. Each service has its own Dockerfile and its own Python or Node dependencies. The README directs users to three separate sub-READMEs for installation details, and the compose file does not document resource requirements.

CPU-based Faster Whisper transcription will be slow on large video libraries. The README does not document a batch ingestion pipeline or a queue; it is not clear whether the system handles concurrent ingestion requests gracefully.

The project has no GitHub releases. Versioning is tracked only in `CHANGELOG.md`. Users who need a stable, pinned release for reproducible deployments will need to pin a specific commit themselves.

The visual understanding model `llava-llama3` is listed in the README with an unchecked checkbox, meaning it was not complete at the last push date of 2026-04-09. Any workflow that depends on visual scene understanding rather than OCR alone should verify the current state of that component before building on it.

## Comparison with Perplexica

Perplexica is an open-source AI search engine that also runs locally and builds answers from retrieved web content. The key difference in approach is scope: Perplexica focuses on web search with a text-first pipeline, while QMedia is designed from the start for mixed-media private archives. Perplexica does not include video transcription or image embedding components. QMedia, by contrast, does not emphasize live web search as its primary ingestion path, even though the README lists Google search as one supported source. The two tools solve adjacent problems for different primary use cases.

## License and Maintenance Status

QMedia is released under the MIT license, which permits use, modification, and distribution without restriction. The last push to the repository was on 2026-04-09, roughly five and a half months before today. That is within the six-month window, so the repository is not technically stale, but given that the project has no releases and several features remain unfinished, adopters should review the commit log and open issues before committing to it as a production dependency.

The repository has no documented upgrade path between states. Because the asset storage is mounted as a host volume in the compose configuration, data survives container rebuilds, but schema changes in the vector store or embedding models would require manual re-ingestion.

## Conclusion

QMedia suits content creators who need to search and query a private library of images, videos, and articles without relying on cloud APIs. Anyone unwilling to run three Python/Node services locally, or whose video library exceeds what a CPU-based Whisper instance can process in reasonable time, should weigh that operational cost first. Before committing, verify that the LLaVA visual model listed in the README is marked complete, since the checkbox for it was unchecked at the time of the last push on 2026-04-09.

## FAQ

### What is QMedia?

QMedia is an open-source multimodal AI content search engine designed for content creators. It indexes text, images, and short videos in a local RAG pipeline and answers natural-language queries with source attribution.

### How does QMedia compare to Plex for media management?

QMedia and Plex serve different purposes. Plex organizes and streams media files; QMedia extracts and indexes the content of those files so you can ask questions about them. QMedia is a search and Q&A tool, not a media server.

### Can QMedia run fully offline without sending data to external APIs?

Yes. The README describes full local deployment covering the web app, RAG server, and LLM server. All model inference runs inside the three local services using Ollama, CLIP, BGE, and Faster Whisper.

### Which video features in QMedia are complete?

According to the README, video transcription via Faster Whisper and LLM-based short video summarization are listed as implemented. Highlight detection, style classification, and detailed content breakdown are listed as future work with unchecked checkboxes.

## Sources

- [Issues](https://github.com/QmiAI/Qmedia/issues)
- [License: MIT](https://github.com/QmiAI/Qmedia/blob/main/LICENSE)
- [QmiAI/Qmedia on GitHub](https://github.com/QmiAI/Qmedia)
- [README](https://github.com/QmiAI/Qmedia/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/qmiai-qmedia
