Model or dataset
llmware-ai/llmware avatar
llmware-ai/llmware

llmware: local RAG pipelines and a 300+ model catalog in Python

Unified framework for building enterprise RAG pipelines with small, specialized models

14,843 stars2,936 forksPythonApache-2.0

At a glance

What is it?
llmware packages document parsing, text chunking, embeddings and small local models behind one Python interface, aimed at teams that want retrieval augmented generation to run on a laptop, an AI PC or a self-hosted server. The trade-off is a wide, fast-moving surface area and configuration that lives in your code rather than in a served endpoint.
Who is it for?
llmware fits teams that need document parsing, text chunking, embeddings and small local models behind one Python API, and who are willing to pin a version and read the source when a method is undocumented. It is the wrong choice if you want a served inference endpoint you can call over HTTP from any language, or if you need a stable API contract across minor releases.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 136 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem llmware targets: retrieval that never leaves the machine

Most RAG stacks assume you will call a hosted model. You chunk documents with one library, embed them with another, store vectors in a third, and send the assembled prompt to an API. Four dependencies, four failure modes, and a data path that leaves your network. llmware's README frames the project as a unified framework for "knowledge-based local, private, secure LLM-based applications", optimized for AI PC and local laptop, edge and self-hosted deployment across Windows, Mac and Linux.

The intended user is an engineer building an internal document assistant: policies, filings, contracts, support tickets. The README lists pdf, pptx, docx, xlsx, txt, csv, md, json/jsonl, wav, png, jpg and html as ingestable formats, which is a wider net than most retrieval libraries cast. Audio and image files appearing in that list is a signal about the audience: teams whose knowledge is not all text.

The second half of the pitch is the model catalog. The README describes 300+ models in quantized, optimized formats, including 50+ finetuned SLIM, Bling, Dragon and Industry-Bert models, with support for GGUF, OpenVINO, ONNXRuntime, ONNXRuntime-QNN, WindowsLocalFoundry and Pytorch. Cloud models from OpenAI, Anthropic and Google are also supported, so the local-first stance is a default rather than a restriction.

That combination is the actual claim: one framework where the retrieval half and the inference half are designed together, so the context you retrieve is in the shape the model expects. Plenty of projects do one half well and leave the join to you. Whether the join here is worth the fixed pipeline is the question the rest of this review tries to answer.

How the pipeline is assembled: Library, embedding, Query, Prompt

The architecture is four objects in sequence, and the README's own examples make the data flow explicit.

A Library is the container. `Library().create_new_library("my_library")` creates both a text collection in a database and a file resource directory, which the README describes as `llmware_data/accounts/{library_name}`. Ingestion is a single call: `lib.add_files("/folder/path/to/my/files")`. Files are routed by extension to the correct parser, then parsed, text chunked and indexed. That routing-by-extension design is why the format list matters so much: an unsupported extension has nowhere to go.

Embeddings are installed onto a library rather than configured globally. `lib.install_new_embedding(embedding_model_name="mini-lm-sbert", vector_db="milvus", batch_size=500)` attaches one embedding model and one vector store. The README shows a second call adding `industry-bert-sec` with `chromadb` to the same library, so a library can carry multiple embeddings over different vector databases. That is unusual and useful for comparing retrieval quality on the same corpus, but it also means the embedding choice is per library, not per query, unless you name it explicitly.

Query is where that naming shows up. `Query(lib)` uses the library's default; `Query(lib, embedding_model_name="mini_lm_sbert", vector_db="milvus")` targets a specific pair. The README lists text, semantic, hybrid, metadata and custom filter queries, with `text_query`, `semantic_query` and `text_query_with_document_filter` as examples. Prompt then takes retrieved chunks and packages them as model-ready context, including in batches, per the README's description of adding a file to a prompter.

The model layer sits behind `ModelCatalog`. `ModelCatalog().list_all_models()` enumerates the catalog, `ModelCatalog().load_model("llmware/bling-phi-3-gguf")` returns a model object, and that object exposes both `inference` and `stream`. The point of the indirection is that the calling code does not change when the underlying runtime does: a GGUF build, an OpenVINO build and a cloud API model are reached through the same two methods. That is the strongest part of the design, because it means swapping runtimes is a string change rather than a rewrite.

Installing llmware and running a first library

The repository carries a `setup.py` that reads the version from `llmware/__init__.py` and publishes the package as `llmware` on PyPI, so installation is a normal pip install. The README's Python badge lists 3.10 through 3.14, while `setup.py` classifiers list 3.9 through 3.13; the two do not agree, and that discrepancy is worth knowing before you pick an interpreter.

Install from PyPI:

bash
pip install llmware

Create a library and point it at a folder of mixed documents. The README gives this exact sequence:

python
from llmware.library import Library

lib = Library().create_new_library("my_library")
lib.add_files("/folder/path/to/my/files")
lib.install_new_embedding(embedding_model_name="mini-lm-sbert", vector_db="milvus", batch_size=500)

After `add_files` returns, the library card should reflect parsed documents and text chunks; `Library().get_library_card("my_library")` returns metadata including documents, text chunks, images, tables and the embedding record, according to the README.

Query the result:

python
from llmware.retrieval import Query
from llmware.library import Library

lib = Library().load_library("my_library")
q = Query(lib)
results = q.semantic_query("what is the retention policy", result_count=10)

To generate an answer rather than retrieve chunks, load a model through the catalog and hand the context to a Prompt. The README uses `llmware/bling-tiny-llama-v0` for this:

python
from llmware.prompts import Prompt

prompter = Prompt().load_model("llmware/bling-tiny-llama-v0")
response = prompter.prompt_main("what is the retention policy?", context="Insert Sources of information")

The repository root also contains `welcome_to_llmware.sh` and `welcome_to_llmware_windows.sh`, which the layout suggests are setup scripts for the two platforms, though the README does not document what they install.

Where llmware gets awkward: parsers, embeddings and version drift

The first real constraint is the parser matrix. `add_files` routes by file extension, so coverage is only as good as the parsers shipped in your installed version. A scanned PDF with no text layer, a spreadsheet with merged cells, or a format not on the README's list will produce something, but the README does not describe fallback behaviour or how to detect a parse that silently yielded nothing. For a compliance corpus, that is the failure mode to test for first: ingest a file you know is hard, then inspect the library card and the chunk text before trusting the index.

The second is embedding migration. Adding a second embedding is a documented one-liner, but the README is silent on what happens to an existing embedding when you change the chunking parameters or re-ingest a document. There is no documented rollback path, and no documented re-embedding workflow. If your corpus changes often, plan for the library to be rebuilt rather than patched, and budget the embedding compute accordingly.

The third is version drift. The repository is not archived, and the last push was on 2026-05-17, with v0.4.6 released on 2026-04-14. The `setup.py` classifier still reads "Development Status :: 4 - Beta", and the classifier list and README badge disagree about supported Python versions. Treat the API as moving: pin the version, and read `llmware/` in the checkout when a method's behaviour is not in the docs.

Finally, llmware is a Python library, not a service. There is no documented REST server, so anything non-Python has to go through your own wrapper. The README's homepage points to a documentation site, and the repository ships a `docs/` directory, but the README itself does not describe a hosted or containerized deployment mode. If your platform team expects a container image with a health endpoint, that is work you are adding, not work llmware does.

llmware against LangChain: one opinionated pipeline or many integrations

The comparison people search for is llmware versus LangChain, and the difference is architectural rather than feature-by-feature.

LangChain is a composition layer. Its value is the breadth of integrations: any vector store, any model provider, any loader, wired together by you. The cost is that you own the glue, the version compatibility between integrations, and the debugging when a chain misbehaves three abstractions deep.

llmware inverts that. Library, Query and Prompt are fixed objects with fixed responsibilities, and the model catalog normalizes access across GGUF, OpenVINO, ONNXRuntime and API models behind `ModelCatalog().load_model(...)`. You get less choice about how the pieces connect and more certainty about what a call does. The README's framing of "the smallest possible compute footprint" is the design constraint driving that: a fixed pipeline is easier to run on a laptop than an assembled one.

The honest split is this. If your retrieval logic is unusual, or you need to swap vector stores per tenant, or your team already has LangChain code in production, llmware's fixed pipeline will feel like a cage. If you want document parsing, chunking, embedding and inference to work together on day one without choosing five libraries, llmware removes that work. Note also that llmware supports cloud models from OpenAI, Anthropic and Google, so the local-first stance is a default, not a wall. A team can prototype locally and point the same Prompt object at a hosted model later without rewriting the retrieval half.

Licence and the cost of keeping up

llmware is Apache-2.0, and the repository carries both a `LICENSE` and a `NOTICE` file at the root. Apache-2.0 permits commercial use and modification and includes a patent grant, with the usual obligations around retaining notices and stating changes. The `NOTICE` file is worth reading before redistribution, since it can carry attribution terms that travel with the code. This is a description of the licence text, not legal advice; get your own review before shipping a derived product.

The models are a separate question. The README points to Hugging Face for the model families, and each model on that hub carries its own licence, which may differ from Apache-2.0. A quantized GGUF build of a base model inherits the base model's terms, not llmware's. Check per model, not per framework, especially before shipping a product that embeds weights.

Upgrade cost is dominated by the surface area. The catalog spans 300+ models and multiple inference runtimes, so a version bump can move model names, parser behaviour or embedding defaults. The `setup.py` version is read from `llmware/__init__.py`, so the installed version is unambiguous once you pin it. Pin it in your requirements file, and check the release notes for v0.4.6 and the two releases before it before moving. Because the project is a library rather than a service, an upgrade is a dependency change in your application, which means your test suite is the only thing standing between a version bump and a behaviour change in production.

Editorial conclusion

llmware fits teams that need document parsing, text chunking, embeddings and small local models behind one Python API, and who are willing to pin a version and read the source when a method is undocumented. It is the wrong choice if you want a served inference endpoint you can call over HTTP from any language, or if you need a stable API contract across minor releases. Before adopting it, install the pinned version from PyPI, run the Library and Query example from the README against a folder of your own PDFs, and confirm that the parsers you need for those file types are present in the version you installed.

Frequently asked questions

What does LLM stand for?

The README does not expand the acronym; it uses LLM throughout when describing model-based applications and the model catalog. No definition appears in the repository files.

What are the top 5 LLM models?

No ranking is published. The README states that llmware's catalog holds 300+ models, including 50+ finetuned SLIM, Bling, Dragon and Industry-Bert models, and that cloud models from OpenAI, Anthropic and Google are also supported.

What are LLM products?

The README does not answer this in general terms. It describes llmware itself as having two components: a model catalog and a RAG pipeline for connecting knowledge sources to generative AI models.

What does LLM stand for in ChatGPT?

The README does not cover ChatGPT or expand the acronym. It only notes that llmware supports leading cloud models from OpenAI, Anthropic and Google alongside its local model catalog.

Official sources

  1. License: Apache-2.0
  2. llmware-ai/llmware on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/llmware-ai-llmware.svg)](https://hysenlabs.com/projects/llmware-ai-llmware)