Model or dataset
amaiya/onprem avatar
amaiya/onprem

OnPrem.LLM: Running LLM Pipelines on Non-Public Data Without a Cloud Round Trip

A toolkit for applying LLMs to sensitive, non-public data in offline or restricted environments

845 stars57 forksJupyter NotebookApache-2.0

At a glance

What is it?
OnPrem.LLM is a Python toolkit that wraps local and cloud LLM backends behind one interface for document ingestion, question answering, extraction and sandboxed agents. It is built for teams whose documents cannot leave the building, and its sparse-vectorstore path is the part worth understanding before you commit.
Who is it for?
Adopt OnPrem.LLM if you need document question answering, extraction or classification over files that cannot leave your network, and you want one Python interface across Ollama, llama.cpp, vLLM, Transformers and hosted providers. Skip it if you need a stable API surface: pyproject.toml still classifies the project as Development Status 3 - Alpha, and the install extras mean the dependency set you actually pull in depends on which optional features you enable.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Who OnPrem.LLM is actually for

The README describes OnPrem.LLM as a toolkit for applying LLMs to sensitive, non-public data in offline or restricted environments, and credits privateGPT as its main inspiration. That framing matters more than the feature list. This is not a general-purpose LLM framework. It is a document intelligence layer for people who already know they cannot send files to a hosted API, and who want the same code to keep working on the day a provider is approved for one narrow workload.

The audience is narrow but real: engineers in regulated industries, government-adjacent work, legal and medical document processing, and anyone whose data governance review ends with "the file never leaves the subnet." The README explicitly names AWS GovCloud LLMs among supported backends, which tells you the author is thinking about restricted environments rather than hobby projects.

The cloud-capable part is not a contradiction. The design keeps the local path as the default and treats hosted providers as a switch you flip per LLM object, so a pipeline can be written once and pointed at different backends per environment.

The mechanism: one LLM object over many backends, and two vectorstore strategies

The architecture is thin by design. A single LLM class takes a model string such as 'ollama/llama3.2' or 'anthropic/claude-sonnet-4-5-20250929', and the backend is selected from that string. The README lists llama_cpp, transformers, Ollama, vLLM, OpenAI and Anthropic among supported backends. The pyproject.toml dependency list shows litellm, langchain, langchain-openai, langchain-huggingface and langchain_litellm, so the provider abstraction is largely delegated rather than hand-rolled.

The document path is where the project makes its own choices. Ingestion goes through utils.download and llm.ingest, and retrieval defaults to a Chroma dense vectorstore installed via the chroma extra. The alternative is SparseStore, which the README describes as supporting environments with modest computational resources and enabling RAG without having to store embeddings in advance. That is the most consequential design decision in the project. A dense store requires an embedding pass over every chunk before the first query, which is where the GPU hours and the disk go. A sparse store shifts that cost to query time.

Beyond retrieval, there are named pipelines for information extraction, summarization, classification, question answering and agents, plus a YAML-configured workflow layer introduced in v0.19.0 and a visual workflow builder documented separately. Metadata-based query routing arrived in v0.21.0, letting a query be directed at a subset of the corpus rather than the whole index.

Installing OnPrem.LLM and running a first local query

The README says to install PyTorch first, then the package itself. The base install does not include the vectorstore, so the quick start example uses the chroma extra. Note that the README writes the extra as `pip install[chroma]` in one place; the working form follows the standard extras syntax shown in the quick start block.

bash
pip install onprem[chroma]

If you plan to use the agent features, there is a separate extra, and pyproject.toml defines it as the patchpal package:

bash
pip install onprem[agent]

The quick start then pulls a model through Ollama and constructs an LLM against it. This is the shortest path to a working local backend, provided Ollama is already running on the machine:

bash
ollama pull llama3.2
python
from onprem import LLM, utils

llm = LLM('ollama/llama3.2')
result = llm.prompt('Give me a short one sentence definition of an LLM.')

For retrieval, the README downloads a PDF into a directory and ingests the whole directory. The call takes a path, not a file, so the directory is the unit of ingestion:

python
utils.download('https://www.arxiv.org/pdf/2505.07672', '/tmp/my_documents/paper.pdf')
llm.ingest('/tmp/my_documents')
result = llm.ask('What is OnPrem.LLM?')

Structured output uses a Pydantic model. The README example extracts a value and a unit from a sentence, and prints 35 and mph respectively:

python
from pydantic import BaseModel, Field

class MeasuredQuantity(BaseModel):
    value: str = Field(description="numerical value")
    unit: str = Field(description="unit of measurement")

structured_output = llm.pydantic_prompt('He was going 35 mph.', pydantic_model=MeasuredQuantity)

One practical note: the README states that llama-cpp-python is optional if you use Ollama, or if you pass model_id to use Hugging Face Transformers instead, or if you are talking to an external REST API. That means the install burden depends entirely on which backend you pick, and the README points Windows users at a separate MSWindows.md file for llama-cpp-python, implying the CPU path there needs extra steps.

The sandboxed agent, and why the sandbox is the interesting part

The AgentExecutor, added in v0.22.0, runs an agent against a model and a task description. The README example uses model='openai/gpt-5-mini' with sandbox=True, and the task asks the agent to walk a directory, extract headings, count words per file, and write an index file.

python
from onprem.pipelines import AgentExecutor

executor = AgentExecutor(model='openai/gpt-5-mini', sandbox=True)
result = executor.run("""
Search this directory for all .md files and:
1. Extract all headings (# ## ###)
2. Count total words in each file
3. Create an index file 'documentation_index.md' with:
   - List of all markdown files
   - Word count for each
   - Main topics covered (from headings)
""")

The task is file mutation, not just reading, and that is exactly the scenario where an unsandboxed agent is dangerous. A model that decides to clean up the directory, or misreads a path, can destroy work. The sandbox flag is the acknowledgement that agent output is untrusted code execution. The README does not document what the sandbox actually enforces, which is a gap worth noting before you point an agent at a directory you care about. The agent extra pulls patchpal, which is presumably the sandboxing mechanism, but the README does not say so explicitly.

Where OnPrem.LLM is the wrong tool

The project is classified in pyproject.toml as Development Status 3 - Alpha. That is the author's own label, and it should shape how you treat the API. Pipelines that exist today may be renamed or restructured; the changelog shows a steady cadence of feature releases (v0.23.6 in August 2026, v0.23.5 and v0.23.4 before it), which is good for capability and bad for interface stability.

The dependency surface is the second concern. The base install pulls unstructured[all-docs], PyMuPDF, sentence_transformers, transformers, langchain and a long list of supporting libraries. That is a large footprint for a tool whose selling point is running in constrained environments. If your restricted environment has a strict package allowlist or an air-gapped mirror that is updated infrequently, satisfying this dependency tree is the hard part of the deployment, not the LLM itself.

Finally, this is not the right tool if you want a served API with an HTTP contract and a stable schema. The README describes a Python toolkit and a Streamlit-based web UI, not a general inference server. If you need to expose document QA to non-Python callers, you are building that layer yourself. And if your documents are already public and cost is the only concern, the local-first design buys you nothing over a hosted API.

How it compares to privateGPT and to a plain LangChain stack

The README states the project was inspired largely by privateGPT, and the two share a goal: local document question answering. The difference visible in the documentation is scope. privateGPT is oriented around the ingest-then-chat loop. OnPrem.LLM adds named pipelines for extraction, summarization and classification, a YAML workflow layer, metadata-based query routing, Pydantic structured outputs, and the sandboxed AgentExecutor. It also treats cloud providers as first-class backends rather than an afterthought, which privateGPT's local-first framing does not emphasize.

The other comparison is to assembling LangChain yourself. OnPrem.LLM is built on LangChain components, so the underlying primitives are the same. What you get by using the toolkit is the assembled pipeline, the document loaders tuned for formats like .msg and Excel via extract-msg and openpyxl, and the SparseStore option. What you give up is control over the exact chain structure and the ability to pin a minimal dependency set. If your team already has a LangChain codebase with a retrieval pattern it trusts, adopting OnPrem.LLM means either rewriting onto its pipeline classes or running two retrieval stacks side by side.

Licence, maintenance and the upgrade cost you should budget for

OnPrem.LLM is licensed under Apache-2.0, declared in pyproject.toml with license-files pointing at the LICENSE file. Apache-2.0 is permissive and includes an explicit patent grant, which is generally the friendlier option for corporate use than a copyleft licence. That said, the toolkit orchestrates third-party components, and their licences are separate from this one. llama-cpp-python, Chroma, Elasticsearch client libraries, boto3 and the rest each carry their own terms, and some model weights carry their own licences on top. Checking the licence of the model you intend to run is a separate exercise from checking the licence of this package.

The last push to the repository was on 2026-09-05, and the most recent tagged release is v0.23.6 from 2026-08-25. The release cadence through 2026 has been roughly monthly, and the news entries in the README show a steady stream of additions: workflows, asynchronous prompts, provider-implemented structured outputs, query routing, the AgentExecutor. Upgrading is not free. Each release has added surface area, and the optional extras in pyproject.toml mean an upgrade can pull in new transitive dependencies if you have the all extra installed. Pinning to a specific version and reading CHANGELOG.md before bumping is the practical approach. There is no documented migration guide, so the changelog is the only signal about what changed between versions.

Editorial conclusion

Adopt OnPrem.LLM if you need document question answering, extraction or classification over files that cannot leave your network, and you want one Python interface across Ollama, llama.cpp, vLLM, Transformers and hosted providers. Skip it if you need a stable API surface: pyproject.toml still classifies the project as Development Status 3 - Alpha, and the install extras mean the dependency set you actually pull in depends on which optional features you enable. Before committing, verify your target backend works in your environment, and check whether the SparseStore path removes the need to precompute embeddings for your corpus, because that decision drives both memory use and ingest time.

Frequently asked questions

What is OnPrem.LLM?

It is a Python-based toolkit for applying large language models to sensitive, non-public data in offline or restricted environments, according to the README. It is designed for fully local execution but also supports cloud providers such as OpenAI and Anthropic.

How do I install OnPrem.LLM?

The README says to install PyTorch first, then run pip install onprem. RAG with the default Chroma vectorstore needs the chroma extra, and AI agents need the agent extra.

Which LLM backends does OnPrem.LLM support?

The README lists llama_cpp, transformers, Ollama, vLLM, OpenAI and Anthropic among supported backends. The backend is chosen from the model string passed when constructing an LLM, for example 'ollama/llama3.2'.

Does OnPrem.LLM require a GPU or large amounts of memory?

The README states that modules like SparseStore support environments with modest computational resources, enabling RAG without having to store embeddings in advance. The default Chroma dense vectorstore does require embeddings to be computed and stored.

What is the SparseStore in OnPrem.LLM?

It is described in the README as an option for environments with modest computational resources that enables RAG without having to store embeddings in advance. The documentation points to an advanced example on NSF awards for a worked case.

Is OnPrem.LLM production-ready?

pyproject.toml classifies the project as Development Status 3 - Alpha, which is the author's own label. The release cadence through 2026 has been roughly monthly, and the README does not document a migration guide between versions.

Official sources

  1. amaiya/onprem on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/amaiya-onprem.svg)](https://hysenlabs.com/projects/amaiya-onprem)