Minions: A Protocol for Reducing Cloud LLM Costs by Reading Long Contexts Locally
Big & Small LLMs working together
At a glance
- What is it?
- Minions is a Python framework from Stanford's Hazy Research lab that routes long-context reading tasks to a small on-device model and reserves the frontier cloud model for synthesis and reasoning. The last push was on 2026-03-12, more than six months before this article was written.
- Who is it for?
- Minions is worth evaluating for workloads that require processing long documents where the bulk of the work is extraction or summarization rather than complex reasoning. The cost savings depend on the length of the local context and the quality gap between the on-device model and the frontier model; the Minions paper at arXiv:2502.15964 provides the experimental analysis.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Cost Problem Minions Addresses
Frontier LLM APIs charge per token, and long-context tasks are expensive because the document content counts toward the input token budget. A legal review, a technical audit of a large codebase, or summarizing a lengthy research report can require tens of thousands of tokens of context, most of which is the document being analyzed rather than the instruction or the question.
The observation behind Minions is that most of the work in long-context tasks is reading and extracting information, which a small on-device model can do adequately. The synthesis step, which requires reasoning across the extracted information to produce an answer, benefits most from a frontier model's capabilities. By splitting the task across a local model and a cloud model, you can keep the expensive API calls short while still getting high-quality final answers.
The README states that the protocol enables small on-device models to collaborate with frontier models in the cloud, and that reading long contexts locally can reduce cloud costs with minimal or no quality degradation. The paper at arXiv:2502.15964 provides the empirical analysis supporting that claim. Minions is a demonstration of the protocol, not a production service.
Two Protocols: Minion and Minions
The repository implements two distinct protocols. The Minion protocol (singular) is a simple delegation: the local model reads the context and answers a specific question, then passes the result to the cloud model for final synthesis. This is a single round of local-then-cloud communication.
The Minions protocol (plural) is a multi-round version where the cloud model decomposes the task into subtasks, delegates each subtask to local model workers, and iterates based on the partial results. The local models work in parallel on different parts of the context, and the cloud model synthesizes across all the worker outputs. The tokasaurus server is recommended for the Minions protocol because its high-throughput serving handles the parallel worker load more efficiently than a single-request server.
The README also describes Secure Minions, a separate protocol that adds end-to-end encryption to the local-remote communication. The secure chat documentation is in the secure/ subdirectory and requires additional dependencies via pip install -e ".[secure]". The Secure Minions Chat Blogpost from 2025-05-12 on the HazyResearch site describes the security model.
Installing Minions and Starting a Local Model Server
Clone the repository and install the package in editable mode:
git clone https://github.com/HazyResearch/minions.git
cd minions
pip install -e . # installs the minions package in editable modeFor Apple Silicon users who want to use the MLX-LM backend:
pip install -e ".[mlx]"The README states it has been tested on Mac and Ubuntu with Python 3.10 to 3.11, and that Python 3.13 is not supported. A virtual environment is recommended to avoid dependency conflicts; the README provides a conda example:
conda create -n minions python=3.11A local model server must also be installed and running before Minions can use a local model. Ollama is the recommended option for hardware without a dedicated GPU. Lemonade is recommended for AMD CPUs, GPUs, and NPUs. Tokasaurus is recommended for NVIDIA GPUs, particularly when using the multi-worker Minions protocol. Install tokasaurus with:
pip install tokasaurusAt least one cloud provider API key must also be set as an environment variable. The README lists OpenAI, TogetherAI, DeepSeek, and MiniMax among the supported cloud providers.
Local Model Server Options and Their Trade-offs
Minions separates the local model server from the framework itself, which means the quality and speed of the local model processing depends on the server you choose and the model you run through it.
Ollama is the broadest compatibility option. It runs on any hardware including CPU-only machines and supports the Flash Attention optimization on supported hardware. The README notes that enabling Flash Attention on macOS requires setting an environment variable before starting the Ollama app. Ollama's throughput is lower than dedicated GPU servers, which matters more for the Minions plural protocol than the Minion singular protocol.
Lemonade is the option for AMD hardware, supporting APU configurations documented in the RyzenAI documentation. The README notes that Lemonade does not support the Minion-CUA protocol at this time.
The Cartesia-MLX backend is available for Apple Silicon specifically, requiring XCode, the command line tools, nanobind, and the Cartesia Metal backend before the MLX package itself. This is the most involved setup of the local server options.
For llama-cpp-python users, the README provides a client initialization example showing how to load a GGUF model file and configure GPU layer offloading, as well as how to load models directly from Hugging Face by repository ID and file pattern.
Limitations and Cases Where Minions Is the Wrong Tool
Minions is designed around a specific cost-reduction strategy: read locally, synthesize in the cloud. It is not a general-purpose agent framework. If your task requires the frontier model to reason over the full document text rather than over extracted summaries, the quality gap introduced by local reading may be significant.
The quality of the local model determines how much information is preserved in the extraction step. A small model that misses nuanced passages or makes extraction errors will degrade the quality of the final answer regardless of how good the cloud synthesis model is. The paper at arXiv:2502.15964 provides analysis of the quality trade-off at specific model sizes, but the results are specific to the experimental setup.
As an alternative framework for document-heavy workloads, LlamaIndex is a widely used Python library for building RAG pipelines that retrieves relevant chunks before sending them to a language model. The key difference in approach is that LlamaIndex uses retrieval to limit context rather than local model reading, and it has a much larger ecosystem of integrations and production deployments. Minions is research software demonstrating a specific protocol; LlamaIndex is a production-oriented framework with active maintenance and community support.
Maintenance State and Dependencies
The last push to the repository was on 2026-03-12, which is more than six months before this article was written. The repository has no GitHub releases and no changelog. The setup.py version is pinned at 0.1.0. These indicate a research demonstration repository that has not been developed into a maintained production library.
The dependency list in setup.py and requirements.txt is extensive, including several GPU attestation packages (nv-attestation-sdk, nv-local-gpu-verifier, azure-security-attestation, azure-identity) that are part of the Secure Minions extension and are not required for basic use. Installing the full requirements.txt on a machine without NVIDIA GPU drivers will cause errors. The README's installation steps use pip install -e . rather than installing from requirements.txt, which is the safer starting point.
The license is MIT. The paper citation is available in the README for academic use. Teams evaluating Minions for production use should verify whether the repository has been updated before committing to it, given the current maintenance gap.
Editorial conclusion
Minions is worth evaluating for workloads that require processing long documents where the bulk of the work is extraction or summarization rather than complex reasoning. The cost savings depend on the length of the local context and the quality gap between the on-device model and the frontier model; the Minions paper at arXiv:2502.15964 provides the experimental analysis. The last push was on 2026-03-12, which is more than six months before this article was written. Teams considering Minions for production use should verify that the repository is being maintained before building workflows around it. The secure chat extension requires significant additional setup and is documented separately in the secure/ subdirectory.
Frequently asked questions
What does Minions actually do to reduce cloud LLM costs?
Minions routes the context-reading portion of a task to a small local model rather than sending the full document to the cloud API. Only the extracted information and the final synthesis request go to the cloud model, which reduces the number of tokens billed by the cloud provider.
What local model servers does Minions support?
Minions supports Ollama for CPU or general-purpose GPU setups, Lemonade for AMD CPUs and GPUs, and tokasaurus for NVIDIA GPUs. MLX-LM and Cartesia-MLX are also supported for Apple Silicon, and llama-cpp-python is supported for GGUF models.
Is Minions production-ready?
The README describes the repository as a demonstration of the protocol rather than a production service. The last push was on 2026-03-12, more than six months before this article was written, and the version is 0.1.0 with no GitHub releases. It is research software.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hazyresearch-minions)