Model or dataset
llama-farm/llamafarm avatar
llama-farm/llamafarm

LlamaFarm: a local AI platform for RAG, classifiers and OCR

Deploy any AI model, agent, database, RAG, and pipeline locally or remotely in minutes

837 stars59 forksPythonApache-2.0

At a glance

What is it?
LlamaFarm bundles a FastAPI server, a Celery RAG worker and a Universal Runtime for HuggingFace inference behind one llamafarm.yaml. It is Apache-2.0 and ships desktop builds for macOS, Windows and Linux.
Who is it for?
Adopt LlamaFarm if you want a self-hosted RAG and document-processing stack where every setting lives in one llamafarm.yaml and models can be swapped between Ollama, vLLM and an OpenAI-compatible endpoint. Do not adopt it if you need a stable API surface: the project is at v0.0.34, and the README does not document rollback or a version compatibility matrix.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 112 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap LlamaFarm fills between a model and a working application

Running a model locally is the easy part. The work that follows is document ingestion, embedding, retrieval, reranking, OCR, classification and a way to reach all of it from an application. LlamaFarm packages that layer. The README describes it as an open-source AI platform that runs entirely on your hardware, and lists RAG, custom classifiers trained with SetFit from 8 to 16 examples, anomaly detection with 12 or more algorithms, tool calling over the Model Context Protocol, OCR and named entity recognition as the capabilities it covers.

The audience is narrower than the tagline suggests. The repository ships a desktop app for macOS, Windows and both x86_64 and ARM64 Linux, which is aimed at people who want to click through a UI. It also ships a Go and Python codebase with Nx task orchestration, which is aimed at engineers who intend to run the services themselves. Those two audiences get different setup paths and different failure modes. The desktop download is the friendlier entry point; the CLI plus Nx path is the one that exposes the architecture.

Three services, one config file, and where the ports land

The architecture table in the README lists three components. The Server is a FastAPI REST API that also serves the Designer web UI and handles project management, on port 14345. The RAG Worker is a Celery worker for asynchronous document processing and has no fixed port in that table. The Universal Runtime handles ML model inference, embeddings, OCR and anomaly detection on port 11540.

Configuration is centralised. The README states that all configuration lives in llamafarm.yaml, with no scattered settings or hidden defaults. That file carries a runtime block where each named model declares a provider, a model identifier and a base_url. Because the provider is a field rather than a build-time choice, the same project can point at a local runtime, at Ollama, or at any OpenAI-compatible endpoint by editing YAML.

The RAG path is a four-step flow. A dataset is created and linked to a processing strategy and a database. Files in PDF, DOCX, Markdown or TXT are uploaded, and processing runs automatically unless the --no-process flag is passed. Manual processing exists for cases where auto-processing was deliberately skipped, which the README gives as large batches. Queries then run as semantic search with optional metadata filtering. The split between upload and process is the design decision worth noticing: it lets a large ingestion job be staged without blocking on the worker.

Installing the lf CLI and running a first RAG query

The README gives three install routes. The desktop app needs no additional setup. The CLI route installs a binary called lf. The source route uses Nx. Start with the CLI on macOS or Linux:

bash
curl -fsSL https://raw.githubusercontent.com/llama-farm/llamafarm/main/install.sh | bash

On Windows the equivalent is a PowerShell one-liner, and the README also points at the releases page for a direct download. After installation, create a project and start the services:

bash
lf init my-project
lf start

The first command generates llamafarm.yaml. The second starts the services and opens the Designer UI, which the README places at http://localhost:14345. If that page loads, the Server service is up; it does not by itself prove the Universal Runtime is reachable on 11540.

Chat is the fastest way to confirm a model is wired correctly. The README shows both an interactive form and a one-off message:

bash
lf chat
lf chat "Hello, LlamaFarm!"

For a RAG workflow, create a dataset bound to a processing strategy and a database, then upload documents. The README notes that uploads auto-process by default:

bash
lf datasets create -s default -b main_db research
lf datasets upload research ./papers/*.pdf
lf rag query --database main_db "Your query"
lf rag health

What you should see is the upload triggering the Celery worker, and the query returning semantic matches from main_db. The README's own example is truncated at the point where it would explain the --no-process case, so check docs/ before assuming the flag's exact behaviour.

Where LlamaFarm is the wrong tool

Version numbers are the first limitation. The most recent release listed is v0.0.34, dated 2026-05-22, and the last push to main was on 2026-06-10. A 0.0.x line means the maintainers have not promised API stability. If you are building against the FastAPI server on 14345 rather than the lf CLI, expect the surface to move between releases. The README does not document a rollback procedure or a version compatibility matrix between CLI and server, so a mismatched pair is a real risk.

The second limitation is operational. Three services, one of them a Celery worker, means you are running a queue and a broker. That is a reasonable architecture for batch document processing and a poor fit for a single-user script that just wants to embed a folder of text. If your problem is one model and one prompt, the desktop app already covers it and the server plus worker plus runtime stack is overhead you will pay for on every upgrade.

The third is scope creep in the runtime. The Universal Runtime advertises text generation, embeddings, OCR through Surya, EasyOCR, PaddleOCR or Tesseract, document extraction, classification, NER, reranking and anomaly detection. Each of those pulls its own model weights and dependencies. A team that only needs embeddings is still installing the surface area of the OCR and vision stack. The README does not describe a minimal install profile that trims these, so plan for the full dependency set.

LlamaFarm compared with wiring Ollama and LangChain yourself

The obvious alternative is assembling the same pipeline from Ollama plus a framework such as LangChain or LlamaIndex. The difference is where the integration work sits. With the hand-built route, you own the ingestion pipeline, the vector store schema, the worker queue and the HTTP layer between them. LlamaFarm supplies those as named services and exposes them through lf subcommands and a REST API, which trades flexibility for a fixed shape.

A second alternative is a hosted platform. The README's own framing of LlamaFarm is enterprise AI capabilities on your own hardware, no cloud required, and it lists no API costs and offline operation once models are downloaded as properties of that choice. A hosted service will not match that, but it also will not ask you to run a Celery worker. The decision is whether data residency and per-token cost matter more than operational burden.

A third comparison is at the model layer rather than the platform layer. LlamaFarm is provider-agnostic by configuration. The README shows the same runtime block pointed at the Universal Runtime on http://127.0.0.1:11540/v1, at Ollama on http://localhost:11434/v1, and at an OpenAI-compatible endpoint such as https://api.openai.com/v1 with an api_key read from an environment variable. That means adopting LlamaFarm does not lock you to local inference; it locks you to the llamafarm.yaml shape. If you later move to vLLM, the README lists a VLLM_BASE_URL in .env.example pointing at port 8000, and the provider block is what changes.

Licence, upgrade cost and what the repository does not say

LlamaFarm is Apache-2.0, and the repository carries a NOTICES file alongside the LICENSE, which is the pattern you expect when third-party components are vendored or redistributed. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, but it also requires that you preserve notices and state changes. The practical implication for a fork is that you inherit the obligation to keep attribution intact, including whatever the NOTICES file enumerates. That is a description of the licence text, not advice about your situation; have counsel read it if you plan to redistribute.

Upgrade cost is dominated by the release cadence. Three releases are listed between 2026-04-28 and 2026-05-22, roughly every two to three weeks. The repository carries a release-please-config.json and a .release-please-manifest.json, so versioning and changelog generation are automated, and CHANGELOG.md is the place to look before bumping. The last push was on 2026-06-10, so the tree has been quiet since then relative to that earlier cadence.

What the README does not document is as important as what it does. There is no migration guide between llamafarm.yaml versions, no statement about which Python and Go versions the released binaries were built against beyond the Python 3.10+ and Go 1.24+ badges, and no rollback instructions. The .env.example file is thorough about provider endpoints and API keys, and lists OPENAI_API_KEY, ANTHROPIC_API_KEY, TOGETHER_API_KEY, GROQ_API_KEY and COHERE_API_KEY with their base URLs, plus OLLAMA_HOST, OLLAMA_PORT, VLLM_HOST, VLLM_PORT, TGI_HOST and TGI_PORT. That file is the most reliable single reference for which providers the platform expects to talk to.

Editorial conclusion

Adopt LlamaFarm if you want a self-hosted RAG and document-processing stack where every setting lives in one llamafarm.yaml and models can be swapped between Ollama, vLLM and an OpenAI-compatible endpoint. Do not adopt it if you need a stable API surface: the project is at v0.0.34, and the README does not document rollback or a version compatibility matrix. Before committing, verify that the Universal Runtime starts on port 11540 on your hardware, and check the CLI reference in docs/ for the exact dataset flags, since the README truncates its own RAG example.

Frequently asked questions

What is LlamaFarm used for?

The README describes it as an open-source AI platform that runs on your own hardware for RAG applications, custom text classifiers, anomaly detection and document processing. It bundles a FastAPI server, a Celery RAG worker and a Universal Runtime for model inference, embeddings and OCR.

How do I install LlamaFarm?

The README gives three routes: download the desktop app for macOS, Windows or Linux; run the install script for the lf CLI on macOS and Linux or the PowerShell equivalent on Windows; or clone the repository and use Nx to start the server, rag and universal-runtime services. The CLI route is a single curl command piped to bash.

Does LlamaFarm require an internet connection or API keys?

The README states that it works offline once models are downloaded and that there are no per-token API costs with open-source models. API keys are only needed if you configure a cloud provider; .env.example lists optional keys for OpenAI, Anthropic, Together, Groq and Cohere.

Which model providers can LlamaFarm use?

The runtime block in llamafarm.yaml takes a provider field. The README shows three: universal for the built-in runtime on port 11540, ollama for GGUF models on port 11434, and openai for any OpenAI-compatible endpoint such as vLLM or Together. .env.example also lists vLLM on port 8000 and TGI on port 8080.

What ports does LlamaFarm use?

The architecture table gives port 14345 for the FastAPI server and Designer web UI, and port 11540 for the Universal Runtime. The RAG worker has no port listed because it is a Celery worker. Ollama, vLLM and TGI defaults appear in .env.example on 11434, 8000 and 8080 respectively.

Official sources

  1. License: Apache-2.0
  2. llama-farm/llamafarm on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/llama-farm-llamafarm.svg)](https://hysenlabs.com/projects/llama-farm-llamafarm)