# Sparrow: document intelligence that stays on your GPUs

> Sparrow is KataNA ML's GPL-3.0 platform for enterprise document intelligence: REST APIs turn invoices, receipts, statements and tables into schema-validated JSON, powered by Vision LLM pipelines that run on your own infrastructure across MLX, vLLM, Ollama or Mistral OCR. An agent framework with Prefect monitoring orchestrates multi-step workflows on top.

**katanaml/sparrow** — Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM

- Repository: https://github.com/katanaml/sparrow
- Website: https://sparrow.katanaml.io
- Stars: 5,227 · Forks: 519
- Language: Python
- License: GPL-3.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/katanaml-sparrow

## Documents in, validated JSON out, on your own hardware

Sparrow is an API-first platform built for enterprise document intelligence, combining accurate structured extraction from documents, invoices, statements and tables, with workflow agents and decision agents. The positioning against the usual shape of this category is explicit: everything runs on your own infrastructure with no external API calls or cloud dependencies, so a pipeline of sensitive financial documents never leaves the machines you control. Input is universal across the genre's real workloads, invoices, receipts, forms, bank statements and tables, in images and multi-page PDFs, and output is clean structured data validated against a JSON schema. The platform layer adds the enterprise furniture, RESTful APIs for integration, rate limiting and usage analytics, with the codebase under GPL-3.0 and commercial licensing available for organizations that need it.

## One API surface, five inference backends

The backend story is the platform's most practical property: MLX on Apple Silicon, vLLM on NVIDIA, Ollama, Hugging Face and Mistral OCR as a cloud option, all presenting the same API surface, so switching hardware is a flag rather than a rewrite. Local Vision LLM support names a current model roster, Mistral, Qwen 3.6, DeepSeek OCR, dots.ocr and Gemma 4 among them, and the documentation is explicit that you download the model first with MLX, vLLM or Ollama before pointing Sparrow at it. The backend selection happens through repeated --options flags, mlx, ollama, vllm or mistral, alongside the pipeline choice. For a team whose available hardware changes between a MacBook prototype and a GPU server deployment, this is the difference between a demo and a product, and the cloud Mistral OCR backend remains available as the one deliberately external path.

## A 30-second setup with a poppler footnote

The quickstart is specific about its prerequisites: Python 3.12.10 or newer managed through pyenv, macOS for the MLX backend or Linux and Windows for the others, and a GPU with enough memory for the selected Vision LLM. The setup itself is conventional and complete:

```bash
pyenv install 3.12.10
pyenv global 3.12.10
```

```bash
python -m venv .env_sparrow_parse
source .env_sparrow_parse/bin/activate  # Linux/Mac
```

```bash
pip install -r requirements_sparrow_parse.txt
```

```bash
python api.py
```

Two platform details reward attention. macOS needs poppler installed through Homebrew for PDF processing. And the requirements file itself must match the platform, sparrow-parse with the mlx extra on macOS, plain sparrow-parse on Linux and Windows to skip the MLX libraries, a fork in the setup the README explains rather than hides. A separate environment_setup.md carries the complete instructions for the multi-pipeline layouts.

## A schema you write, a verdict you can check

The first documented extraction shows the whole contract in one command: you write a JSON schema for the fields you want, choose pipeline, backend and model, and point at a file:

```bash
./sparrow.sh '[{"instrument_name":"str", "valuation":0}]' \
  --pipeline "sparrow-parse" \
  --options mlx \
  --options mlx-community/Qwen2.5-VL-72B-Instruct-4bit \
  --file-path "data/bonds_table.png"
```

The answer comes back as data plus a verdict:

```json
{
  "data": [
    {"instrument_name": "UNITS BLACKROCK...", "valuation": 19049},
    {"instrument_name": "UNITS ISHARES...", "valuation": 83488}
  ],
  "valid": "true"
}
```

The schema is the integration contract on both ends, the model is told what shape to produce and the caller is told whether automatic validation passed, which is what lets extraction feed a data pipeline rather than a human inbox. The example itself, a bonds table with fund names and valuations, is representative of the financial statements the platform targets.

## Parse, Instructor and Agents as interchangeable pipelines

The architecture table reads as a small catalog. Sparrow ML LLM is the main API engine running the document processing pipelines. Sparrow Parse is the Vision LLM library for structured JSON extraction, the pipeline the quickstart exercises. Sparrow OCR handles text recognition as preprocessing, feeding cleaner input onward. Sparrow Agents provides workflow orchestration for complex multi-step processing, and Sparrow UI wraps the same API in a web interface with drag and drop upload, real-time results, JSON-based data query and structured output. The pluggability claim means these mix and match per task, Vision LLM extraction where layout matters, the text-focused Sparrow Instructor where it does not, and agent pipelines where the work is a sequence rather than a step. Instruction calling extends the platform beyond extraction entirely, text processing, validation and decision making through an instruction inference API served by models like Gemma, Mistral and Qwen 3.6. The web UI deserves its line in the architecture too, because it exercises the identical REST surface a backend integration would, making it both an operator tool and a live demonstration of the API contract.

## Agents you can watch, through Prefect

The agent framework is where Sparrow stops being an extraction service and becomes a workflow platform. It orchestrates multi-step workflows with custom agents, decision agents that judge rather than just transform, and what the documentation calls robust error handling, the difference between an agent pipeline that fails loudly on step four and one that silently produces half a document. The distinctive choice is observability: visual monitoring runs through Prefect, the Python workflow orchestration library, so agent executions are tracked as flows with their own dashboard rather than as log lines. A built-in dashboard and agent workflow tracking surface the same information for operators. For enterprise deployment, being able to show an auditor which agent transformed which document, and where each step succeeded or failed, is not a nicety, and building it on a known orchestrator rather than a bespoke UI is the kind of decision that ages well.

## A monorepo of three trees and a quickening cadence

The repository divides into three top-level trees, sparrow-data holding the parse and OCR libraries, sparrow-ml holding the LLM engine and the agents, and sparrow-ui holding the interface, with each pipeline recommended its own virtual environment in the setup guide, .env_sparrow_parse and siblings, an honest acknowledgment that the dependencies differ enough to be separated. Release history shows a platform mid-acceleration, v0.4.4 in September 2025, then v0.5.0 on 2026-05-26 and v0.6.0 on 2026-06-05 ten days later, with the last push on 2026-09-14 and a CHANGELOG tracking the sequence. Against managed document AI services, the comparison is structural rather than qualitative: those services extract through a vendor's API, while Sparrow's README stakes its ground on running with no external API calls at all, and which side of that line an enterprise needs is usually decided by its compliance department before its engineering team is consulted.

## Conclusion

Use Sparrow when document extraction must run inside your infrastructure, with no external API calls and the same API surface across Apple Silicon, NVIDIA and local inference backends. Use a managed document AI service when you prefer vendor-operated extraction and accept documents leaving your environment. Verify first that your GPU memory fits the chosen Vision LLM, that Python 3.12.10 and the platform-correct requirements file are in place, and note that GPL-3.0 governs the codebase with commercial licensing explicitly available for enterprise needs.

## FAQ

### What is Sparrow, the document AI platform?

Sparrow is KataNA ML's GPL-3.0 licensed platform for enterprise document intelligence, turning invoices, receipts, statements, forms and tables into validated JSON through REST APIs. It runs on your own infrastructure with ML, LLM and Vision LLM pipelines and no external API calls.

### Which hardware does Sparrow run on?

It supports MLX on Apple Silicon, vLLM on NVIDIA GPUs, Ollama, Hugging Face and the Mistral OCR cloud backend, all behind the same API surface. The quickstart requires Python 3.12.10 or newer and a GPU with enough memory for the chosen Vision LLM.

### How do you extract data with Sparrow?

Call sparrow.sh with a JSON schema describing the desired fields, select the pipeline and backend with --pipeline and --options flags, and pass the file path. The response returns the extracted data plus a valid flag reflecting automatic schema validation.

## Sources

- [katanaml/sparrow on GitHub](https://github.com/katanaml/sparrow)
- [License: GPL-3.0](https://github.com/katanaml/sparrow/blob/main/LICENSE)
- [Project website](https://sparrow.katanaml.io)
- [README](https://github.com/katanaml/sparrow/blob/main/README.md)
- [Releases](https://github.com/katanaml/sparrow/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/katanaml-sparrow
