Model or dataset
aryn-ai/sycamore avatar
aryn-ai/sycamore

Sycamore (sycamore-ai): an LLM document processing engine for ETL and RAG

🍁 Sycamore is an LLM-powered search and analytics platform for unstructured data.

608 stars72 forksPythonApache-2.0

At a glance

What is it?
Sycamore is an Apache-2.0 Python framework that turns PDFs, images and reports into structured, searchable chunks and loads them into vector or hybrid search engines. The catch is that its best partitioning path depends on Aryn's hosted DocParse service.
Who is it for?
Adopt Sycamore if your pipeline is Python, your corpus is heavy on PDFs and scanned images, and you are willing to run a Ray-backed DocSet job and either use Aryn DocParse or run the Aryn Partitioner locally. Skip it if you need a pure-Python, dependency-light parser, if your documents are already clean text, or if you cannot accept a hosted partitioning call in your data path.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 61 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Sycamore solves, and who ends up using it

Most RAG pipelines break at the same place: the document. A PDF with a two-column layout, an embedded table and a chart does not turn into clean text with a naive extractor, and the chunks that come out carry the damage into retrieval. Sycamore is aimed squarely at that step. The README describes it as an "open source, AI-powered document processing engine for ETL, RAG, LLM-based applications, and analytics on unstructured data", and the repository topics list dataprep, etl, information-retrieval and semantic-search alongside ai and llm. The audience is therefore not end users but engineers building ingestion pipelines: people who need to partition, enrich, chunk and load a document corpus, and who are comfortable writing Python transforms rather than configuring a no-code connector.

The second audience is teams that already have a vector store and are unhappy with what is in it. Sycamore's pitch is that better partitioning produces better chunks, and the README claims the Aryn DETR model, trained on "80k+ enterprise documents", can lead to "6x more accurate data chunking and 2x improved recall on hybrid search or RAG when compared to alternate systems". Those numbers are the vendor's own, stated without a published methodology in the README, so treat them as a hypothesis to test on your corpus rather than a settled result.

The DocSet abstraction and the Ray execution model

The central abstraction is the DocSet, described in the README as "a scalable and robust abstraction for document processing" with "powerful high-level transformations in Python for data processing, enrichment, and cleaning". A DocSet is a distributed collection of document records; you apply functions to it and the framework handles execution. The README calls this a "functional programming approach" that lets you "rapidly customize and experiment with your chunking".

The execution backend is Ray, listed as a feature ("Scalable Ray backend") and visible in the repository structure: lib/ holds the library, apps/ holds services including an integration app, and the top-level Makefile only recurses into apps. That split matters operationally. You are not calling a parser function and getting a string back; you are submitting a job that Ray distributes, which is what makes large corpora tractable and also what makes small ones feel heavy.

The data flow runs in one direction. Documents enter through crawlers or local paths, are partitioned into labelled elements, then pass through transforms such as table extraction, OCR, visual summarization and LLM-powered UDFs. Embeddings are generated with a model of your choice, and the result is written to a target store. The README names the supported destinations explicitly: OpenSearch, ElasticSearch, Pinecone, DuckDB, Qdrant and Weaviate. The repository confirms this with separate example files for DuckDB, Pinecone, Qdrant, Weaviate and Neo4j, plus S3 and HTML ingestion examples. Partitioning itself is delegated: Sycamore "leverages Aryn DocParse", a serverless GPU API that returns partitioned output as JSON, with the option to "run the Aryn Partitioner locally" instead.

Installing sycamore-ai and running a first DuckDB job

The README states that Sycamore currently runs on Linux and Mac OS. Installation is a single pip command:

bash
pip install sycamore-ai

Database connectors are not bundled. They ship as Python extras, and the README's example adds DuckDB:

bash
pip install sycamore-ai[duckdb]

The supported connector extras named in the README are `duckdb`, `elasticsearch`, `opensearch`, `pinecone`, `qdrant` and `weaviate`. The pyproject.toml in the repository goes further for development, listing extras including `eval`, `neo4j`, `local-inference`, `legacy-partitioners`, `anthropic`, `google-genai` and `iceberg`. Note that `neo4j` and `local-inference` appear in the development dependency group rather than in the README's connector list, so confirm the extra you need against the published package before pinning it.

To use the hosted partitioner you need an Aryn DocParse key, obtained by signing up at the link the README gives. If you would rather not send documents to a hosted service, the README says you can run the Aryn Partitioner locally. The repository also ships a compose.yaml that starts a full local stack. Its own comments describe the sequence: a crawler downloads one sample PDF, the demo UI starts, OpenSearch logs heavily, and then the Sycamore container downloads PyTorch files before running Ray. The UI is served on port 3000, mapped through the `UI_PORT` variable, and OpenSearch is exposed through `OPENSEARCH_PORT`. The file documents two recovery commands, `docker compose run reset` to reset the stack and `docker compose run sycamore_crawler_http_sort_all` to fetch more benchmark files. Expect the first run to be slow for reasons unrelated to Sycamore's own code.

Where Sycamore is the wrong tool

The dependency surface is the first real constraint. Python support is pinned in pyproject.toml to `>=3.11,<3.14`, so anyone on 3.10 or on a newer interpreter is out. The Ray backend and the PyTorch downloads the compose file describes mean the install is not lightweight, and the README's own platform statement stops at Linux and Mac OS. Windows users are not addressed.

The second constraint is the hosted dependency. The headline partitioning quality comes from Aryn DocParse, a serverless API. Running the Aryn Partitioner locally is offered as an alternative, but the README does not document parity between the two paths, so you cannot assume the local model reproduces what the hosted service returns. If your documents cannot leave your network and the local path underperforms, you are choosing between a weaker partitioner and a different project.

The third is scope. Sycamore is an ingestion and preparation engine. It loads into vector and hybrid search stores, and the README mentions "an OpenSearch hybrid search and RAG engine for testing", but the query-serving story is thin here. If you want a complete retrieval application with an API and a UI out of the box, you are assembling it yourself.

Finally, there is no documented rollback. The README does not cover what happens when a job writes partial results to a target index, and the compose file's reset command resets the local Docker stack, not a production store. Plan idempotency into your own pipeline.

How Sycamore differs from Unstructured and LlamaIndex

Unstructured takes a document-centric approach: you call a partitioning function on a file and get elements back, synchronously, in your own process. Sycamore's DocSet is the opposite bet. It treats the corpus as the unit of work and distributes it across Ray, which is why the README frames DocSets as removing "the undifferentiated heavy lifting of reliably loading chunks". For a few hundred PDFs, Unstructured's model is simpler to reason about. For a corpus that keeps growing, the distributed model is the one that survives.

LlamaIndex overlaps at the edges but starts from the other end. Its readers and node parsers are components inside an orchestration framework aimed at querying; Sycamore is a data preparation engine that happens to ship a test search engine. The practical difference shows up when you hit a bad table: in Sycamore the fix is a transform applied across the DocSet, and in a query-first framework the fix tends to be a different parser configuration.

The honest comparison is against building it yourself. If your documents are clean HTML or Markdown, the repository's own examples/html_ingest.py and examples/markdown.py suggest Sycamore can handle them, but the vision model and DocParse machinery are doing nothing for you. The value is concentrated in the messy formats.

Maintenance, licensing and the cost of upgrading

Sycamore is licensed Apache-2.0, which permits commercial use and modification. The licence covers the code in this repository; it does not cover the Aryn DocParse service, which is a separate hosted product with its own sign-up and terms. Running the Aryn Partitioner locally is the path that keeps you inside the open source boundary, and the licence implications of that choice are the ones worth confirming with your own counsel rather than inferring from the repository.

The release cadence is uneven. v0.1.32 landed on 2025-04-21, v0.1.33 on 2025-07-22, and v0.1.34 on 2026-05-31, roughly a ten-month gap before the most recent tag. The last push to the default branch was on 2026-08-01, so commits have continued since that release. The version numbers themselves are the real upgrade signal: everything is still 0.1.x, which means no stability guarantee and a reasonable expectation of breaking changes between minor releases. The repository pins `sycamore-ai = "^0.1.32"` in pyproject.toml, so caret semantics apply within the 0.1 line.

Upgrade cost is dominated by two moving parts you do not control: the Aryn partitioning service and the Ray version underneath. The compose file pins images through environment variables such as `UI_VERSION`, `OPENSEARCH_VERSION`, `RPS_VERSION` and `JUPYTER_VERSION`, which means the whole stack is versioned together and a bump is a coordinated change rather than a pip upgrade. Budget for re-running a sample of your corpus after each bump and comparing chunk output, because the README offers no compatibility table.

Editorial conclusion

Adopt Sycamore if your pipeline is Python, your corpus is heavy on PDFs and scanned images, and you are willing to run a Ray-backed DocSet job and either use Aryn DocParse or run the Aryn Partitioner locally. Skip it if you need a pure-Python, dependency-light parser, if your documents are already clean text, or if you cannot accept a hosted partitioning call in your data path. Before committing, verify three things: which extras your target database needs, how the local partitioner performs on your own PDFs, and what the Aryn DocParse key and quota terms look like for your volume.

Frequently asked questions

How do I install sycamore-ai?

Run pip install sycamore-ai on Linux or Mac OS. Database connectors are installed as extras, for example pip install sycamore-ai[duckdb]; the README lists duckdb, elasticsearch, opensearch, pinecone, qdrant and weaviate.

Does sycamore-ai require an Aryn API key?

To use Aryn DocParse you sign up through the link in the README and use the resulting API key. The README also states you can choose to run the Aryn Partitioner locally instead, though it does not document parity between the two paths.

What Python versions does sycamore-ai support?

The pyproject.toml pins Python to >=3.11,<3.14. The README states the project currently runs on Linux and Mac OS.

What vector databases can sycamore-ai load into?

The README names OpenSearch, ElasticSearch, Pinecone, DuckDB, Qdrant and Weaviate. The repository also contains example files for Neo4j and BigQuery ingestion.

Is sycamore-ai actively maintained?

The repository is not archived and the last push to the default branch was on 2026-08-01. The most recent release is v0.1.34 from 2026-05-31, following a roughly ten-month gap after v0.1.33.

Official sources

  1. aryn-ai/sycamore on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/aryn-ai-sycamore.svg)](https://hysenlabs.com/projects/aryn-ai-sycamore)