# Chonkie: a Python chunking library for RAG ingestion pipelines

> Chonkie splits text for retrieval-augmented generation with seven chunker types, refineries and a self-hosted REST API. It is small, MIT licensed, and opinionated about installing only what you use.

**feyninc/chonkie** — 🦛 CHONK docs with Chonkie ✨ — The lightweight ingestion library for fast, efficient and robust RAG pipelines

- Repository: https://github.com/feyninc/chonkie
- Website: https://docs.chonkie.ai
- Stars: 4,779 · Forks: 362
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/feyninc-chonkie

## The problem Chonkie targets: chunking code you keep rewriting

Every retrieval pipeline needs the same unglamorous step. Take a document, cut it into pieces small enough to embed and retrieve, keep enough context that the pieces still mean something, and repeat that consistently across formats. Most teams write this once, then rewrite it when a new file type or a new embedding model arrives. Chonkie is aimed squarely at that repetition.

The README frames it as a chunking library that "just works", and the pyproject description calls it "the no-nonsense chunking library". The intended audience is Python developers building RAG ingestion who do not want to depend on a large framework for one stage of the pipeline. The package is MIT licensed and requires Python 3.10 or newer, so it drops into existing services rather than dictating an application structure.

## Seven chunkers, refineries, and a pipeline object

The mechanism is a chunker object you call on text. The README's basic example imports RecursiveChunker, instantiates it with no arguments, calls it on a string, and iterates the result. Each returned chunk exposes at least text and token_count, so downstream code can filter or log by size without re-tokenizing.

The chunker table lists seven options with short aliases. TokenChunker splits into fixed-size token chunks. FastChunker does SIMD-accelerated byte-based chunking and ships in the default install. SentenceChunker splits on sentences. RecursiveChunker splits hierarchically using customizable rules. SemanticChunker splits on semantic similarity and is credited in the README as inspired by Greg Kamradt's work. LateChunker embeds text first and then splits it, which the table says produces better chunk embeddings. CodeChunker handles source code.

The Pipeline object is the more interesting piece of architecture. It chains stages: chunk_with, then refine_with, then run. The README example stacks a recursive chunk at tokenizer gpt2 with chunk_size 2048 and a markdown recipe, a semantic chunk at 512, an overlap refinery with context_size 128, and an embeddings refinery using sentence-transformers/all-MiniLM-L6-v2. The same pipeline exposes arun for asynchronous execution. That design means the split-then-adjust sequence lives in configuration rather than in your own glue code, and the same pipeline can be replayed through the REST API as a stored object.

## Installing Chonkie and running your first chunk

The README gives two basic install paths. The plain install pulls the default dependency set, which the chunker table says already includes FastChunker. The uv variant is offered as the faster option.

```bash
pip install chonkie
```

```bash
uv pip install chonkie
```

After installing, the README's first example is three steps: import a chunker, instantiate it, call it. Run this and you should see the text and token count printed per chunk.

```python
from chonkie import RecursiveChunker

chunker = RecursiveChunker()
chunks = chunker("Chonkie is the goodest boi! My favorite chunking hippo hehe.")

for chunk in chunks:
    print(f"Chunk: {chunk.text}")
    print(f"Tokens: {chunk.token_count}")
```

If you want the pipeline form instead, the README chains chunk_with and refine_with calls and then calls pipe.run(texts=...). The result is a doc object whose chunks attribute you iterate. The README does not document what happens when a refinery's embedding model is unavailable, so treat that path as something to test yourself before it reaches production. The README also points to docs.chonkie.ai for per-chunker install instructions, under a stated rule of minimum installs: install only the extras your chosen chunker needs, and avoid the all extra in production.

## Self-hosting the REST API with Docker

Chonkie can run as a self-hosted API rather than an imported library. The README shows installing the extra set, then starting the server through the CLI, and notes uvicorn as a direct alternative. Interactive documentation is served at /docs.

```bash
pip install "chonkie[api,semantic,code,catsu]"
chonkie serve --port 3000 --reload --log-level debug
```

The docker-compose.yml in the repository maps port 8000, mounts ./data into the container, and sets DATABASE_URL to sqlite+aiosqlite:////app/data/chonkie.db, which means pipelines persist to a SQLite file on that mounted volume. It also lists optional provider keys for the embeddings refinery: OPENAI_API_KEY, COHERE_API_KEY, VOYAGE_API_KEY and MISTRAL_API_KEY, all defaulting to empty. CORS_ORIGINS defaults to *, which allows every origin; that default is worth changing before exposing the service. The healthcheck polls http://localhost:8000/health every 30 seconds after a 15 second start period.

Pipelines are the API's reusable unit. The README creates one with a POST to /v1/pipelines carrying a name and a steps array of chunk and refine entries, then lists them with a GET on the same path. Because those rows live in SQLite, the ./data volume is the thing to back up.

## Where Chonkie is the wrong tool

The API server is single-node by design. Pipelines are stored in a local SQLite database, so horizontal scaling means either sharing a file or accepting per-instance state. There is no mention of an external database backend in the compose file or the README.

The semantic and late chunkers depend on embeddings. That means a model or a provider API key, extra latency, and a cost that scales with document volume. If your documents are short, uniform, and already well delimited, a token or sentence chunker will do the job and the semantic path only adds moving parts. Similarly, the README explicitly warns against the all extra in production environments, so a team that installs everything to avoid deciding has taken on dependencies it may never call.

The library is Python only. The pyproject classifiers list Python 3.10 through 3.13 and no other language. A Node or Go service cannot embed it directly; it would have to call the REST API, which reintroduces the operational questions above.

## How Chonkie differs from general text splitters

The obvious alternative is LangChain's text splitter family, which lives inside a much larger orchestration framework. The difference is scope. LangChain gives you splitters alongside chains, agents, memory and retrievers, and you adopt the framework to get the splitter. Chonkie is the splitter, plus refineries and a pipeline object, and nothing above that. If your application already uses LangChain for orchestration, adding Chonkie means a second dependency for a stage you already have covered. If you are assembling your own retrieval stack and only need the ingestion step, Chonkie's narrower surface is the point.

A second comparison is writing the splitter yourself. The README's own framing, about making your gazillionth chunker, is an admission that hand-rolled splitting is common. What Chonkie adds over a local function is the refinery layer: overlap and embeddings refinement as composable stages, plus the same configuration expressed as an API-stored pipeline. Whether that abstraction earns its place depends on how many chunking variants you actually maintain.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-02, so the project is being worked on. The most recent release listed is v1.7.0 from 2026-07-07, following v1.6.8 on 2026-06-01. That cadence suggests releases arrive faster than the version numbers move, with patch releases between minor bumps.

Upgrade cost concentrates in two places. The default dependency set includes chonkie-core, tokie, numpy, httpx, tenacity and tqdm, with chonkie-core pinned at a minimum of 0.10.2; a major bump there is the thing most likely to break chunker behaviour. The extras are the second axis, since each chunker pulls its own dependencies and the README advises installing only those you use.

The licence is MIT, which permits commercial use and modification. That is a statement about the licence text, not legal advice; if you redistribute the library or bundle it into a product, read the LICENSE file in the repository yourself.

## Conclusion

Adopt Chonkie if you are building a Python RAG ingestion path and want chunking, overlap and embedding refinement behind one interface instead of hand-rolled splitting code. Skip it if your pipeline is not Python, or if you cannot accept that semantic and late chunking pull in embedding models and provider API keys. Verify first that the extras you need are the ones the docs name for your chunker, and confirm the API container's SQLite path and health endpoint fit your deployment before you build an image around them.

## FAQ

### What is Chonkie?

Chonkie is a Python chunking library for RAG ingestion pipelines, MIT licensed and requiring Python 3.10 or newer. It provides several chunker types, refinement steps such as overlap and embeddings, and an optional self-hosted REST API.

### How do I install Chonkie?

The README gives pip install chonkie, or uv pip install chonkie for the faster installer. It recommends installing only the extras your chosen chunker needs and warns against the all extra in production environments.

### Is Chonkie an alternative to LangChain text splitters?

It overlaps with them but covers less ground. LangChain bundles splitters inside a larger orchestration framework, while Chonkie provides chunkers, refineries and a Pipeline object for the ingestion stage only, and nothing above it.

### Can I run Chonkie as a server instead of a library?

Yes. The README documents installing chonkie[api,semantic,code,catsu], starting it with chonkie serve, and running it via Docker with docker compose up. The compose file maps port 8000 and stores pipelines in a SQLite file under the mounted ./data volume.

## Sources

- [feyninc/chonkie on GitHub](https://github.com/feyninc/chonkie)
- [License: MIT](https://github.com/feyninc/chonkie/blob/main/LICENSE)
- [Project website](https://docs.chonkie.ai)
- [README](https://github.com/feyninc/chonkie/blob/main/README.md)
- [Releases](https://github.com/feyninc/chonkie/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/feyninc-chonkie
