# Adaptive Chunking: picking a splitting strategy per document instead of one global default

> Ekimetrics' LREC 2026 implementation scores four chunking methods against five intrinsic metrics and keeps the winner for each document. It ships as a Python package from source, not from PyPI.

**ekimetrics/adaptive-chunking** — Adaptive Chunking: automatically select the best chunking method per document for RAG. Accepted at LREC 2026.

- Repository: https://github.com/ekimetrics/adaptive-chunking
- Stars: 394 · Forks: 42
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ekimetrics-adaptive-chunking

## The problem: one splitter for every document in the corpus

Most RAG pipelines pick a chunk size once, wire it into an ingestion script, and never revisit it. The README states the premise plainly: no single chunking method works best for every document. A legal contract, a technical manual and a sustainability report have different paragraph structures, different table densities and different coreference patterns, and a 1100-token recursive split will treat all three the same way. Adaptive Chunking is aimed at teams that already accept chunking as a tunable stage and want the choice made per document rather than per pipeline. The intended audience is research-leaning: the package is the official implementation of a paper accepted at LREC 2026, the classifiers in pyproject.toml mark it as Alpha and Intended Audience :: Science/Research, and the repository ships a poster.pdf and a data/clair/ directory of 33 pre-parsed documents for reproducing the published numbers. If you want a drop-in replacement for a LangChain text splitter with no evaluation step, this is more machinery than the problem needs.

## How selection works: candidates in, five intrinsic scores out

The design is a comparison loop, not a learned model. Four candidate methods are registered by default: Recursive at a 1100-token target, Recursive at 600, Page splitting with post-processing to enforce size constraints, and LLM Regex, which asks a language model to generate document-specific split patterns. Each candidate produces a chunk list for the document, and each chunk list is scored by five intrinsic metrics that need no ground-truth answers: Size Compliance, Intrachunk Cohesion, Contextual Coherence, Block Integrity, and Filtered Missing Reference Error, which counts coreference chains (entity-pronoun pairs) broken across chunk boundaries. The method with the best score wins for that document. Both halves are meant to be extended: the README says you can register any callable that takes text and returns a list of chunks, and metrics.py holds the scoring functions. The recursive splitter lives in src/adaptive_chunking/splitters.py while the other three live under src/adaptive_chunking/paper/splitters.py, a split that tells you which parts are treated as the framework and which as the experiment harness. The reported intrinsic results put Adaptive Chunking at a mean of 91.07 across the five metrics, ahead of LLM regex at 89.80 and LangChain recursive at 88.62, with Semantic at 76.49 and Sentence at 73.26. Note what that ordering implies: the adaptive layer is a modest gain over a well-configured recursive split, not a step change, and the paper reports Wilcoxon p < 0.001 against all methods on those intrinsic metrics.

## Installing adaptive-chunking and chunking a folder of PDFs

There is no PyPI release. The badge line in the README is commented out, the repository has no releases, and pyproject.toml still carries version 0.1.0, so installation means cloning and installing in editable mode. Python 3.11 or newer is required.

```bash
git clone https://github.com/ekimetrics/adaptive-chunking.git
cd adaptive-chunking
pip install -e ".[dev]"
```

The extras are worth reading before you run anything, because they are not interchangeable. The core install gives you the splitter and the metrics. PDF and Excel parsing sits behind the parsing extra, coreference resolution behind coref, and the full paper reproduction behind paper.

```bash
pip install -e ".[parsing]"
```

Some metrics depend on spaCy models, which are downloaded separately from the package.

```bash
python -m spacy download en_core_web_sm
```

With the parsing extra in place, the README's quick start parses and chunks in one call. The function returns a list of dictionaries, and the documented keys are doc_name, chunk_index, chunk_text, chunk_pages, titles_context and chunk_len.

```python
from adaptive_chunking import chunk_files

chunks = chunk_files("path/to/pdfs/", chunk_size=600, chunk_overlap=50)

for chunk in chunks:
    print(chunk["doc_name"], chunk["chunk_index"], chunk["chunk_len"])
```

The default parser is Docling. Swapping it is a constructor argument rather than a config key.

```python
from adaptive_chunking import chunk_files
from adaptive_chunking.parsing import PyMuPDFParser

chunks = chunk_files("path/to/pdfs/", parser=PyMuPDFParser())
```

If you only want the splitter, RecursiveSplitter takes chunk_size, chunk_overlap, separators, merging and min_chunk_tokens, and exposes split_text. The merging parameter is documented with a value of "small_only" and min_chunk_tokens with 100, which is the configuration the README shows. Calling this directly skips the adaptive selection entirely: you get one strategy, applied uniformly.

## What the metrics do not tell you, and where the package stops

The five metrics are intrinsic, which is the point and the limitation. They score a chunking output on its own terms: are chunks inside the size bounds, do a chunk's sentences cohere with its embedding, does each chunk resemble its surrounding window, are paragraphs and tables left intact, are entity-pronoun pairs kept together. None of that measures whether a retriever surfaces the right passage for a real question. The RAG evaluation pipeline exists for that, with hybrid retrieval and answer correctness scoring, but it lives in the paper reproduction path and pulls in the paper extra: torch, langchain, haystack-ai, openai, deepeval, groq and more. Teams that install only the core package get the selection logic without the evidence that the selection helped on their own queries. The second boundary is cost. LLM Regex is one of the four default candidates, and by construction it calls a language model during chunking, so per-document selection over a large corpus is not free. The README gives no per-document latency or token budget for that path, and the repository's own scale claim is 33 documents across 3 domains at roughly 1.18M tokens, which is a benchmark corpus, not a production ingestion backlog. Third, the package targets documents, not arbitrary text: the parsing backends are PDF and Excel oriented, so a pipeline chunking HTML, chat logs or source code is outside what the parsing extra covers. Finally, development status is Alpha in pyproject.toml, and the last push to the repository was on 2026-07-06. Treat the API as something that can move.

## Against LangChain's recursive splitter and semantic chunking

The honest comparison is with the tools this package benchmarks against. LangChain's recursive splitter is a single deterministic function: you choose separators and a size, and every document gets the same treatment. It is fast, has no model dependencies, and the paper's own numbers show it scoring 88.62 on the intrinsic mean, close behind adaptive selection. If your corpus is homogeneous, the adaptive layer is buying you a few points of Size Compliance and Block Integrity for a lot of extra dependencies. Semantic chunking takes the opposite route from this project: it splits wherever consecutive sentence embeddings diverge past a threshold, so boundaries follow meaning rather than structure. The paper scores it at 76.49 intrinsic mean, with Size Compliance at 48.1 standing out as the weak spot, which is the expected failure of a method that ignores token budgets. Page splitting is the crude baseline at 59.1 retrieval completeness, and it survives in the candidate set because page boundaries are cheap and sometimes correct. LLM Regex is the closest in spirit to adaptive selection, since it also varies per document, but it spends a model call to produce split patterns rather than scoring candidates that already exist. The difference that matters: Adaptive Chunking does not generate a new splitting method, it chooses among existing ones and gives you a score for the choice.

## Licence, dependencies and the cost of keeping up

The package is MIT, and pyproject.toml declares license = "MIT" with the OSI classifier to match. One extra is not: the README labels the coref extra, which installs maverick-coref, as CC BY-NC-SA 4.0 and non-commercial. If you install ".[coref]" or ".[paper]", which depends on it, you are pulling a non-commercial component into your environment even though the surrounding code is MIT. That is a licensing question for your own counsel, not something the README resolves. On upgrade cost: there are no releases to track, so staying current means pulling main and reading commits. The dependency set is heavy and version-pinned in places, with torch==2.6.0 and torchvision==0.21.0 in the paper extra, and core dependencies that include sentence-transformers, spaCy, scikit-learn and tiktoken. Those are the packages most likely to break an editable install on a new Python minor version, and the classifiers only claim 3.11 through 3.13. The repository carries a CITATION.cff, a NOTICE file and an SBOM.md, which is more provenance paperwork than most research code ships with, and it signals the project expects to be cited rather than depended on.

## Conclusion

Adopt it if you already treat chunking as a tunable stage and want a scored, per-document choice rather than one global splitter; the default recursive splitter and the five metrics are usable without touching the paper pipeline. Do not adopt it if you need a published PyPI release, a stable API, or a system that runs without sentence-transformers and spaCy models on disk. First verify that pip install -e ".[parsing]" resolves Docling and PyMuPDF on your platform, that python -m spacy download en_core_web_sm completes, and that the coreference extra's CC BY-NC-SA 4.0 licence is acceptable for your use, since that extra is the one component that is not MIT.

## FAQ

### What is adaptive chunking in ekimetrics/adaptive-chunking?

It is a framework that evaluates several chunking strategies against five intrinsic quality metrics and automatically selects the best one for each document, so a RAG pipeline does not have to commit to a single splitter. The README describes it as the official implementation of a paper accepted at LREC 2026.

### How do I install adaptive-chunking?

Clone the repository and install it in editable mode, since there is no published PyPI release. The README shows pip install -e ".[dev]" for the full development setup, with parsing, coref and paper extras available separately, and some metrics additionally need python -m spacy download en_core_web_sm.

### Which chunking methods does adaptive-chunking compare?

Four by default: Recursive with a 1100-token target, Recursive with a 600-token target, Page splitting with post-processing for size constraints, and LLM Regex, which asks a language model to generate document-specific split patterns. The README states you can register any callable that takes text and returns a list of chunks.

### Is adaptive-chunking available on PyPI?

No. The PyPI badge in the README is commented out, the repository has no releases, and pyproject.toml still lists version 0.1.0, so installation is from a git clone with pip install -e.

## Sources

- [ekimetrics/adaptive-chunking on GitHub](https://github.com/ekimetrics/adaptive-chunking)
- [Issues](https://github.com/ekimetrics/adaptive-chunking/issues)
- [License: MIT](https://github.com/ekimetrics/adaptive-chunking/blob/main/LICENSE)
- [README](https://github.com/ekimetrics/adaptive-chunking/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ekimetrics-adaptive-chunking
