semchunk: hierarchical text chunking for Python RAG pipelines
A fast, lightweight and easy-to-use Python library for splitting text into semantically meaningful chunks.
At a glance
- What is it?
- semchunk splits text into token-bounded chunks while keeping local semantic context, and works with any tokenizer. Here is how it installs, what the hierarchical algorithm does, and where it falls short.
- Who is it for?
- Adopt semchunk if you already have a tokenizer or token counter in your pipeline and need chunks that respect token limits without cutting sentences arbitrarily; the pip install is one line and the chunkerify() signature accepts Tiktoken, Transformers or a plain callable. Do not adopt it if you need a parser that understands document structure such as Markdown headings or PDF layout, because semchunk operates on plain text and the README documents no such handling.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 109 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem semchunk solves for RAG builders
Most text splitting code counts characters or splits on blank lines. That works until a chunk boundary lands mid-sentence and the retrieval step returns a fragment that reads like a ransom note. semchunk is a Python library for splitting text into smaller chunks while preserving as much local semantic context as possible, and it does that under a hard token budget rather than a character budget.
That distinction matters because token counts are what embedding models and language models actually consume. A chunker that respects characters will produce chunks of wildly different token lengths depending on the language and the vocabulary. semchunk takes a tokenizer or a token counter as an argument and treats its output as the unit of measurement.
The audience is narrow and identifiable: engineers building retrieval augmented generation pipelines, document ingestion jobs, or any preprocessing step that feeds text into a model with a context limit. If you are chunking a handful of paragraphs once, a splitter from the standard library is enough. If you are chunking a corpus and the chunk boundaries affect retrieval quality, the token-aware approach is the point.
How the hierarchical chunking algorithm works
The README describes semchunk as powered by a novel hierarchical chunking algorithm and claims 15% better RAG performance than its closest competitors, pointing to a benchmarks section. The README does not publish the benchmark methodology in the excerpt available, so treat that number as the project's own claim rather than an independently reproduced result.
What the API reveals about the mechanism is more useful. `chunkerify()` returns a callable. You give it a string and it returns a list of chunks. You give it a list of strings and it returns a list of lists. The tokenizer you pass in is the measuring instrument, and `chunk_size` is the ceiling each chunk must stay under.
The AI-powered path is a second layer. Passing a `chunking_model` such as `kanon-2-enricher` routes the split through an Isaacus enrichment model, which requires the Isaacus SDK and an `ISAACUS_API_KEY` environment variable. That is a network call per chunking operation, which changes the failure profile: the deterministic local path cannot fail on a timeout, and the AI path can.
Two more arguments shape the output. `overlap` accepts a ratio below 1 or an absolute token count at or above 1, and `offsets` returns the character positions of each chunk alongside the text. Offsets are what you need if you plan to highlight retrieved passages in an original document, and the README's quickstart shows both returned together.
Installing semchunk with pip and running a first chunk
The package installs from PyPI. The README gives `pip install semchunk` and notes that `uv` works as well. There is also a conda-forge channel, and a Rust port named `semchunk-rs` maintained by a separate contributor.
pip install semchunkThe quickstart constructs a chunker by passing a model name, an encoding name, a tokenizer object, or a function. The README's example uses a deliberately tiny chunk size of 4 tokens so the output fits on one screen. Note the comment in that example: semchunk does not know how many special tokens your tokenizer adds to every input, so you may want to deduct that number from your chunk size.
import semchunk
chunk_size = 4
text = 'The quick brown fox jumps over the lazy dog.'
chunker = semchunk.chunkerify('gpt-4', chunk_size)
assert chunker(text) == ['The quick brown fox', 'jumps over the', 'lazy dog.']Passing a list instead of a string changes the return shape, and `progress=True` turns on a progress bar. For large batches the README shows `processes=2` to enable multiprocessing, which is the setting to reach for when a single-threaded loop over thousands of documents becomes the bottleneck.
chunks = chunker(text, offsets=True, overlap=0.5)The line above returns the chunks plus their offsets, with consecutive chunks overlapping by half a chunk. That is the configuration most retrieval pipelines want, because a sentence split across two chunks is still fully present in at least one of them.
The special-token gap and other places semchunk bites
The most concrete limitation is documented by the project itself. semchunk does not know how many special tokens your tokenizer prepends or appends, so a `chunk_size` of 512 can produce inputs your model sees as 514 tokens. The README's advice is to deduct the special token count from your chunk size. That is a manual step, and getting it wrong produces truncation errors that surface far from the chunking code.
The second limitation is the AI-powered path. It requires the Isaacus SDK, an API key, and network access. The README documents the setup but does not document rollback, retry behaviour, or what happens when the enrichment service is unavailable. If your ingestion job runs in an environment without outbound network access, the local tokenizer path is the only usable one.
The third is scope. semchunk takes text. It does not parse Markdown structure, PDF layout, or HTML, and the README documents no such handling. If your source documents have meaningful structure that should survive chunking, you need a converter upstream. Docling is named in the README as a project that uses semchunk, which suggests the intended division of labour: another tool converts, semchunk splits.
Finally, `chunk_size` defaults to the tokenizer's `model_max_length` minus the tokens produced by an empty string, and raises a `ValueError` when that cannot be determined. Passing an explicit size avoids the guesswork.
semchunk compared with recursive character splitting
The obvious alternative is a recursive character splitter, the kind used in LangChain and similar frameworks. The approach differs in what it measures. A recursive splitter walks a list of separators, usually paragraphs then sentences then words, and stops when the piece fits under a character limit. It is deterministic, fast, and has no tokenizer dependency at all.
The difference shows up when your corpus is not English prose. A character limit that yields 400 tokens of English can yield far more tokens of a language with a denser tokenization, and the model truncates. semchunk inverts the relationship: you set the token budget, and the algorithm finds boundaries that fit. That is the trade you are making, and it costs you a tokenizer call per chunk.
A second alternative is to skip chunking libraries and split on document structure yourself, using headings or sections as boundaries. That preserves semantics better than any statistical method when the structure exists, but it gives you no control over chunk size, and a single long section still needs a fallback. semchunk can serve as that fallback, or as the primary splitter when structure is absent.
Maintenance status, licence and upgrade cost
The repository is not archived, and the last push was on 2026-06-13. The most recent release listed is v4.1.0 from 2026-06-12, following v4.0.0 in March 2026 and v3.2.5 in October 2025. The version in `pyproject.toml` is 4.1.1, which is ahead of the latest release tag in the release list, so check PyPI before pinning.
The major version bump from 3 to 4 within roughly eight months is the upgrade cost to plan for. The README documents the current `chunkerify()` signature with keyword-only arguments after `chunk_size`, including `chunking_model`, `isaacus_client`, `tokenizer_kwargs`, `max_token_chars`, `memoize` and `cache_maxsize`. If you pinned an earlier major version, that signature is where breakage will appear. `memoize` defaults to True, which means the chunker caches results; `cache_maxsize` controls the bound. Long-running processes that chunk unbounded distinct texts should set `cache_maxsize` deliberately rather than relying on the default.
Licensing is MIT, stated in `pyproject.toml` and in the LICENCE file at the repository root. That permits commercial use and modification with attribution. The AI-powered path is separate: the Isaacus SDK and API key are governed by Isaacus's own terms, not by semchunk's MIT licence, so the two need to be evaluated independently. Python 3.10 or newer is required; classifiers list support through 3.14.
Editorial conclusion
Adopt semchunk if you already have a tokenizer or token counter in your pipeline and need chunks that respect token limits without cutting sentences arbitrarily; the pip install is one line and the chunkerify() signature accepts Tiktoken, Transformers or a plain callable. Do not adopt it if you need a parser that understands document structure such as Markdown headings or PDF layout, because semchunk operates on plain text and the README documents no such handling. Before committing, verify two things yourself: how your tokenizer's special tokens shift the effective chunk_size, and whether the AI-powered path is worth the extra install of isaacus plus an ISAACUS_API_KEY. The README states semchunk does not know how many special tokens your tokenizer adds, so deduct them from chunk_size and re-check your chunk boundaries on your own corpus.
Frequently asked questions
What is chunking with an example?
Chunking is splitting a text into smaller pieces. The README's example chunks 'The quick brown fox jumps over the lazy dog.' with a chunk size of 4 tokens and returns ['The quick brown fox', 'jumps over the', 'lazy dog.'].
What is a good example of chunking?
The README's quickstart is the example to follow: construct a chunker with semchunk.chunkerify() by passing a model name, an encoding name, a tokenizer, or a token counting function, then call it on a string or a list of strings. Passing offsets=True and overlap=0.5 returns chunks with their offsets and half-chunk overlap.
What is chunking used for?
In semchunk's case, chunking prepares text for retrieval augmented generation and other model inputs that have a token limit. The chunker splits a document into pieces that each stay under chunk_size while keeping local semantic context, and can return offsets so you can map a chunk back to its position in the original text.
What is a good chunk size?
The README does not prescribe a size, but it warns that semchunk does not know how many special tokens your tokenizer adds, so you may want to deduct that number from your chunk size. If chunk_size is omitted it defaults to the tokenizer's model_max_length minus the tokens from an empty string, or raises a ValueError.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/isaacus-dev-semchunk)