chonkiejs: TypeScript text chunking for RAG pipelines
🦛 CHONK your texts with Chonkie ✨ Type-friendly, light-weight, fast and super-simple chunking library
At a glance
- What is it?
- chonkiejs ports the Python chonkie chunking library to TypeScript with a zero-dependency core. It is a good fit when you need chunkers inside a Node or browser app and a poor fit if you need the full Python feature set.
- Who is it for?
- Adopt chonkiejs if your retrieval pipeline runs in TypeScript and you want chunking in-process rather than a Python sidecar, and start with RecursiveChunker at a chunkSize you can measure against your own retrieval evaluation.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What chonkiejs solves, and who it is for
Retrieval augmented generation pipelines need text split into pieces small enough to embed and retrieve. Doing that splitting inside a TypeScript application usually means either calling out to Python or writing a splitter by hand. chonkiejs exists to remove both options. The README says the library was built while developing a TypeScript web app that needed fast, on-the-fly text chunking for RAG, and that the existing libraries the authors tried were either too heavy or not flexible enough. The result is a port of the Python chonkie library, not a binding to it, with type safety added for TypeScript developers.
The audience is narrow and specific. If your application is written in TypeScript and you want chunking to happen in the same process, this is the target case. If your pipeline already lives in Python, the original chonkie library is the direct route and this port adds nothing. The repository is a pnpm workspace with packages under packages/, and the published entry point is the @chonkiejs/core package on npm. The README notes plainly that the library is still under active development and not at feature parity with the Python original.
The chunker classes and the create-then-chunk pattern
Every chunker follows one shape: a static async create method returns an instance, then an async chunk method returns an array of chunks. The README gives the RecursiveChunker example with chunkSize 512, and each chunk exposes text and tokenCount. That uniformity is the main design decision in the library. You can swap a RecursiveChunker for a SentenceChunker without rewriting the calling code, because the surface is identical.
The chunkers differ in how they decide where a boundary falls. TokenChunker cuts at fixed token counts with an optional overlap. RecursiveChunker walks a hierarchy of rules, described in the README as paragraphs, then sentences, then punctuation, then words, then characters, which makes it the general-purpose choice. SentenceChunker groups whole sentences up to a token budget, with delimiters defaulting to '. ', '! ', '? ' and a newline, and an includeDelim option that attaches the delimiter to the previous chunk, the next chunk, or neither. SemanticChunker takes a required embeddings function and splits at points where similarity between sliding sentence windows drops below a threshold, with a similarityWindow default of 3. CodeChunker uses tree-sitter to split source code along AST boundaries, and TableChunker splits markdown or HTML tables into sub-tables that repeat the header row.
One detail matters more than the rest. The default tokenizer is the string 'character', which means token counts are character-based unless you supply something else. The README lists @chonkiejs/token as the package that adds HuggingFace tokenizer support, and it depends on @huggingface/transformers. So the zero-dependency claim in the package table applies to @chonkiejs/core only, and it holds only while you stay on the character tokenizer.
Installing chonkiejs and chunking a first document
Installation is a single npm command against the core package. The README gives no other prerequisite for basic use.
npm install @chonkiejs/coreWith the package installed, the README's usage example creates a RecursiveChunker with a chunk size of 512 and iterates the returned chunks, printing each chunk's text and token count.
import { RecursiveChunker } from '@chonkiejs/core';
const chunker = await RecursiveChunker.create({
chunkSize: 512
});
const chunks = await chunker.chunk('Your text here...');
for (const chunk of chunks) {
console.log(chunk.text);
console.log(`Tokens: ${chunk.tokenCount}`);
}Because create is async, it must be awaited before chunk is called. The README's CodeChunker example is the exception worth noting: it states that chunk is synchronous after create(), so a code chunker does not need an await on the second call. If you want a tokenizer other than the character default, the README points to @chonkiejs/token for HuggingFace tokenizer support, which is a separate install and brings @huggingface/transformers with it. The SemanticChunker path is different again: it requires you to pass an embeddings function, either a callback that maps an array of strings to an array of number arrays, or an object exposing an embed method.
Where chonkiejs is the wrong tool
The clearest limitation is stated by the project itself. The README says the library is a port rather than a binding, is still under active development, and is not yet at feature parity with the Python chonkie library. Anyone whose work depends on a specific chunker or option that exists in Python should check the TypeScript package before assuming it is present. The README does not enumerate which features are missing, so that check has to be done against the package's exports rather than the documentation.
The character tokenizer default is the second constraint. If your embedding model has a hard token limit and you are counting characters, chunkSize 512 does not mean 512 model tokens. The README documents @chonkiejs/token as the route to real tokenization, which means the lightweight path and the accurate-token-count path are different installs with different dependency footprints.
SemanticChunker has a sharper edge. It requires an embeddings function and calls it while chunking, so chunking cost scales with your corpus and your embedding provider, not with CPU time alone. For a one-off split of a static document set, that is an expensive way to draw boundaries. RecursiveChunker or SentenceChunker will usually do the job at a fraction of the cost, and the semantic split only pays off when the text has topic shifts that a punctuation rule cannot see. The README does not document any caching or batching behaviour for the embeddings callback, so treat that cost as yours to manage.
chonkiejs against the Python chonkie library
The obvious alternative is the original chonkie library in Python, and the difference is not cosmetic. The Python library is the source of the design; chonkiejs is a reimplementation of it, and the README is explicit that parity has not been reached. Choosing between them is mostly a question of where the rest of your pipeline runs. If your retrieval stack, embedding calls and evaluation harness are already Python, adding a Node process to chunk text buys you nothing and costs you a service boundary.
The second alternative is the @chonkiejs/cloud package in this same repository. The README's package table describes it as cloud-based chunkers, including Semantic, Neural and Code, reached through api.chonkie.ai, and it depends on @chonkiejs/core. The trade-off is direct: local chunking keeps text inside your process with zero dependencies, while the cloud package sends text to an external API and can offer chunker types the local package does not implement. For regulated content or air-gapped environments, the local core package is the only one of the two that applies. The README does not document pricing, rate limits or data retention for the cloud endpoint.
Maintenance, releases and the MIT licence
The repository is not archived, and the last push was on 2026-09-10, so the project is being worked on. Recent releases are versioned per package rather than as a single monorepo version: token-v0.0.4 on 2026-07-07, core-v0.0.11 on 2026-06-18, and core-v0.0.10 on 2026-05-12. The zero-point-zero version numbers are worth taking at face value. Under semantic versioning, anything before 1.0.0 can break in a minor release, and the README's own note about being under active development and short of feature parity reinforces that reading.
Upgrade cost is partly structural. The repository uses changesets, with scripts for changeset, version and release, so version bumps and changelogs are generated from changeset files rather than written by hand. The build is scoped: the root build script cleans package dist directories and then builds only packages/core, and the test script builds core and runs its tests with vitest. If you depend on @chonkiejs/cloud or @chonkiejs/token, note that the root scripts as written do not build them, so their release cadence is something to track separately. The licence is MIT, which is permissive and places few obligations on how you redistribute the library. That is a statement about the licence text in the repository, not legal advice; check the terms against your own distribution model.
Editorial conclusion
Adopt chonkiejs if your retrieval pipeline runs in TypeScript and you want chunking in-process rather than a Python sidecar, and start with RecursiveChunker at a chunkSize you can measure against your own retrieval evaluation. Do not adopt it if you need the full Python feature set, because the README states the port is not at feature parity, or if you need a tokenizer other than the default character tokenizer without pulling in @chonkiejs/token and its HuggingFace dependency. Before committing, verify that the chunker you intend to use is exported from @chonkiejs/core at the version you install, and check whether your token accounting should come from a real tokenizer instead of the character default.
Frequently asked questions
How do I install chonkiejs?
The README gives a single command, npm install @chonkiejs/core, for the local chunking package. The cloud chunkers and HuggingFace tokenizer support live in separate packages, @chonkiejs/cloud and @chonkiejs/token.
Does chonkiejs support semantic chunking?
Yes. SemanticChunker is exported from @chonkiejs/core and splits text where embedding similarity between sliding sentence windows drops below a threshold. It requires you to supply an embeddings function, either as a callback or as an object with an embed method.
Is chonkiejs the same as the Python chonkie library?
No. The README states that chonkiejs is a port of the Python chonkie library to TypeScript rather than a binding, and that it is still under active development and not at feature parity with the original.
Which tokenizer does chonkiejs use by default?
The default is the character tokenizer, passed as the string 'character' in chunker options, so token counts are character-based. The README points to the @chonkiejs/token package for HuggingFace tokenizer support.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/feyninc-chonkiejs)