wtpsplit ships two model families, and the smaller ones score higher
Toolkit to segment text into sentences or other semantic units in a robust, efficient and adaptable way.
At a glance
- What is it?
- Sentence segmentation from the SaT and WtP research code, packaged under the name of the older model. The published scores put the -sm variants above their full-size counterparts, and nothing on the page explains what sm stands for.
- Who is it for?
- This is a research codebase that happens to be packaged well enough to use, and the package is the better half: three install extras for the ONNX and AITune paths, a small API surface, and a published model table you can check before downloading weights. What to verify before building on it.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 11, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The package is named after the model the page tells you not to start with
The repository ships two implementations and recommends one of them. SaT, Segment Any Text, is the current approach and is marked state-of-the-art and encouraged. WtP, Where's the Point, is the previous version and is described as maintained for reproducibility. The project states plainly that the namesake WtP is kept for consistency and that the newer SaT work covers 85 languages at higher performance and lower compute cost.
So the distribution name, the import name and the recommended model all disagree. The package is wtpsplit, the import is from wtpsplit, the repository is segment-any-text/wtpsplit, and the checkpoint the usage example loads is sat-3l. A separate README_WTP.md sits at the root for the older line, and the paper links point at an arXiv entry for SaT and an ACL anthology entry for WtP. Anyone arriving from the paper title will search for the tool by the older name and land on a repository whose recommended model family has a different name, which is a small piece of friction that costs nothing to fix and is worth knowing about before you file an issue about it.
The -sm models are smaller and score higher, and the abbreviation is never expanded
The model table publishes an English score and a multilingual score for eleven checkpoints, and the pairs behave in a way nobody explains. The 1-layer model scores 88.5 and 84.3, while the 1-layer sm variant scores 88.2 and 87.9, losing a little on English and gaining more than three points on multilingual. The same inversion repeats upward: 3 layers scores 93.7 and 89.2 against 96.5 and 93.5 for the sm version, 6 layers scores 94.1 and 89.7 against 96.9 and 95.1, and 12 layers scores 94.0 and 90.4 against 97.4 and 96.0 for its sm twin, which is the top score in both columns.
So the recommended choice for speed, the 3-layer models, is also the weakest English performer among the 3-layer and larger entries, while the recommended best model is the 12-layer sm variant. That is a defensible set of tradeoffs, but the page never says what sm abbreviates, and the only place it appears in the usage section is a bare comment:
# use our '-sm' models for genThe scores themselves are described as macro-average F1 across a set of corpora, and that sentence stops after the word a. The eight corpora and the 85 languages are cited to the paper rather than enumerated here, so the table can be compared internally but not audited from this page alone.
optimize() has to come after to() and half(), and the first split pays for it
The faster PyTorch path is a compiler call rather than a flag, and the file is specific about how to use it. Call optimize() after to() and half(), so that the compiled graph is built for your device and dtype, and expect the first split afterwards to be slow while graphs are built. The call carries a comment naming the default backend, inductor, with dynamic shapes on.
Three caveats come with it. It works only with the SaT and WtP PyTorch checkpoints and is not available when a model is loaded through ort_providers for ONNX. The backend name accepts synonyms, inductor and torchinductor. And because chunk length and the last batch size change between calls, dynamic shapes are the default rather than an option: the reduced-overhead mode using CUDA graphs is only faster when every forward pass has the same shapes, otherwise the guidance is to stay on the default or switch to the no-cudagraphs tuning mode. On NVIDIA hardware there is one extra switch, setting float32 matmul precision to high before inference.
The TPU path is a footnote by comparison. Move the model with to("xla:0") and pass pad_last_batch=True into split when you do.
The ONNX section spells the provider argument two different ways
Two code blocks in the ONNX section use the same idea and different names. The performance example constructs the model like this:
sat = SaT("sat-3l-sm", ort_providers=["CUDAExecutionProvider", "CPUExecutionProvider"])The LoRA instructions later load a merged model with onnx_providers instead:
sat = SaT(<OUTPUT_DIR>, onnx_providers=["CUDAExecutionProvider", "CPUExecutionProvider"])Both blocks appear in the same section of the same page, and only the first spelling matches how the argument is introduced everywhere else. Whichever one your installed version accepts, the second is worth testing before you rely on the LoRA path, because the argument name is the only thing separating those two lines.
The LoRA flow itself has three parameters to choose: run the export script with use_lora set to true and an output directory, point lora_path at a local module or use style_or_domain and language to pull one from the model hub, then load the exported directory with the merged weights already inside it. The extras are named in the packaging as onnx-gpu and onnx-cpu, and the two installations are mutually exclusive in practice, which is why the plain onnxruntime requirement sits commented out in the dependency list with a note about conflicts between onnxruntime and onnxruntime-gpu.
The 144 ms against 94.9 ms number comes from one sentence repeated a thousand times
The ONNX section includes a timing comparison, and it is worth reading what it actually measures:
>>> from wtpsplit import SaT
>>> texts = ["This is a sentence. This is another sentence."] * 1000
# PyTorch GPU
>>> model_pytorch = SaT("sat-3l-sm")
>>> model_pytorch.half().to("cuda");
>>> %timeit list(model_pytorch.split(texts))
# 144 ms ± 252 μs per loop (mean ± std. dev. of 7 runs, 10 loops each)The input is one sentence pair repeated a thousand times, so the measurement covers throughput on trivially uniform text with the model already resident. The method is stated properly, seven runs of ten loops with a mean and a standard deviation, and the standard deviations are small next to the difference. What is not stated is the hardware, the batch shape or the backend, and the ONNX line's own timing comment ends without closing its parenthesis. The prose around it is candid about the gap, calling the PyTorch number quite fast already before the comparison. Treat the ratio as an indication that the ONNX path is worth trying, not as a benchmark you can quote.
Packaging is spread over three files and the dev group installs from an index
The build is defined in setup.py, which pins version 2.2.2 and requires Python 3.9 or newer. The install list is transformers, huggingface-hub, numpy, scikit-learn, tqdm, skops, pandas and mosestokenizer, and there are four extras: adapters, onnx-gpu, onnx-cpu and aitune. One detail in that list is a comment rather than a requirement, because onnxruntime is deliberately not a hard dependency since it conflicts with the GPU build.
Notably absent from both that list and the development requirements is any PyTorch entry, even though every usage example calls half() and moves the model to a device. What the development requirements do contain is a set of exact pins that do not match the runtime floors: transformers 4.29.2 against a floor of 4.22.2, scikit-learn 1.2.2, numpy 1.23.5, huggingface-hub 0.25.2 and protobuf 3.20.
The tool configuration lives in pyproject.toml, which contains no build settings at all: black and ruff both at line length 120, a ruff rule selection of four error families, E741 ignored, and a per-file exemption for the long lines in test.py. A comment explains that this is the pre-0.16 default because ruff 0.16 turns on 413 rules by default while CI installs the newest ruff. And the dependency group named dev contains a single entry, wtpsplit at 2.1.7 or newer, which means a development environment installs the published package rather than the checkout it sits in.
AITune adds a second index and benchmarks your GPU before it runs
The last backend is the one that changes your install. It comes from a separate index on NVIDIA's package host:
pip install wtpsplit[aitune] --extra-index-url https://pypi.nvidia.comThe extra pulls aitune and requests, and it needs Linux, CUDA and that separate install step, none of which applies to the other three extras. What it buys is automatic backend selection: it benchmarks TensorRT, Torch-TensorRT and Torch Inductor on the machine it is running on and picks a fast path. The default strategy is first-wins, trying TensorRT then Torch-TensorRT then Inductor, and there is a documented alternative that tunes Inductor only, with a batch count cap, for teams that do not want the TensorRT dependency.
The rest of the repository is a research layout rather than a library one. Alongside the package directory there are a demo script for length-constrained segmentation and a matching test file at the root, a configs directory, a scripts directory that holds the ONNX export script, test.py as the entry point, release.sh, RELEASE.md and a second README for the older model. The demo and its test sitting at the top level, next to the release script, is a reasonable choice for a paper repository and a slightly odd one for something people pip install.
Editorial conclusion
This is a research codebase that happens to be packaged well enough to use, and the package is the better half: three install extras for the ONNX and AITune paths, a small API surface, and a published model table you can check before downloading weights. What to verify before building on it. The model table is the reason to pick a checkpoint, and it contradicts the intuition that bigger is better, so pick from the numbers rather than the layer count. The API has one inconsistency worth checking against your installed version, since the provider argument is spelled two ways in the documentation. And the development environment installs the package from an index rather than the checkout, so a local test run can be measuring a published release instead of your working tree.
Frequently asked questions
What does wtpsplit do and which models does it load?
It segments text into sentences or other semantic units, and the recommended family is SaT with checkpoints such as sat-3l and sat-12l-sm. The older WtP models are kept for reproducibility, and SaT covers 85 languages at lower compute cost than the previous version.
How do you install wtpsplit with ONNX or AITune support?
The base install is pip install wtpsplit. ONNX comes as two mutually exclusive extras, wtpsplit[onnx-gpu] and wtpsplit[onnx-cpu], and AITune is installed separately with pip install wtpsplit[aitune] --extra-index-url https://pypi.nvidia.com, which needs Linux and CUDA.
Which wtpsplit model should be used for speed?
The page recommends the 3-layer models as a tradeoff between speed and performance and names the 12-layer sat-12l and sat-12l-sm as the best. In the published table the sm variants score above their full-size counterparts, with sat-12l-sm at 97.4 English and 96.0 multilingual.
When should wtpsplit optimize() be called?
After to() and half(), so the compiled graph matches your device and dtype, and expecting the first split to be slow while graphs are built. It works with the PyTorch checkpoints only, not with ONNX providers, and dynamic shapes are on by default because chunk length and last batch size vary between calls.
Can wtpsplit be used without a GPU?
Yes. There are CPU extras for ONNX, the CPU execution provider is listed alongside the CUDA one, and batching multiple texts into a single split call is the documented way to get better throughput on any device.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/segment-any-text-wtpsplit)