# catalyst: A Fast, Pure-C# NLP Library for .NET 9 and .NET 10

> catalyst is a C# NLP library from curiosity-ai, inspired by spaCy's pipeline design and targeting the .NET 9 and .NET 10 runtimes. It provides tokenization above 1 million tokens per second, three separate named entity recognition strategies, and NuGet-distributed language model packages from the Universal Dependencies project. The primary trade-off is that it targets only modern .NET runtimes and does not support LLM-style transformer models.

**curiosity-ai/catalyst** — 🚀 Catalyst is a C# Natural Language Processing library built for speed. Inspired by spaCy's design, it brings pre-trained models, out-of-the box support for training word and document embeddings, and flexible entity recognition models.

- Repository: https://github.com/curiosity-ai/catalyst
- Stars: 860 · Forks: 86
- Language: C#
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/curiosity-ai-catalyst

## What catalyst Is and Who It Targets

Most production NLP pipelines in the Python world use spaCy or Hugging Face Transformers. C# developers working in .NET have historically had to call Python services or use thin wrappers around Java-based libraries. catalyst fills this gap: it is a pure C# NLP library that runs anywhere .NET Core runs, including Windows, Linux, macOS, and ARM devices.

The library targets .NET 9 and .NET 10. Support for .NET Standard 2.1 and .NET 5 through .NET 8 has been dropped in the current version. This means catalyst is the right choice for new projects or projects upgrading to modern .NET, and the wrong choice for applications that must remain on older runtimes.

The target use case is production NLP workloads: tokenization of large text volumes, named entity recognition with trainable models, part-of-speech tagging, word and document embedding, and language detection. It is explicitly not a library for conversational AI, LLM inference, or fine-tuning transformer models.

## Non-Destructive Tokenization and the Performance Design

catalyst's tokenizer is non-destructive, meaning the original text is preserved and token boundaries are tracked by offset rather than by copying substrings. The README claims tokenization above 1 million tokens per second on a modern CPU. More than 99.9% of the tokenizer is RegEx-free by design, which eliminates the performance variability and backtracking overhead that comes with heavy RegEx use in NLP pipelines.

Binary serialization uses MessagePack throughout. This keeps model loading fast and model files compact, with no dependency on JSON or XML deserialization during inference. Models can also be loaded from and stored to streams, which makes it practical to load language models from embedded resources or remote storage rather than always from disk:

```csharp
using(var f = File.OpenWrite("my-pattern-spotter.bin"))
{
    await isApattern.StoreAsync(f);
}
```

The performance claims come from the library's documentation rather than from independent benchmarks. The design choices (RegEx-free tokenization, MessagePack serialization, offset-based token representation) are architectural decisions that support high throughput, but actual latency in a given application will depend on the complexity of the pipeline and the hardware available.

## Named Entity Recognition: Three Separate Strategies

catalyst provides three distinct approaches to named entity recognition, each suited to different trade-offs between setup cost, coverage, and adaptability.

The gazeteer approach uses a lookup table (Spotter) to match entities against a predefined list. This is fast and requires no training, but coverage is limited to what is in the list.

The rule-based approach uses PatternSpotter, which matches sequences of tokens by combining syntactic patterns. A pattern is defined by chaining PatternUnit constraints:

```csharp
var isApattern = new PatternSpotter(Language.English, 0, tag: "is-a-pattern", captureTag: "IsA");
isApattern.NewPattern(
    "Is+Noun",
    mp => mp.Add(
        new PatternUnit(P.Single().WithToken("is").WithPOS(PartOfSpeech.VERB)),
        new PatternUnit(P.Multiple().WithPOS(PartOfSpeech.NOUN, PartOfSpeech.PROPN, PartOfSpeech.AUX, PartOfSpeech.DET, PartOfSpeech.ADJ))
));
```

The perceptron-based approach uses AveragePerceptronEntityRecognizer, which learns from annotated examples and generalizes beyond fixed lists and patterns. Training this model requires labeled data but produces a model that handles unseen text better than either of the static approaches.

## Getting Started: Language Packages, DiskStorage, and the NLP Pipeline

Language-specific data and pre-trained models are distributed as separate NuGet packages. The core package is Catalyst, and each language adds a Catalyst.Models.* package. Both must be on compatible versions: upgrading one without the other prevents the language package from loading.

A minimal pipeline for English requires registering the language, configuring storage, and running a document through the pipeline:

```csharp
Catalyst.Models.English.Register();
Storage.Current = new DiskStorage("catalyst-models");
var nlp = await Pipeline.ForAsync(Language.English);
var doc = new Document("The quick brown fox jumps over the lazy dog", Language.English);
nlp.ProcessSingle(doc);
Console.WriteLine(doc.ToJson());
```

Calling Pipeline.ForAsync triggers lazy model loading: models are downloaded from the online repository on first use and cached to disk under the path given to DiskStorage. Subsequent calls load from disk without a network request.

For large document collections, the pipeline supports C# lazy evaluation and native multi-threading through the IEnumerable overload:

```csharp
var docs = GetDocuments();
var parsed = nlp.Process(docs);
```

Processing happens across threads as the enumerable is consumed, without requiring explicit parallelism code.

## Training FastText Word Embeddings

catalyst includes out-of-the-box support for training FastText and StarSpace word and document embeddings. A FastText model is trained by constructing the model, setting its type and loss function, and passing a stream of processed documents:

```csharp
var nlp = await Pipeline.ForAsync(Language.English);
var ft = new FastText(Language.English, 0, "wiki-word2vec");
ft.Data.Type = FastText.ModelType.CBow;
ft.Data.Loss = FastText.LossType.NegativeSampling;
ft.Train(nlp.Process(GetDocs()));
ft.StoreAsync();
```

The FastText integration is part of the catalyst package, not a separate library. Pre-trained embedding models are listed as coming soon in the README; the current release supports training from scratch.

For fast embedding search, the curiosity-ai team has also released a C# implementation of the Hierarchical Navigable Small World (HNSW) algorithm on NuGet, and a C# implementation of UMAP (Uniform Manifold Approximation and Projection) for dimensionality reduction on GitHub. The README mentions these as companion tools, not as dependencies of catalyst itself.

## Breaking Changes in the Current Version and the spaCy Comparison

The current version includes two significant breaking changes. First, ILemmatizer now takes ReadOnlySpan<char> instead of IToken. The README explains that the previous interface allocated 72 bytes per token due to boxing, because Token is a struct. Any custom ILemmatizer implementation must be updated. The English.Map.ToAmerican and ToBritish methods narrow to the same span signature, though a string still works through implicit conversion.

Second, netstandard2.1, netcoreapp3.1, and net5.0 through net8.0 are no longer built. catalyst and all Catalyst.Models.* packages now target net9.0 and net10.0 only. Projects on .NET 8 or earlier cannot upgrade to the current version without also upgrading their target framework.

Comparing catalyst to spaCy, which the README cites as its design inspiration: spaCy is a Python NLP library maintained by Explosion AI, well known for its production-ready pipeline design, fast tokenizer, and pre-trained statistical models. It runs in Python and requires no .NET runtime. catalyst takes the same pipeline concept and implements it in pure C# for the .NET ecosystem. The choice between them is determined almost entirely by which runtime the rest of your application uses: Python teams use spaCy; .NET teams use catalyst.

The last push to the repository was on 2026-09-24. The project is MIT-licensed.

## Conclusion

catalyst is the right library for .NET 9 and .NET 10 applications that need fast, production-grade NLP without leaving the C# ecosystem. It is not suitable for projects running on .NET Standard, .NET 5 through .NET 8, or for use cases that require transformer-based or LLM-style models. Before adopting it, run the language model upgrade check: catalyst language packages and the core library must be upgraded together, because a Catalyst.Models.* package built against an older catalyst version will not load.

## FAQ

### Which .NET versions does catalyst support?

The current version targets net9.0 and net10.0. Support for .NET Standard 2.1 and .NET 5 through .NET 8 has been removed. The README states that catalyst and its Catalyst.Models.* language packages must be upgraded together, since a language package built against an older version of catalyst will not load.

### How are language models installed in catalyst?

Language models are distributed as separate NuGet packages named Catalyst.Models.*. Registering the language with a call like Catalyst.Models.English.Register() and setting a DiskStorage path triggers lazy model download from the online repository on first use. Subsequent pipeline calls load from the local cache.

### Does catalyst support transformer or LLM-style models?

The README does not describe transformer-based or LLM-style model support. catalyst's capabilities cover tokenization, named entity recognition (gazeteer, pattern, and perceptron), part-of-speech tagging, language detection, lemmatization, and FastText and StarSpace word embeddings. It is designed for traditional NLP pipelines, not for large language model inference.

## Sources

- [curiosity-ai/catalyst on GitHub](https://github.com/curiosity-ai/catalyst)
- [Issues](https://github.com/curiosity-ai/catalyst/issues)
- [License: MIT](https://github.com/curiosity-ai/catalyst/blob/master/LICENSE)
- [README](https://github.com/curiosity-ai/catalyst/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/curiosity-ai-catalyst
