# kreeben/resin: a vector space search engine and key/value column store in C#

> Resin is a C# reboot of an older search engine project, combining a vector space search engine with an append-only key/value column store. It is for engineers who want to build lexicons and string-analysis pipelines without pulling in a large dependency tree.

**kreeben/resin** — Language model search engine built on a vector database and an anything key/value store.

- Repository: https://github.com/kreeben/resin
- Stars: 577 · Forks: 41
- Language: C#
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kreeben-resin

## What problem Resin addresses, and for whom

Resin is a reboot of an earlier project by the same author, and the README says so up front rather than presenting it as a clean-sheet design. The stated scope is three things stacked: a vector space search engine, a vector database, and an "anything key/value store". The phrase "anything" is doing real work here. Values are ReadOnlySpan<byte>, so the store does not care whether those bytes are a float vector, a UTF-8 string, or a serialized struct. The README claims it "can produce large language models out of strings and large anything models out of byte arrays", which is an ambitious framing for what the repository actually exposes: a column store, text analysis utilities, and a command line tool for lexicons.

The audience is narrow and identifiable. This is for a C# developer who needs on-disk key/value storage with predictable append behaviour, and who wants string and vector analysis in the same package rather than stitching together a database driver and a numeric library. The highlights list says "clean, dependency light design that is easy to extend", and the top-level repository layout supports that claim: a src/ directory, a couple of .bat launchers, and nothing else. If you are not in the .NET ecosystem, there is no binding mentioned and no HTTP service described, so the practical answer is that Resin is not for you.

## Columns, three streams, and why values are never rewritten

The storage model is the most concrete part of the README. A column is backed by three files with fixed roles. The .key stream holds sorted key representations in fixed-size slots, with sizeof(long) per entry for page-level storage, serialized in page batches. The .adr stream holds Address structs aligned with those key entries; each Address carries an Offset and a Length. The .val stream holds the actual value bytes.

Reads work by looking up the key in the column-wide snapshot, following the address entry, and slicing the value stream. Writes go through one of two methods. TryPut inserts only if the key does not already exist in the column-wide snapshot, returns false when it does, and triggers page serialization when the page fills. PutOrAppend takes the other path: if the key exists anywhere in the column, no new key is stored, and instead the new value is linked into the value stream using a fixed-size LinkedAddressNode. The original value stays first, appended values follow in insertion order, and the address entry for the key points at the list head once linking is active.

The design decision worth pausing on is that .val is append-only. Existing bytes are never modified in place. New values go at the end, which keeps previously written offsets valid, and the README notes the benefits directly: stable offsets make address caching safe, appending scales linearly, and historical values stay intact. Linking does not rewrite the old value either; a LinkedAddressNode header is appended and the previous node's NextOffset is patched. That is a coherent trade-off. You pay in read complexity, since a Get on a linked key has to walk a chain and concatenate bytes, and you gain a value stream with no fragmentation and no in-place mutation. For a store intended to be read heavily and appended to over time, the choice is defensible.

## The TKey contract is stricter than it first looks

The README devotes a full subsection to TKey restrictions, which is a signal that this is where implementations break. TKey must be a value type implementing both IEquatable<TKey> and IComparable<TKey>. Ordering and equality must be stable across sessions, because the column-wide key snapshot is built with BinarySearch over sorted keys. CompareTo must define a strict total order consistent with Equals. Those two requirements are easy to satisfy individually and easy to violate jointly, and the failure is not a compile error.

Page-level storage narrows further. It operates on long keys. double and float are stored through their IEEE bit representations, int and long are stored directly, and every other TKey type is hashed via GetHashCode() to a long. That last clause is the one to read twice. A custom struct key is not ordered on disk by its CompareTo; it is hashed. The README says collisions affect page-level operations for non-primitive keys, and recommends numeric primitives for deterministic ordering and lookup. If you arrive with a composite key struct and assume the column snapshot's sort order governs everything, the hashing step will surprise you. Columns are effectively sets of keys, since duplicate keys are prevented by both write methods, and the README points out that this enables union, intersection and joins across columns. That is a genuine capability, but it rests on the key contract holding.

## Building and running the Wikipedia command line tool

The README does not give a top-level install command. What it gives is a pointer: use Resin.KeyValue for on-disk structures and read/write sessions, use Resin.TextAnalysis for StringAnalyzer, VectorOperations and similarity tooling, and use Resin.WikipediaCommandLine for building and validating lexicons, with detailed CLI usage in src/Resin.WikipediaCommandLine/README.md. The repository root contains resin.bat, resin-test.bat and wikipedia.bat, which are the launchers the project ships.

The README does not document a NuGet package name or a version, so the traceable route is to clone the repository and build the solution under src/. The README itself does not spell out clone commands, so treat the repository URL and project path as something to confirm against your checkout.

The batch launchers at the root are the entry points the project provides. On Windows you would invoke them directly; the README does not describe a Linux or macOS equivalent.

```bash
wikipedia.bat
```

For programmatic use, the README names the three namespaces to reference. A first real use is a column you write keys into and read back, with the two write methods behaving differently on a repeated key. The README documents the signatures as TryPut(TKey key, ReadOnlySpan<byte> value) and PutOrAppend(TKey key, ReadOnlySpan<byte> value).

On the read side, Get returns the concatenated bytes of all linked values, and GetMany returns the same concatenation plus a count. A key that does not exist yields ReadOnlySpan<byte>.Empty, and GetMany reports count = 0. That empty-span convention means a missing key and a present-but-empty value are not distinguishable through Get alone.

## Where Resin is the wrong tool

The README describes storage semantics and a command line tool. It does not describe a query language, a ranking function, an index build pipeline, or an HTTP interface. For a project whose first sentence calls itself a search engine, that is a large gap. You can store vectors and you can call VectorOperations, but the README does not show how a query is expressed or how results are ordered. Anyone expecting to point Resin at a corpus and get ranked results back will be writing that layer themselves.

The key contract is the second sharp edge. Non-primitive keys are hashed to long for page-level operations, so a custom struct with a well-defined CompareTo still loses that ordering below the column snapshot, and GetHashCode collisions degrade page-level behaviour. If your keys are naturally strings, the README's recommendation of numeric primitives does not map onto your data without a mapping step you would have to design.

Third, the project is a reboot. The README links the previous repository at a specific commit and says this is a fresh start. There are no releases listed, and the README offers no migration notes from the old project. The last push to the repository was on 2026-06-12. There is no documented rollback story for the append-only value stream either; the README explains why values are never rewritten, not how you would remove one.

## How it differs from Lucene.NET and a plain embedded store

The obvious comparison in the .NET space is Lucene.NET. The difference is in what each one owns. Lucene.NET is an inverted-index library: you hand it documents, it builds term dictionaries and postings lists, and it gives you scoring, analyzers and a query parser. Resin inverts that. Its primary artefact is a column store where each column is effectively a set of keys, with union, intersection and joins available across columns, and the text side is StringAnalyzer plus VectorOperations for bags of words and chars. You compose retrieval from set operations and vector comparison rather than from a query parser. If your retrieval problem is naturally expressed as set algebra over keyed columns, Resin's model is closer to the problem than bending Lucene's document model to it. If you need scoring out of the box, Lucene.NET is the shorter path, and Resin's README does not claim to replace it.

The second comparison is a generic embedded key/value store such as SQLite or an LSM-tree library. Those give you a mutable store with a delete operation and a query language. Resin gives you an append-only value stream with stable offsets and a linking mechanism for multiple values per key, and no documented delete. The advantage is that addresses stay valid and caching is safe; the cost is that the store grows with every append and the README does not describe compaction.

## Licence and the cost of tracking a reboot

Resin is MIT licensed, which is permissive and imposes no copyleft obligation on your own code. The practical implication is that you can vendor the source or ship it inside a closed product, provided you keep the copyright notice and permission text with the distribution. That is a summary of the licence's shape, not legal advice; read the LICENSE file in the repository before you rely on it.

Upgrade cost is harder to estimate because there is nothing to upgrade from in a formal sense. No releases are listed, so consumption is by source or by building the solution yourself. The README points at src/Resin.WikipediaCommandLine/README.md for CLI detail, which means the CLI contract lives in a second document that can drift from the root one. The TKey restrictions are the part most likely to bite on an upgrade: they are stated as requirements rather than as versioned behaviour, so a change in how non-primitive keys are hashed or how the column snapshot is sorted would be a silent break for anyone relying on the current ordering. Pin to a commit rather than to a branch if you depend on it.

## Conclusion

Adopt Resin if you are working in C# and want a dependency-light column store plus string and vector analysis primitives you can extend, particularly for lexicon building or bag-of-words style work. Do not adopt it if you need a documented query language, a stable public API surface, or a search server you can point at a corpus and forget about; the README describes the storage semantics in detail but leaves the search layer at the level of "vector space search engine" with no query syntax shown. Before committing, verify three things: that your TKey type satisfies the struct, IEquatable and IComparable restrictions with a strict total order consistent with Equals, that you are comfortable with non-primitive keys being hashed to long for page-level storage, and that the Resin.WikipediaCommandLine README covers the corpus format you actually have.

## FAQ

### What is kreeben/resin?

It is a C# vector space search engine, vector database and key/value store, described in the README as a reboot of an older Resin search engine and machine learning project. It ships key/value column storage, text analysis utilities, and a command line tool for building and validating lexicons.

### What happens if I write the same key twice with kreeben/resin?

TryPut returns false and writes nothing when the key already exists in the column-wide snapshot. PutOrAppend stores no new key and instead links the new value into the value stream with a LinkedAddressNode, leaving the original value first and appending subsequent values in insertion order.

### Which key types can I use with kreeben/resin columns?

TKey must be a struct implementing IEquatable<TKey> and IComparable<TKey>, with CompareTo defining a strict total order consistent with Equals. Page-level storage works on long keys: double and float use their IEEE bit representations, int and long are stored directly, and any other type is hashed via GetHashCode.

### How do I install kreeben/resin?

The README does not give an install command or a package name. It points to Resin.KeyValue, Resin.TextAnalysis and Resin.WikipediaCommandLine under src/, and the repository root ships resin.bat, resin-test.bat and wikipedia.bat as launchers, with CLI detail in src/Resin.WikipediaCommandLine/README.md.

### Does kreeben/resin modify values in place?

No. The README states the .val stream is append-only: existing bytes are never modified in place, new values are written at the end, and linking appends LinkedAddressNode headers rather than rewriting existing values. The stated benefits are stable offsets for address caching and linear append scaling.

## Sources

- [Issues](https://github.com/kreeben/resin/issues)
- [kreeben/resin on GitHub](https://github.com/kreeben/resin)
- [License: MIT](https://github.com/kreeben/resin/blob/main/LICENSE)
- [README](https://github.com/kreeben/resin/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kreeben-resin
