# Kalosm: Local AI Inference for Rust with Structured Generation

> Kalosm is an Apache-licensed Rust library that lets applications run language, audio, and image models locally, including a structured generation engine that constrains output to native Rust types. The repository also contains Fusor, a compiler-backed inference runtime for CPUs and WebGPU that is still in early development.

**floneum/kalosm** — Instant, controllable, local pre-trained AI models in Rust

- Repository: https://github.com/floneum/kalosm
- Website: http://floneum.com/kalosm
- Stars: 2,229 · Forks: 135
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/floneum-kalosm

## What Kalosm Provides for Rust Applications

Kalosm addresses a specific gap: most production AI inference tooling targets Python, leaving Rust developers who want to run models locally without a clean, typed API. The library wraps pre-trained language, audio, and image models behind a consistent Rust interface, handling model downloads, tokenisation, and inference internally. Applications add Kalosm as a dependency and call high-level methods without managing model files or C bindings directly.

The scope covers three modalities. For text, Kalosm supports Llama, Mistral, Phi, and BERT. For audio, it wraps Whisper for transcription from microphone or file input. For images, it wraps Segment Anything for segmentation. Every model in the current table supports quantisation and GPU acceleration, according to the README model matrix.

Beyond raw inference, Kalosm provides utilities for the full retrieval pipeline: extracting context from txt, html, docx, md, and pdf files; chunking that context; and searching it through vector database integrations. Web crawling and scraping are also included. These utilities position Kalosm as a base layer for building local retrieval-augmented applications in Rust.

## The Structured Generation Engine

Structured generation is one of the features that distinguishes Kalosm from a plain model wrapper. The library uses a custom parser engine and sampler so that any Rust type annotated with the Parse and Schema derive macros can be used as an output target, constraining what the model is allowed to generate token by token.

The README shows a Character struct with fields for name, age, and description, each annotated with a regex pattern or numeric range. Kalosm then ensures the model only produces output that satisfies those constraints:

```rust
use kalosm::language::*;

/// A fictional character
#[derive(Parse, Schema, Clone, Debug)]
struct Character {
    /// The name of the character
    #[parse(pattern = "[A-Z][a-z]{2,10} [A-Z][a-z]{2,10}")]
    name: String,
    /// The age of the character
    #[parse(range = 1..=100)]
    age: u8,
    /// A description of the character
    #[parse(pattern = "[A-Za-z ]{40,200}")]
    description: String,
}
```

The README notes that beyond regex, custom grammars can be provided to constrain output to arbitrary structures including JSON, HTML, and XML. The library also claims that structure-aware acceleration makes constrained generation faster than unconstrained generation, though the README does not give benchmark figures.

## Starting a Kalosm Project

Kalosm requires a stable Rust toolchain installed via rustup. The quickstart in the README builds a chatbot backed by a Phi-3 model.

Create a new project and add dependencies:

```sh
cargo new kalosm-hello-world
cd ./kalosm-hello-world
```

```sh
cargo add kalosm --features llama
cargo add tokio --features full
```

The README provides a complete main.rs that loads the model and enters a chat loop:

```rust
use kalosm::language::*;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
  let model = Llama::phi_3().await?;
  let mut chat = model.chat()
    .with_system_prompt("You are a pirate called Blackbeard");

  loop {
    chat(&prompt_input("\n> ")?)
      .to_std_out()
      .await?;
  }
}
```

Run with the release profile for meaningful inference speed:

```sh
cargo run --release
```

On first run, Kalosm downloads the quantized model checkpoint. The download size for Phi-3 is in the low-gigabyte range given the model sizes listed in the README table. The async runtime requirement means all Kalosm operations are written in async Rust with tokio.

A more complete guide and additional examples covering transcription, image segmentation, and semantic search are in the examples folder and on the Kalosm website.

## Constraints and Cases Where Kalosm Is the Wrong Tool

The Rust-only surface is a hard constraint. Python developers, data scientists using Jupyter notebooks, or teams whose existing AI pipeline is in LangChain or Hugging Face Transformers cannot use Kalosm directly. The library has no Python bindings.

Fusor, the inference backend embedded in the same repository, is explicitly marked in the README as not ready for production. The note reads: 'Fusor is still early in development and is not ready for production use.' Fusor handles GGUF model loading and uses an e-graph compiler to fuse operation chains, targeting CPU and WebGPU execution. Until Fusor matures, the stability guarantee on the inference layer is limited.

Model support is bounded by what is in the current table. The README lists six model families as of the current release. Applications that need a model outside that list, or that require fine-tuned variants or specific quantisation levels not available in the pre-built checkpoints, will need to verify compatibility before starting development. The workspace Cargo.toml pins kalosm at version 0.4.0; API stability across minor versions is not guaranteed for a project at this stage.

## Kalosm Compared to Candle

Candle is a minimalist ML framework for Rust developed at Hugging Face, aimed at inference with low overhead and no Python runtime dependency. The key difference in approach is the level of abstraction. Candle operates closer to the tensor level, giving developers direct control over model architecture and computation, which is useful when working with custom or experimental models. Kalosm operates at the application API level, providing ready-made model wrappers and higher-level utilities such as the structured generation engine, the retrieval pipeline, and the audio transcription interface.

Candle does not include built-in structured generation constrained by Rust's type system; Kalosm's Parse and Schema derive macros are not available in Candle. Candle gives access to a broader set of Hugging Face model checkpoints. Kalosm trades that breadth for a more complete application-layer toolkit. A developer building a custom model architecture would start with Candle. A developer who wants to add a local chatbot or transcription feature to a Rust application with minimal model-level code would start with Kalosm.

## Fusor and the Inference Backend

The repository contains two separate projects. Kalosm is the application interface described above. Fusor, in the fusor/ subdirectory, is an independent workspace and the local inference backend Kalosm's model crates use internally. Its README section describes it as a compiler-backed CPU and WebGPU runtime for quantized ML inference. It loads GGUF model files and uses an e-graph compiler to fuse operation chains into optimized kernels without requiring model authors to write GPU shader code.

The README gives this Rust expression as an example of what Fusor can compile into a single kernel:

```rust
fn exp_add_one(tensor: Tensor<2, f32>) -> Tensor<2, f32> {
  1. + (-tensor).exp()
}
```

Because Fusor excludes itself from the parent workspace with the exclude directive in Cargo.toml, it is versioned and built independently. The 'not ready for production' warning means that teams deploying Kalosm should expect Fusor's API and performance characteristics to change. The last push to the repository was on 2026-09-27, indicating the project is under active development.

## Conclusion

Kalosm is the right choice for Rust developers who need to run language, audio, or image models locally without calling a cloud API, and who want guarantees about output structure that are expressed directly in the type system. Applications that need Python-ecosystem tooling, Hugging Face integrations, or stable production ABI guarantees will find the Rust-only surface a real constraint. Fusor, the underlying inference backend, is explicitly marked as not ready for production. Verify that the model variants you need are in the current model table and that the quantized checkpoints are available before starting integration.

## FAQ

### How do I add Kalosm to an existing Rust project?

The README quickstart shows adding Kalosm with cargo add kalosm --features llama and cargo add tokio --features full. The llama feature flag is required when using Llama or Phi models; other feature flags cover different model families. The workspace Cargo.toml lists the current version as 0.4.0.

### Does Kalosm require a GPU to run models?

The model table in the README marks every listed model as both quantized and GPU-accelerated, but quantized models are designed to run on CPU as well. The Fusor backend targets both CPU and WebGPU, and the README quickstart does not mention a GPU requirement for the Phi-3 chatbot example.

### What is Fusor and how does it relate to Kalosm?

Fusor is a separate compiler-backed inference runtime in the same repository that Kalosm's model crates use as their local backend. It loads GGUF model files and fuses operation chains using an e-graph compiler. The README explicitly marks Fusor as not ready for production use.

## Sources

- [floneum/kalosm on GitHub](https://github.com/floneum/kalosm)
- [License: Apache-2.0](https://github.com/floneum/kalosm/blob/main/LICENSE)
- [Project website](http://floneum.com/kalosm)
- [README](https://github.com/floneum/kalosm/blob/main/README.md)
- [Releases](https://github.com/floneum/kalosm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/floneum-kalosm
