Model or dataset
kkokosa/dotLLM avatar
kkokosa/dotLLM

dotLLM: a C# inference engine for .NET teams running Llama, Mistral, Phi and Qwen

LLM inference engine written in .NET

523 stars59 forksC#GPL-3.0

At a glance

What is it?
dotLLM implements model loading, tokenization, sampling and CPU compute in pure C#, with an optional CUDA backend. It is a preview-stage project, and the documentation is honest about what is not finished yet.
Who is it for?
Adopt dotLLM if you want LLM inference inside a .NET process without a llama.cpp wrapper, and you can work with a 0.1.0 preview that has no continuous batching and no LoRA yet. Do not adopt it if you need stable release numbering or multi-adapter serving today.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 63 days ago.
What is it written in?
Mainly C#, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What dotLLM solves, and who it is aimed at

Most ways to run a language model from .NET end in the same place: a P/Invoke boundary into llama.cpp, or an HTTP call to a Python process. dotLLM takes the other route. The README describes it as a ground-up engine where orchestration, model loading, tokenization, sampling and CPU compute are all implemented in pure C#, with CUDA acceleration through PTX kernels loaded via the CUDA Driver API rather than a bundled native library. That distinction matters if your deployment pipeline already ships .NET assemblies and you would rather not add a second runtime to it.

The target reader is a .NET engineer who needs a model running inside the same process as the rest of the application, or who wants to inspect the internals of inference rather than treat it as a black box. The repository layout supports that reading: DotLLM.Core holds the abstractions (ITensor, IBackend, IModel, ISamplerStep), and each layer above it depends only on the layers below. The samples directory includes console, server, logprobs, tool-calling and interpretability examples, which suggests the project expects people to read and modify code, not just call an endpoint.

How the engine is layered, and what each layer does

The architecture diagram in the README shows five layers. DotLLM.Core defines interfaces, tensor types and configuration. Above it, DotLLM.Models handles GGUF and SafeTensors parsing, DotLLM.Tokenizers implements BPE and SPM, and the Cpu and Cuda projects hold the kernels. DotLLM.Engine owns the KV-cache, scheduler, samplers, constraints and speculative decoding. DotLLM.Server exposes the ASP.NET OpenAI-compatible API.

Two design choices are worth calling out. First, attention variants are parameterized rather than hard-coded: MHA, MQA and GQA are selected through ModelConfig, with an IAttentionStrategy interface for kernel selection, and position encoding (RoPE, ALiBi, absolute, none) is pluggable through IPositionEncoding. Second, sampling is a composable chain of ISamplerStep implementations: repetition penalty, then temperature, then top-k, top-p, min-p, then a categorical sample. That ordering is fixed by the chain, so if you want a different order you are writing your own step rather than reconfiguring a pipeline.

Memory handling is the other visible mechanism. Tensor data lives in unmanaged memory via NativeMemory.AlignedAlloc with 64-byte alignment, and the README states there are no managed heap allocations on the hot path. GGUF files are loaded through MemoryMappedFile, so the operating system demand-pages multi-gigabyte weights instead of reading them eagerly. Quantized formats supported include FP16, Q8_0 and Q4_K_M, with fused scale-by-integer dot-product kernels operating directly on quantized blocks.

Installing dotLLM and running a first prompt

There are three documented install paths. The quickest is the global .NET tool, which requires the .NET 10 runtime. The README gives this sequence, which pulls a small GGUF model and then runs a completion against it:

bash
dotnet tool install -g DotLLM.Cli --prerelease

dotllm model pull QuantFactory/SmolLM-135M-GGUF

dotllm run QuantFactory/SmolLM-135M-GGUF -p "The capital of France is" -n 64

The model pull downloads the weights once and caches them, so the run command only needs the model identifier. The -n flag caps the number of generated tokens at 64. If the tool installs correctly, the run command prints a continuation of the prompt rather than an error about a missing model.

If you would rather not install a .NET runtime, the second path is a self-contained archive from the releases page, one per platform (win-x64, linux-x64, osx-arm64). On Linux or macOS the README unpacks it and invokes the bundled executable directly:

bash
tar -xzf dotllm-<version>-linux-x64.tar.gz
cd dotllm-<version>-linux-x64
./dotllm model pull QuantFactory/SmolLM-135M-GGUF
./dotllm run QuantFactory/SmolLM-135M-GGUF -p "The capital of France is" -n 64

To get an HTTP endpoint instead of a one-shot completion, the README uses the serve subcommand, which starts the OpenAI-compatible API together with a built-in chat UI:

bash
dotllm serve QuantFactory/SmolLM-135M-GGUF

The README does not state the port that serve binds to, so check the console output when the process starts. The third path is referencing the NuGet packages from your own .NET application; the README points to a NuGet Packages section further down, and each project in the layered architecture ships as a separate package so you pull in only the backend you need.

Structured output, speculative decoding and the preview caveats

The serving layer is where dotLLM has the most surface area. The OpenAI-compatible API covers /v1/chat/completions and /v1/completions, with tool calling and streaming. Constrained decoding uses an FSM or pushdown automaton to guarantee output that matches JSON, a JSON Schema, a regex or a grammar. That is a stronger guarantee than post-hoc validation: the sampler cannot emit a token that would break the grammar.

The KV-cache uses PagedAttention with block-level allocation, prefix caching and copy-on-write. Speculative decoding follows draft-verify-accept with KV-cache rollback, but the README limits it to greedy mode today, with non-greedy support planned and tracked in issue #121. If you are running temperature above zero, speculative decoding is not available to you yet.

The status line is the part to read carefully. The project is at Phase 6, and the release history reflects that: v0.1.0-preview.1, preview.2 and preview.3 all shipped in April 2026, and no stable release is listed. Continuous batching is explicitly planned for Phase 9, not implemented, which means concurrent requests are not scheduled at iteration level with preemption and priority queuing. LoRA adapters are also planned, in Phase 7, so runtime adapter loading without weight merging and concurrent multi-adapter serving are not available. Native AOT is described as experimental, and the README asks users to file an issue if an AOT build crashes. In other words, the engine is usable, but the throughput features that make a serving stack competitive are still on the roadmap.

Where dotLLM is the wrong choice

The clearest case against dotLLM is production serving under concurrency. Without continuous batching, the scheduler cannot pack requests from different clients into the same iteration, so a busy endpoint will not reach the utilization that vLLM-style servers get from iteration-level scheduling. The README names continuous batching as a Phase 9 item, so this is a known gap rather than an oversight, but it is a gap you would feel immediately under load.

The second case is model coverage. The README lists Llama, Mistral, Phi, Qwen and DeepSeek as transformer targets, and the status line names Llama, Mistral, Phi and Qwen among supported models. If your model architecture is not on that list, the parameterized TransformerBlock and ModelConfig will not save you without code changes. There is no documented mechanism for loading an arbitrary architecture at runtime.

The third case is anyone who needs a stable, versioned API. Everything published so far carries a preview suffix. If your release process cannot absorb breaking changes between preview builds, wait. A related consideration is the licence: dotLLM is GPL-3.0, which is a copyleft licence, and that is a different proposition from the permissive licences many .NET libraries use. Whether that matters depends on how you link and distribute, and it is a question for your own legal review rather than something the README resolves.

dotLLM compared with llama.cpp bindings

The obvious alternative for a .NET application is a binding over llama.cpp, such as LLamaSharp. The difference is not performance on paper; it is where the boundary sits. A binding keeps the C++ inference core and exposes it through interop, so you inherit llama.cpp's model support and its quantization work, and you also inherit a native binary in your deployment. dotLLM removes that binary. The README is explicit that the CUDA path uses PTX kernels through the CUDA Driver API with no native shared library, and that CPU compute is pure C#.

That trade cuts both ways. In exchange for a single managed deployment story, you take on a younger engine with a narrower model list and no continuous batching. You also gain something a binding cannot offer: the whole engine is C# source you can read and modify, with diagnostic hooks (IInferenceHook) intended for activation capture, logit lens and SAE integration, and a samples/DotLLM.Sample.Interpretability project in the repository. If your interest is in inspecting what happens inside the model rather than only getting tokens out, the interop route gives you far less to work with.

Maintenance, licensing and what to verify before adopting

The repository is not archived, and the last push was on 2026-07-30. The three releases listed are all preview builds from April 2026, so the versioning has not yet reached 0.1.0 stable despite months of commits since. The presence of a CONTRIBUTING file, a CODE_OF_CONDUCT, a SECURITY policy and a CI workflow suggests a project that expects outside contributions, and the docs directory holds a ROADMAP that the README links for the Phase 7 and Phase 9 items.

Upgrade cost is the practical question. Because every published version is a preview, you should expect API and behaviour changes between builds, and the engine layer (ISamplerStep, IBackend, IInferenceHook) is exactly where a preview project is most likely to shift. Pinning a specific preview version and reading the release notes before moving is the only sensible approach with the information available.

On licensing, the project is GPL-3.0, stated in the LICENSE file and shown in the README badge. That is a copyleft licence with distribution obligations that differ from permissive alternatives. Nothing here is legal advice, and the interaction between GPL-3.0 and your own distribution model is worth a review before you ship anything built on it.

Editorial conclusion

Adopt dotLLM if you want LLM inference inside a .NET process without a llama.cpp wrapper, and you can work with a 0.1.0 preview that has no continuous batching and no LoRA yet. Do not adopt it if you need stable release numbering or multi-adapter serving today. Before you commit, verify two things on your own hardware: that your target model architecture appears in the supported list, and that the CUDA path behaves as expected, since the README states the CUDA kernels load through the driver API with no native shared library.

Frequently asked questions

What is dotLLM?

It is an LLM inference engine written natively in C#/.NET, covering model loading, tokenization, sampling and CPU compute, with CUDA acceleration through PTX kernels loaded via the CUDA Driver API. The README describes it as a ground-up engine rather than a wrapper around llama.cpp or Python libraries.

How do I install dotLLM?

The README gives three options: install the DotLLM.Cli global tool with dotnet tool install -g DotLLM.Cli --prerelease (requires the .NET 10 runtime), download a self-contained archive for win-x64, linux-x64 or osx-arm64 from the releases page, or reference the NuGet packages from your own .NET application.

Which models does dotLLM support?

The status line names Llama, Mistral, Phi and Qwen as supported, and the architecture section also lists DeepSeek among transformer targets handled through a parameterized TransformerBlock and ModelConfig. Quantization formats include FP16, Q8_0 and Q4_K_M.

Does dotLLM have an OpenAI-compatible API server?

Yes. The DotLLM.Server layer exposes /v1/chat/completions and /v1/completions through ASP.NET, with tool calling and streaming, and the serve subcommand also starts a built-in chat UI. The README does not state which port the server binds to.

Does dotLLM support continuous batching?

No. The README lists continuous batching, with iteration-level scheduling, preemption and priority queuing, as a Phase 9 item that is planned rather than implemented. LoRA adapters are likewise planned for Phase 7.

What licence does dotLLM use?

GPL-3.0, shown in the README badge and stored in the LICENSE file at the repository root. That is a copyleft licence, so the distribution obligations differ from permissive alternatives.

Official sources

  1. Issues
  2. kkokosa/dotLLM on GitHub
  3. License: GPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kkokosa-dotllm.svg)](https://hysenlabs.com/projects/kkokosa-dotllm)