huggingface/candle: a Rust ML framework for people who want tensors, not a Python runtime
Minimalist ML framework for Rust
At a glance
- What is it?
- Candle is a minimalist machine learning framework for Rust, published as a workspace of crates with CPU, CUDA and Metal backends. It is aimed at Rust developers who want inference and tensor math in-process, and it is a poor fit for anyone who wants a training ecosystem.
- Who is it for?
- Adopt candle if you are shipping Rust binaries that need tensor math or model inference and you are willing to read example source instead of a tutorial. Do not adopt it if you need a training loop, a model zoo you can call by name, or a stable API across versions.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What candle actually solves for a Rust codebase
Most machine learning work happens in Python, which means any Rust service that needs a model has to either shell out to a Python process, run an ONNX runtime, or reimplement inference. Candle offers a fourth option: a tensor library and a set of model implementations that compile into your Rust binary. The README describes it as a "minimalist ML framework for Rust with a focus on performance (including GPU support) and ease of use."
The intended user is a Rust engineer who needs matrix multiplication, autograd-free inference, or a specific model (Whisper, LLaMA, YOLO, Stable Diffusion) inside an application, a CLI, or a WebAssembly target. The repository ships candle-wasm-examples and candle-wasm-tests as workspace members, so browser deployment is a first-class target rather than an afterthought. If your problem is "I want to call a model from Rust without a Python sidecar," candle is aimed at you. If your problem is "I want to fine-tune a model and track experiments," it is not.
The crate layout: core, nn, transformers, and what sits outside the workspace
Candle is a Cargo workspace, and the members list in the root Cargo.toml is the clearest statement of what is supported. Inside the workspace: candle-core, candle-datasets, candle-examples, candle-nn, candle-pyo3, candle-transformers, candle-wasm-examples/*, candle-wasm-tests, and tensor-tools. Outside it, in the exclude list: candle-book, candle-flash-attn, candle-flash-attn-v3, candle-kernels, candle-metal-kernels, and candle-onnx.
That split matters when you plan a build. candle-core holds the Device and Tensor types. candle-nn holds layers and optimizers. candle-transformers holds the model implementations. The excluded crates are the ones that need special build handling: the flash-attention crates, the CUDA kernel generation crate, the Metal kernel crate, and the ONNX loader. The Makefile exposes a clean-ptx target that deletes generated .ptx files and resets candle-kernels/src/lib.rs, which tells you kernel artifacts are generated at build time rather than checked in. Every crate in the workspace carries version 0.11.0 and the licence string "MIT OR Apache-2.0".
One design consequence: because the attention and kernel crates sit outside the workspace, a plain `cargo build` at the root does not compile them. You opt in by depending on them directly.
Installing candle and running your first matmul
The README points to the Installation page of the project's documentation site for getting candle-core set up, and the root Cargo.toml pins candle-core at version 0.11.0. The README's own example is a matrix multiplication, written into myapp/src/main.rs. It creates two random tensors on the CPU device and multiplies them:
use candle_core::{Device, Tensor};
fn main() -> Result<(), Box<dyn std::error::Error>> {
let device = Device::Cpu;
let a = Tensor::randn(0f32, 1., (2, 3), &device)?;
let b = Tensor::randn(0f32, 1., (3, 4), &device)?;
let c = a.matmul(&b)?;
println!("{c}");
Ok(())
}According to the README, `cargo run` should display a tensor of shape `Tensor[[2, 4], f32]`. Note the error type: every tensor operation returns a Result, so you propagate errors with `?` rather than unwrapping. That is the pattern throughout the API.
Moving to a GPU is a one-line change once you have installed candle with CUDA support, and the README shows it as a diff:
- let device = Device::Cpu;
+ let device = Device::new_cuda(0)?;The `0` is the device ordinal. The README does not document a Metal equivalent in the same diff form, so if you are on Apple hardware, check the installation page before assuming the same call shape.
The examples directory is the real documentation
The README lists roughly thirty command-line examples under candle-examples/examples/, covering LLaMA v1 through v3, Falcon, Codegeex4, GLM4, Gemma v1 and v2, RecurrentGemma, Phi-1 through Phi-3, StableLM, Mamba, Mistral, Mixtral, StarCoder and StarCoder2, Qwen1.5, RWKV v5 and v6, Replit-code, Yi-6B and Yi-34B, quantized LLaMA, quantized Qwen3 MoE, Stable Diffusion, Wuerstchen, yolo-v3 and yolo-v8, segment-anything, SegFormer, Whisper, EnCodec, MetaVoice, and Parler-TTS.
That list is broad, and it is also the honest measure of what candle gives you. There is no unified model-loading API described in the README that lets you name a model and get a pipeline. Each example is a directory with its own binary, its own argument parsing, and its own weight-download logic. If you want to run Mixtral, you read the Mixtral example and adapt it. That is a deliberate trade: the framework stays small, and the examples absorb the per-model complexity.
The quantized LLaMA example is worth calling out because the README states it uses the same quantization techniques as llama.cpp, and the quantized Qwen3 MoE example supports gguf quantized models. That gives you a path to running GGUF weights without a C++ dependency, which is a concrete reason to pick candle over writing your own loader.
Where candle is the wrong tool
The README describes candle as an ML framework, but the examples it advertises are overwhelmingly inference: speech recognition, text generation, object detection, image segmentation, text-to-image, text-to-speech. The Mamba entry is explicitly labelled "an inference only implementation." Nothing in the README presents a training loop, a dataset pipeline beyond the candle-datasets crate name, a distributed training story, or a checkpointing format.
So if your task is training or fine-tuning, candle is the wrong tool, and the README does not pretend otherwise. The absence of a published release list in the repository metadata is a second limitation: the workspace version is 0.11.0, and a 0.x version number means the API can change between minor releases. Pinning an exact version is not optional if you have production code.
A third constraint is build surface. The flash-attention and kernel crates are excluded from the workspace and require their own build steps, and the Makefile's clean-ptx target exists specifically because those artifacts go stale. On a CI machine without the right toolchain, you will be debugging the build before you debug your model.
How candle differs from burn and tch-rs
Two other Rust options come up in the same conversation. Burn takes the opposite architectural bet: it treats backends as pluggable and puts training, autodiff, and a unified module system at the centre, which means more abstraction between you and the kernel. Candle keeps the tensor type close to the metal and pushes model-specific code into examples. If you want autodiff and a training story, Burn's design targets that directly; candle's README does not claim it.
tch-rs is a different comparison. It binds to libtorch, so you inherit PyTorch's operator coverage and its C++ shared library as a runtime dependency. Candle has no such dependency: it is pure Rust plus generated kernels, which is why the WebAssembly examples can exist at all. The cost is that you get candle's operator set, not PyTorch's, and you get candle's model implementations, not Hugging Face Transformers' full catalogue. The choice is roughly: bind to a large C++ library, or accept a smaller pure-Rust surface with a smaller model list.
Licence and the cost of tracking 0.x releases
The workspace declares `license = "MIT OR Apache-2.0"`, and the repository carries LICENSE-MIT and LICENSE-APACHE at the top level. That dual licence is the same arrangement Rust itself uses, and it lets a downstream user pick either set of terms. It is permissive in both branches, so it does not impose copyleft obligations on your application. This is a description of what the files say, not legal advice; if your organisation has specific licence review requirements, the two files in the repository root are what your reviewers will want to read.
Upgrade cost is dominated by the 0.x version. The workspace version is 0.11.0, and the Cargo.toml uses path dependencies with matching version strings for every internal crate, so the pieces move together. A minor bump can change tensor method signatures or example argument parsing. The practical approach is to pin the exact version in your own Cargo.toml and treat upgrades as a scheduled task with a diff review, rather than letting `cargo update` pull a new minor version into a build you did not test. The CHANGELOG.md at the repository root is where the project records what changed, and it is the first file to read before bumping.
Editorial conclusion
Adopt candle if you are shipping Rust binaries that need tensor math or model inference and you are willing to read example source instead of a tutorial. Do not adopt it if you need a training loop, a model zoo you can call by name, or a stable API across versions. Before committing, verify three things in the repository: whether the example you need is in the workspace members list or in the excluded list, which backend feature flags your target platform documents, and whether the crate version you pin matches the 0.11.0 workspace version.
Frequently asked questions
What is huggingface/candle?
It is a minimalist machine learning framework for Rust, described in its README as focused on performance including GPU support, and distributed as a Cargo workspace of crates such as candle-core, candle-nn, and candle-transformers.
How do I install huggingface/candle and run a first example?
The README says to install candle-core as described on the project's Installation documentation page, then write a main.rs that builds two tensors on Device::Cpu and calls matmul. Running cargo run should print a tensor of shape Tensor[[2, 4], f32].
Can huggingface/candle run models on a GPU?
Yes. The README states that after installing candle with CUDA support you replace Device::Cpu with Device::new_cuda(0), where 0 is the device ordinal. The README's diff example covers CUDA only.
Does huggingface/candle support training models?
The README does not present a training loop or fine-tuning workflow. Its examples are inference oriented, and the Mamba entry is explicitly described as an inference only implementation.
What licence does huggingface/candle use?
The workspace Cargo.toml declares license = "MIT OR Apache-2.0", and the repository root contains LICENSE-MIT and LICENSE-APACHE.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/huggingface-candle)