Model or dataset
lucasjinreal/Crane avatar
lucasjinreal/Crane

Crane's description promises VLA and its own model list leaves every VLA model unchecked

A Pure Rust based LLM, VLM, VLA, TTS, OCR Inference Engine, powering by Candle & Rust. Alternate to your llama.cpp but much more simpler and cleaner..

488 stars59 forksRustMIT

At a glance

What is it?
A Rust inference engine built on Candle with four GPU backends, where the same Candle crates are depended on twice under two names, two backends are forks the manifest says are meant to disappear, compose.yaml hardcodes host GIDs and starts nothing without a profile, and one changelog entry admits the real 18.7 GB checkpoint was never loaded.
Who is it for?
Crane is worth reading if you want Candle's kernels without writing Candle plumbing, and the parts that are finished are finished properly: ternary GGUF loading with automatic format detection, quantized embedding tables that save 1772 MiB on a 27B Q4_K_M, per-request thinking control, and a changelog with real numbers in it. Two things to weigh first.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The description lists VLA and the model list leaves every VLA model unchecked

The repository description calls Crane an LLM, VLM, VLA, TTS and OCR inference engine. The supported models list is a checklist, and reading it as a checklist rather than a headline changes the picture. Checked are PaddleOCR-v6, the Bonsai 2 Ternary 27B, Qwen3.8-Flash-Next, Qwen 3.6 and 3.8, Qwen 3.5 with an Ornith agentic model, Hunyuan Dense, Gemma 4 for text and vision, Qwen3 VL at 2B and 4B, PaddleOCR VL, Qwen3 and Qwen 2.5 across their size ranges, Moonshine ASR, Silero VAD, and four text-to-speech or audio entries. Unchecked are Qwen3.5-VLA, Qwen3.5-GR00T and Pi0.5, plus a line that says only more to come, and a struck-through line for two TTS projects marked as work in progress. The feature list agrees with the checkboxes rather than the description, listing VLA with the note that it is on the way.

Candle is depended on twice under two names, and two backends are forks

The workspace manifest explains its own dependency shape in comments, which is unusual and useful. Plain `candle-core`, `candle-nn` and `candle-transformers` are pinned at 0.11, and then the same three crates are declared again as `candle-core-default`, `candle-nn-default` and `candle-transformers-default`, each with `package` set back to the upstream name and the same version, used by `crane-core` in place of the bare names. The reason given is that the mainstream Candle keeps its upstream name while only the forks that add a backend Candle lacks yet carry a suffix, and that each fork, along with the SYCL and ROCm special-casing in `candle_backend.rs`, is meant to go away once that backend merges upstream. So two of the four execution backends are patches to Candle rather than features of Candle, and the manifest is candid that they are temporary. A third direct dependency, `candle-metal-kernels`, exists because Candle's own `metal_backend` module re-exports only three items and the fused Metal kernels have to be named explicitly.

compose.yaml carries host GIDs and starts nothing without a profile

The compose file runs one service per GPU backend, selected by profile, and the header comment gives the shape:

bash
COMPOSE_PROFILES=rocm MODEL_DIR=~/models MODEL=Qwen3-4B docker compose up -d
docker compose run --rm crane-bench-rocm gdn_bench 16 512 128 128 100
docker compose run --rm crane-bench-rocm topk_bench 248320 40 200

With no profile selected nothing starts, which is a good default. The ROCm service is where the friction is: it maps `/dev/dri` and `/dev/kfd` from the host and refers to host GIDs in a comment that reads render equals 105 and video equals 39 on this host, with the sentence about resolving names instead cut off partway through. So the file encodes one machine's group IDs. The same section warns against setting `HSA_OVERRIDE_GFX_VERSION` because gfx1151 is supported natively, and a separate note records that YAML merge keys merge mappings rather than lists, so a service that sets its own `volumes` or `environment` has to repeat the anchor's entries.

The server binds [::] because rootless podman forwards localhost as IPv6

One line of the compose file carries a long comment. The container command passes `--host [::]`, and the comment explains that the default 0.0.0.0 is IPv4 only, while rootless podman with pasta forwards the host's localhost, which is `::1`, as IPv6, so a container bound to IPv4 only would refuse the forwarded address. That is a specific, hard-won interoperability fix, and it also means the server listens on both families inside the container whether or not a given user needs it. The rest of the shared anchor is conventional: the model directory is mounted read-only, the port is `${CRANE_PORT:-8080}:8080`, and the model path is passed as `/models/${MODEL:?set MODEL to a path under MODEL_DIR}`, so compose refuses to start rather than serving nothing when the variable is unset.

One changelog entry says the real checkpoint was never loaded

The 15 September entry on KugelAudio is the most useful paragraph in the project because of what it concedes. Metal support is described as verified with new Metal-specific regression tests across the convolution tokenizer, the decoder backbone, the diffusion head and the DPM-Solver++ scheduler, confirming there are no CUDA-only paths left in the port. A new `--quant q4_0` in-situ quantization path is described in detail, reading each tensor from a CPU-scoped `VarBuilder` so the transient high-precision weight never lands on the target device. Then the sentence that matters: loading the real roughly 18.7 GB checkpoint end to end has not been verified on a memory-constrained machine, with a pointer to a doc comment in `example/kugelaudio_simple.rs`. The entry also explains that `embed_tokens` and `lm_head` stay dense, about 1 GB each in half precision, because quantizing their 152064-row shape spiked memory disproportionately.

Calendar versioning with no release, and a different author in the manifest

The workspace version is `26.9.0`, which reads as September 2026 and matches the newest update entry dated 2026.09.21, so the project versions by calendar rather than semver. The repository has no GitHub releases at all, which means there is no tag to install and no changelog attached to a downloadable artefact, only a manifest version and a README. The manifest also names an author who does not match the repository owner: `authors` lists Nicholas Jela with a gmail address, in a repository under the `lucasjinreal` account. Four crates make up the workspace, `crane-core`, `crane`, `crane-serve` and `example`, on edition 2024, with `install.sh` and `publish.sh` at the root and a `docker/` directory, so publishing is scripted even though the release history is empty.

Four backends, one of them a single Intel card

Execution paths are named throughout the project: CPU, NVIDIA CUDA, Apple Metal, and an AMD ROCm path for Strix Halo at gfx1151. A fourth appears in the model list, where Qwen3.8-Flash-Next is described as a 177B mixture-of-experts model with hyper-connections, n-gram PLE and QSA sparse attention, running a GSQ-RCO coder GGUF on a single 32 GB Intel Arc Pro B70 via SYCL. A `build_sycl_rpath.rs` at the repository root exists to make that path link. So the Intel path is evidenced by one card and one model, while the CUDA path carries the bulk of the model list. The performance claim to treat carefully is the one in the README's own words, that Qwen3-VL 2B runs 50 times faster than native PyTorch on M1, M2 and M3, since no benchmark for that figure appears in the repository listing.

Editorial conclusion

Crane is worth reading if you want Candle's kernels without writing Candle plumbing, and the parts that are finished are finished properly: ternary GGUF loading with automatic format detection, quantized embedding tables that save 1772 MiB on a 27B Q4_K_M, per-request thinking control, and a changelog with real numbers in it. Two things to weigh first. The feature set is uneven against the project's own description, since VLA appears in the repository description while every VLA model on the list is unchecked and the feature list says VLA is on the way, so read the checkbox column rather than the tagline. And the ROCm path in `compose.yaml` carries host GIDs from one machine with a comment about it, so it is a starting point to edit rather than a file to run unchanged. There are no GitHub releases, the manifest version is a calendar version reading 26.9.0, and the last commit on the default branch main is dated 1 October 2026.

Frequently asked questions

Does Crane support VLA models?

Not yet, according to its own checklist. The repository description mentions VLA, but Qwen3.5-VLA, Qwen3.5-GR00T and Pi0.5 are all unchecked in the supported models list, and the feature list describes VLA as on the way.

Which GPU backends does Crane support?

CPU, NVIDIA CUDA, Apple Metal, and AMD ROCm for Strix Halo at gfx1151, plus a SYCL path evidenced by a 32 GB Intel Arc Pro B70 running a Qwen3.8-Flash-Next GGUF. A build_sycl_rpath.rs at the repository root exists for the SYCL path.

How do I run crane-serve with Docker?

With compose profiles, one service per backend, setting COMPOSE_PROFILES, MODEL_DIR and MODEL, for example COMPOSE_PROFILES=rocm MODEL_DIR=~/models MODEL=Qwen3-4B docker compose up -d. With no profile selected nothing starts, CRANE_PORT overrides the published port which defaults to 8080, and the model directory is mounted read-only.

What is the version of Crane and are there releases?

The workspace manifest is at 26.9.0, a calendar version matching the September 2026 update entries, and the repository has no GitHub releases. The workspace has four crates, crane-core, crane, crane-serve and example, on edition 2024.

Official sources

  1. Issues
  2. License: MIT
  3. lucasjinreal/Crane on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lucasjinreal-crane.svg)](https://hysenlabs.com/projects/lucasjinreal-crane)