CLI tool
ardanlabs/kronk avatar
ardanlabs/kronk

ardanlabs/kronk: a Go SDK and model server for local inference over llama.cpp and whisper.cpp

Your personal engine for running open source models locally. Use Go for hardware accelerated local inference with llama.cpp and whisper.cpp directly integrated into your Go applications. Kronk provides a high-level API and a model server.

821 stars62 forksGoApache-2.0

At a glance

What is it?
Kronk wraps llama.cpp, whisper.cpp and stable-diffusion.cpp behind Go APIs and an OpenAI-compatible model server, so Go teams can run text, vision, embedding, reranking and transcription workloads without Python in the stack. The trade-off is a hard coupling to native library versions that the project tracks release by release.
Who is it for?
Adopt Kronk if your application is already Go and you want inference inside the process, or if you want an OpenAI-compatible endpoint you host yourself on Linux, macOS arm64 or Windows amd64. Do not adopt it if you need a stable API surface for image generation, since Malina is documented as experimental and not a model server backend, or if you cannot accept that a llama.cpp bump can force a coordinated upgrade across yzma, bucky, malina and Kronk.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Kronk solves for Go teams running local models

Most local inference stacks assume Python. You install a runtime, pull wheels for CUDA or Metal, load a model through a Python binding, and then reach it from Go over HTTP or a subprocess. Kronk removes that layer. The README describes it as "a Go SDK and model server for hardware-accelerated local inference" that provides "high-level Go APIs over native inference engines without requiring Python or a separate model-serving stack".

The audience is narrow and identifiable: Go developers who want inference inside their own binary, and teams that want to host an OpenAI-compatible endpoint without operating a Python service next to it. The README splits these into two columns. Use the SDK when you need inference inside a Go process, direct control over model loading and lifetime, or application-specific caching and concurrency. Use the model server when you need HTTP APIs for one or more clients, OpenAI- and Anthropic-compatible endpoints, browser-based model management, or authentication, rate limiting, metrics and tracing.

That split matters because the server is built on the same public SDKs. If you outgrow the server, the API underneath is the same one you would have called directly. The go.mod file shows the shape of the dependency graph: yzma for the llama.cpp binding, bucky for whisper.cpp, malina for stable-diffusion.cpp, plus OPA, Prometheus, OpenTelemetry, MCP and Badger for the server side.

How Kronk binds llama.cpp, whisper.cpp and stable-diffusion.cpp into one SDK

Kronk is three subsystems with different maturity levels. Kronk itself runs text, vision, embedding and reranking models through llama.cpp and yzma. Bucky provides speech-to-text through whisper.cpp and bucky. Malina is an experimental image-generation SDK built on stable-diffusion.cpp and malina.

The README states that Kronk downloads native libraries compatible with the installed release, and that the first model or SDK example you run can download the compatible native libraries and model files automatically. So the data flow is: your Go code calls the SDK, the SDK calls into a native library through FFI (the jupiterrider/ffi dependency in go.mod), and the native library loads a GGUF or equivalent model file from local storage. The model server wraps the same SDKs and exposes Chat Completions, Responses, embeddings, reranking and audio transcription in OpenAI-compatible form, plus an Anthropic-compatible Messages API.

The SDK feature list is concrete: text generation, streaming, reasoning, tool calls, vision, embeddings, reranking, concurrent processing and incremental message caching. Bucky supports file transcription, translation, channel-separated diarization and live streaming transcription. The examples directory mirrors this, with separate entries for chat, response, embedding, rerank, rag, grammar, concurrency, pool and agent, and a cluster of malina examples covering img2img, controlnet, animatediff, upscale and adetailer.

Installing Kronk and running your first local model

The README recommends Homebrew on macOS and Linux. The fully qualified formula name is deliberate: it trusts only the Kronk formula rather than every item in the tap.

bash
brew install ardanlabs/kronk/kronk

kronk server start

After the server starts, open http://localhost:11435 to manage models and use the browser interface. That port is the default the README gives, not a value you configure away without checking the manual.

If you prefer the Go toolchain, the CLI installs the same way as any Go command on a supported platform.

bash
go install github.com/ardanlabs/kronk/cmd/kronk@latest

kronk server start

The first model or SDK example you run can download the compatible native libraries and model files automatically, so the initial start may take noticeably longer than later ones. For container deployment and persistent storage, the README points to the Container Quick Start in the installation chapter of the manual.

If you want the SDK rather than the server, the repository ships runnable examples driven by make targets. These download compatible libraries and models on their first run.

bash
make example-question  # Ask a local language model a question.
make example-agent     # Run a small coding agent.
make example-vision    # Prompt a vision model with an image.

Expect the first invocation of any of these to fetch both a library bundle and a model file. Nothing in the README suggests an offline-first install path for the examples.

The native library coupling is Kronk's real constraint

Kronk does not ship its own inference kernels. It follows the native engines it integrates, and the README is explicit that upstream changes can require coordinated releases of yzma, bucky, malina and Kronk. The compatibility tables are the practical consequence.

llama.cpp v0.3.0-b10646 maps to yzma v1.25.0 and Kronk 1.32.3 or later. whisper.cpp v1.9.3 maps to bucky v1.1.0 and Kronk 1.31.8 or later. stable-diffusion.cpp master-830-50d6405 maps to malina v1.0.5 and Kronk v1.32.2 or later. Older rows exist for earlier Kronk releases, which is useful when you are pinned to a version and need to know which llama.cpp build it expects.

The README's instruction is unambiguous: use each subsystem's downloader instead of mixing native libraries from unrelated releases, because every Kronk release is bound to known-compatible library versions. If you vendor your own llama.cpp build and point Kronk at it, you are outside the supported combination. That is a legitimate thing to do, but you own the debugging.

This is the failure mode worth naming before adoption. A llama.cpp update that changes an ABI or a symbol can force you to move yzma, bucky, malina and Kronk together. The repository carries a BREAKING_CHANGES.md file for exactly this reason. Teams that treat the native layer as something to pin independently of the Go module will spend time reconciling versions the project already reconciled for them.

Platform support is also uneven by design. Linux on amd64 and arm64 gets CUDA, Vulkan, HIP, ROCm and SYCL. macOS is arm64 only, with Metal. Windows is amd64 only, with CUDA, Vulkan, HIP, SYCL and OpenCL. The README warns that not every backend is available for every SDK or architecture and points to the CLI or SDK library manager as the source of truth for your installed version. If you are on macOS x86_64, this project is not for you.

Where Kronk sits against llama.cpp bindings and Ollama

The obvious alternative is calling llama.cpp directly, either through its own server binary or through a thin binding. That gives you the smallest possible dependency surface and no opinion about model management. You also write the model download, the OpenAI-compatible request shapes, the auth layer, the metrics and the tracing yourself. Kronk's bet is that most Go teams do not want to write those, and the go.mod file shows what it pulled in to avoid it: OPA for policy, Prometheus for metrics, OpenTelemetry for tracing, Badger for storage, MCP for tool integration.

A second comparison is Ollama. Both expose an OpenAI-compatible endpoint and both manage local model files. The difference the README makes visible is the embedding model. Kronk is a Go module you import, with the server built on the same public SDKs, so an application can call inference in-process and skip HTTP entirely. Ollama is a separate daemon you talk to over its API. If your application is Go and you want the model load lifetime under your own control, that distinction is the whole decision. If your application is not Go, or you want one daemon serving several unrelated clients, Kronk's SDK advantage does not apply and its server is just another OpenAI-compatible endpoint.

The README also lists integrations with OpenWebUI, OpenCode and Claude Code, which suggests the server is meant to slot into existing client tooling rather than replace it.

Licence, release cadence and what upgrading costs

Kronk is Apache-2.0. That is a permissive licence with an explicit patent grant, and it does not impose copyleft obligations on your application. It says nothing about the licences of the models you load or of the native libraries Kronk downloads, and those are separate questions you have to answer for your own distribution. This is not legal advice.

The release history in the repository shows three releases within roughly two days at the end of August 2026: v1.32.1, v1.32.2 and v1.32.3. That cadence is consistent with a project that tracks fast-moving upstream engines rather than one that batches changes quarterly. The last push to the default branch was on 2026-08-27.

The upgrade cost is therefore not the Go module bump. It is the native library bundle that comes with it. If you pin Kronk, you also pin yzma, bucky and malina versions, and you inherit whichever llama.cpp, whisper.cpp or stable-diffusion.cpp build those were tested against. Upgrading Kronk means re-downloading those bundles, which the README frames as the supported path. Treat the compatibility table as part of your dependency manifest, not as documentation you read once.

Editorial conclusion

Adopt Kronk if your application is already Go and you want inference inside the process, or if you want an OpenAI-compatible endpoint you host yourself on Linux, macOS arm64 or Windows amd64. Do not adopt it if you need a stable API surface for image generation, since Malina is documented as experimental and not a model server backend, or if you cannot accept that a llama.cpp bump can force a coordinated upgrade across yzma, bucky, malina and Kronk. Verify first that your GPU backend appears in the platform table for your architecture, and check the compatibility table against the exact Kronk version you install before pinning anything in production.

Frequently asked questions

What is ardanlabs/kronk?

It is a Go SDK and model server for hardware-accelerated local inference, providing high-level Go APIs over llama.cpp and whisper.cpp without requiring Python or a separate model-serving stack. The model server exposes OpenAI-compatible Chat Completions, Responses, embeddings, reranking and audio transcription endpoints, plus an Anthropic-compatible Messages API.

How do I install ardanlabs/kronk?

The README recommends Homebrew on macOS and Linux with the fully qualified formula name ardanlabs/kronk/kronk, or go install github.com/ardanlabs/kronk/cmd/kronk@latest on a supported platform. Both paths are followed by kronk server start, after which the browser interface is available at http://localhost:11435.

Which platforms and GPU backends does Kronk support?

Linux on amd64 and arm64 supports CUDA, Vulkan, HIP, ROCm and SYCL; macOS arm64 supports Metal; Windows amd64 supports CUDA, Vulkan, HIP, SYCL and OpenCL. The README states that not every backend is available for every SDK or architecture and points to the CLI or SDK library manager as the source of truth for your installed version.

Is the Malina image generation part of Kronk stable?

No. The README marks Malina as experimental, states that its public API is subject to change, and notes that it is not yet a Kronk model-server backend. It is currently SDK-only.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ardanlabs-kronk.svg)](https://hysenlabs.com/projects/ardanlabs-kronk)