Model or dataset
kaito-project/aikit avatar
kaito-project/aikit

AIKit: build, fine-tune and ship LLMs as container images

🏗️ Fine-tune, build, and deploy open-source LLMs easily!

539 stars57 forksGoMIT

At a glance

What is it?
AIKit turns a declarative YAML file into an OCI image that serves an OpenAI-compatible API, using BuildKit under the hood. It fits teams who want model distribution to look like image distribution, and it will frustrate anyone expecting a hosted service.
Who is it for?
AIKit suits platform teams who already ship container images and want model serving to follow the same registry, signing and SBOM path, plus anyone who needs an OpenAI-compatible endpoint on a laptop without a GPU. It is the wrong tool if you want a managed inference service, or if your deployment is air-gapped and you rely on the pre-made runner images, which the README says download models at startup and are not air-gapped by default.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap AIKit fills between a model file and a running endpoint

Downloading a GGUF file is easy. Getting that file to a colleague, a CI job, or a Kubernetes node in a way that is reproducible, signed and versioned is the part that usually turns into shell scripts. AIKit's answer is to treat the model as a container image. The README describes three capabilities: inference through LocalAI, fine-tuning through Unsloth, and packaging models as OCI artifacts that can be pushed to any OCI-compliant registry. The intended user is not a researcher tuning hyperparameters. It is an engineer who already has a registry, a signing policy and a deployment pipeline, and wants a language model to flow through that pipeline like any other artifact.

The pitch rests on a narrow dependency: Docker or Podman. The README states that no GPU, internet access or additional tools are needed beyond one of those two runtimes, which is a real constraint in the other direction too. If your environment cannot run containers, AIKit has nothing to offer you.

BuildKit frontend, Aikitfile, and what actually lands in the registry

The mechanism is visible in the repository layout. The Dockerfile builds ./cmd/frontend into a static binary and ships it in a scratch image with an ENTRYPOINT of /bin/aikit. That binary is a BuildKit frontend, and go.mod pins github.com/moby/buildkit v0.33.0 alongside github.com/modelpack/model-spec v0.0.7. So the flow is: you write a declarative Aikitfile, docker buildx build reads it through the AIKit frontend, and the output is an image or an OCI artifact rather than a running process. The Makefile shows this in practice, with TEST_FILE defaulting to test/aikitfile-llama.yaml and build-test-model invoking docker buildx build -f ${TEST_FILE}.

The frontend approach has consequences worth naming. Because the build runs inside BuildKit, caching, multi-platform output and provenance attestations come from BuildKit rather than from AIKit itself. The README lists SBOMs, provenance attestations and signed images under supply chain security, and the Makefile's build-base target passes --sbom=true --push. That is the strongest argument for the design: you inherit an ecosystem you probably already operate. The cost is that debugging a failed build means reading BuildKit output, and the Aikitfile is a build input, not a runtime config you can edit on a live deployment.

Running a pre-made model in two commands

The README's quick start does not require building anything. It runs a published image that already contains the weights and a LocalAI server. This is the fastest way to see whether the tool fits, and it works on AMD64 and ARM64 CPUs because the images are multi-platform.

Start the Llama 3.1 8B image in the background and publish port 8080:

bash
docker run -d --rm -p 8080:8080 ghcr.io/kaito-project/aikit/llama3.1:8b

Navigate to http://localhost:8080/chat for the WebUI, as the README instructs. The same port serves the OpenAI-compatible API, so any OpenAI client can point at it. The README gives this curl example, and the response echoes the model name llama-3.1-8b-instruct with a normal chat completion body:

bash
curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
    "model": "llama-3.1-8b-instruct",
    "messages": [{"role": "user", "content": "explain kubernetes in a sentence"}]
  }'

For GPU hosts the Makefile shows the shape of the change rather than the README: run-test-model-gpu adds --gpus all, and run-test-model-rocm passes --device /dev/kfd --device /dev/dri with video and render group membership. Those flags are for locally built test images, so treat them as a reference for the device plumbing, not as a documented interface for the published tags.

Where AIKit stops being the right tool

The air-gapped story is the clearest limitation, and the README states it plainly: air-gapped inference works with self-hosted or local registries when model content and dependencies are baked or mirrored ahead of time, and runner images that download models at startup are not air-gapped by default. That sentence splits the pre-made catalogue into two classes, and the README does not label which tag is which. If your cluster has no egress, you must confirm per image whether the weights are inside the layers or fetched on first boot. Getting this wrong produces a container that starts, logs a download attempt, and fails in a way that looks like a network fault rather than a packaging choice.

There are two other boundaries. Fine-tuning is delegated to Unsloth, so AIKit's contribution is the interface and the packaging, not the training loop; if you need distributed training across many nodes, this is not that. And the whole thing assumes a container runtime. A bare-metal Python environment where you cannot install Docker or Podman is out of scope, no matter how small the model.

AIKit versus a plain Ollama or llama.cpp setup

The obvious alternative is running llama.cpp or Ollama directly. Both give you a local server and, in Ollama's case, a model pull command. The difference is what you are distributing. Ollama manages a local model store on the host; the model is state that lives outside your deployment tooling. AIKit produces an artifact in a registry, so the model version is an image tag you can pin in a Kubernetes manifest, sign, and attach an SBOM to. The Makefile reflects that orientation: the targets are build-aikit, build-test-model, build-base, and the KIND_VERSION and KUBERNETES_VERSION variables sit next to them because the intended destination is a cluster.

That is a genuine trade-off rather than a clear win. Pulling a multi-gigabyte image on every node is heavier than a shared model cache, and rebuilding for every model change is slower than editing a config file. AIKit is the better fit when model provenance and deployment reproducibility matter more than iteration speed. If you just want a chat model on your laptop, the two-command quick start is convenient, but Ollama would be equally convenient and less opinionated about registries.

Maintenance, licensing and the version you pin

The repository is not archived, and the last push was on 2026-09-08. The release history shows v0.22.1 on 2026-08-08, v0.21.0 on 2026-03-17 and v0.20.3 on 2025-12-04, so releases arrive in bursts rather than on a schedule. One inconsistency is worth flagging before you file a bug: the Makefile sets VERSION := v0.21.0 while the newest release is v0.22.1, so the version string baked into locally built binaries may lag the published tag.

AIKit itself is MIT licensed, which is permissive and imposes little on your own code. The models are a different matter, and the README's pre-made table lists the licence per model rather than implying one: the Llama entries point at the Meta Llama licence, Mixtral at Apache 2.0, Phi 4 at MIT, and Gemma has its own terms. Packaging weights into an image does not change the terms attached to those weights. Check the licence of the specific model you bake before you push it to a shared registry, and treat the per-model links in the README as the starting point rather than the final word.

Editorial conclusion

AIKit suits platform teams who already ship container images and want model serving to follow the same registry, signing and SBOM path, plus anyone who needs an OpenAI-compatible endpoint on a laptop without a GPU. It is the wrong tool if you want a managed inference service, or if your deployment is air-gapped and you rely on the pre-made runner images, which the README says download models at startup and are not air-gapped by default. Before adopting it, verify three things against your own registry: whether the pre-made image you picked is baked or a runner, which model licence applies to the weights inside it, and whether your Kubernetes cluster can pull the artifact format you produced.

Frequently asked questions

What is AIKit?

AIKit is a platform for hosting, deploying, building and fine-tuning large language models, with three stated capabilities: inference through LocalAI, fine-tuning through Unsloth, and packaging models as OCI artifacts. It is written in Go and exposes a BuildKit frontend so models can be built and distributed as container images.

Do I need a GPU to run AIKit?

No. The README states that no GPU, internet access or additional tools are needed beyond Docker or Podman, and the quick start runs the Llama 3.1 8B image on a local machine. GPU acceleration is optional and the documentation covers NVIDIA CUDA and AMD ROCm support.

How do I install AIKit?

There is no package to install for the pre-made models. The README's quick start is a single docker run command against ghcr.io/kaito-project/aikit/llama3.1:8b, which serves a WebUI and an OpenAI-compatible API on port 8080. Building your own images uses docker buildx build with an Aikitfile, as the Makefile's build-test-model target shows.

Can AIKit run without internet access?

The README says air-gapped inference works with self-hosted or local registries when model content and dependencies are baked or mirrored ahead of time, and warns that runner images which download models at startup are not air-gapped by default. You have to check each image to know which category it falls into.

Is AIKit compatible with OpenAI API clients?

Yes. The README describes LocalAI as providing a drop-in replacement REST API that is OpenAI API compatible, and lists clients such as Kubectl AI and Chatbot-UI as examples. The curl example posts to /v1/chat/completions and returns a standard chat completion response.

Official sources

  1. kaito-project/aikit on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kaito-project-aikit.svg)](https://hysenlabs.com/projects/kaito-project-aikit)