Model or dataset
signerlabs/Klee avatar
signerlabs/Klee

signerlabs/Klee: a native macOS AI chat app with local MLX inference

A native macOS AI chat app powered by MLX. 100% local inference on Apple Silicon, no cloud required. Built with ShipSwift.

1,766 stars136 forksSwiftLicense varies

At a glance

What is it?
Klee is a SwiftUI app that runs 4-bit MLX models on Apple Silicon, with no account and no cloud. It is a good fit for Mac owners who want tool-calling chat that stays on the machine, and a poor fit for anyone without 16 GB of RAM or a need for Linux and Windows.
Who is it for?
Adopt Klee if you have an Apple Silicon Mac with at least 16 GB of RAM and want a local chat app whose tool calls (file_read, file_write, shell_exec) stay on the machine. Do not adopt it if you need Windows or Linux, or if you want a large model on a 16 GB machine, where the README's table caps you at the 8B to 12B class.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Swift, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What signerlabs/Klee actually is and who it is for

Klee is a macOS application, not a library or a server. The README describes it as "A native macOS AI agent that runs entirely on your Mac. No cloud, no account, no subscription." Inference runs through MLX, Apple's array framework for Apple Silicon, via the mlx-swift-lm package. The repository is a single Xcode project (Klee.xcodeproj) with a Klee/ source directory, so the deliverable is an .app bundle you drag into Applications.

The target user is narrow and specific: someone with an Apple Silicon Mac, macOS 15.0 or later, and 16 GB of RAM at minimum. If you meet that, Klee gives you a chat window where the model can read files, write files, list directories, delete files, fetch web pages and run shell commands, all without an API key. If you do not meet it, there is no fallback path in the README. There is no CPU mode, no remote backend, no Linux build.

How MLX inference and tool calling work inside Klee

The mechanism is a local model file plus a Swift tool-call loop. Klee downloads 4-bit quantized weights from the mlx-community organization on HuggingFace, caches them in ~/.klee/models/, and loads them with mlx-swift-lm for inference on the GPU. The README states that models "persist across app restarts," and that interrupted downloads "resume automatically." Nothing in that path touches a remote inference endpoint.

The tools are declared natively rather than through MCP. The README is explicit: "Klee uses native tool calling (mlx-swift-lm ToolCall API) -- no MCP, no Node.js, no external processes." The seven tools are file_write, file_read, file_list, file_delete, web_search, web_fetch and shell_exec. Six of them operate on the local filesystem or a local shell; only web_search and web_fetch leave the machine, and web_search requires a Jina AI API key that you paste into the sidebar.

That design has a consequence worth stating plainly. A model with file_delete and shell_exec in its toolset is not a chatbot. The README does not document a sandbox, a confirmation prompt per tool call, or a permission model for the filesystem. The only stated bound on shell_exec is a "30s timeout." Treat the tool surface as the real security boundary, because the documentation does not describe another one.

Installing Klee from the signed .dmg and running a first chat

Klee is not on the App Store. The README says it is "distributed directly as a signed macOS app (Developer ID)." The install is three steps: download the .dmg from the Releases page, drag Klee into Applications, and open it. If Gatekeeper blocks the first launch, the README's instruction is to go to System Settings > Privacy & Security and click "Open Anyway."

There is no installer command to run. The only shell commands in the README belong to the build-from-source path, which needs Xcode 16+ and macOS 15.0+:

bash
 git clone https://github.com/signerlabs/Klee.git
 cd Klee
 open Klee.xcodeproj

After opening the project, select the Klee scheme and build and run with Cmd+R. The README notes the SPM dependency (mlx-swift-lm) resolves automatically on the first build.

For the normal app path, the first real use is model selection. On launch, Klee detects system RAM and shows compatible models. Tap the download button next to one. On a 16 GB machine the README's table recommends Qwen 3.5 9B (~6 GB), Qwen 3 8B (~4.3 GB), Gemma 3 12B (~8 GB) or DeepSeek R1 8B (~4.6 GB). Once the download finishes, select the model and type. You should see tokens stream in as they are generated, and for reasoning models an inline thinking card you can collapse.

Web search is optional and off by default. Get a free key at jina.ai, click the sidebar toggle at the top-right, enable Web Search, and paste the key. Until you do, web_search and web_fetch have nothing to call.

The 16 GB floor and the model table are the real constraints

The hard limit is memory. Klee requires 16 GB of RAM, and the README's table ties model choice to that number. On 16 GB you are choosing between 4.3 GB and 8 GB of weights. On 32 GB you can reach Qwen 3.5 27B (~16 GB) or Qwen 3.5 35B MoE (~20 GB). The 70 GB Qwen 3.5 122B model needs 96 GB or more. There is no documented offloading strategy for running a larger model on a smaller machine, so the table is a ceiling, not a suggestion.

This is the case where Klee is the wrong tool. If your Mac has 8 GB of RAM, Klee is not for you, and the README offers no smaller tier. If you are on an Intel Mac, it is not for you either, since the requirement is Apple Silicon M1 or later. And if your workload depends on the largest open models, a 16 GB laptop will not get you there through Klee.

The second limitation is scope. Klee is macOS-only, closed to plugin authors for now ("Platform modules" are listed as coming soon), and its model catalogue is a fixed set of mlx-community 4-bit builds. There is no documented way to point Klee at an arbitrary local GGUF file or at an Ollama instance. If your models live outside that catalogue, Klee will not see them.

Maintenance is worth a note. The repository is not archived, and the last push was on 2026-03-20. The single release, v1.0.0, is dated 2026-03-19. That is roughly six months before today, so the project is best described as having a stable v1.0.0 rather than a fast-moving branch.

How Klee differs from Ollama and LM Studio

The closest alternatives are Ollama and LM Studio, and the difference is in the packaging, not the inference. Ollama runs as a background service with a CLI and an HTTP API on port 11434, and it is cross-platform: macOS, Linux and Windows. Klee is a single 75 MB SwiftUI app with no background service and no documented API, and it runs only on macOS 15+ with Apple Silicon.

LM Studio offers a GUI for downloading and running local models on macOS, Windows and Linux, and it exposes an OpenAI-compatible server. Klee does not document a server mode, so you cannot point another program at it. What Klee adds instead is a built-in agent loop: the model can call file_read, file_write, file_list, file_delete, web_search, web_fetch and shell_exec directly, with no MCP server and no Node.js process to manage. Ollama and LM Studio give you the model; Klee gives you the model plus a tool harness wired into the chat window.

That trade runs both ways. Choosing Klee means accepting a fixed model catalogue, a macOS-only binary, and a tool surface that the README does not describe as sandboxed. Choosing Ollama or LM Studio means assembling the agent layer yourself, or using a separate client that speaks their API.

Build, upgrade and licence questions

Upgrades follow the same path as install: download a newer .dmg from Releases and replace the app in Applications. The README does not document an in-app updater or a rollback procedure, so if a new build misbehaves, the recovery step is to reinstall an older .dmg from the Releases page. Models live in ~/.klee/models/ and persist across app restarts, so replacing the app does not force a re-download.

Building from source has a higher floor than running the app: Xcode 16+ and macOS 15.0+. Dependencies resolve through SPM on the first build, which means the build depends on network access to the package registry and on the mlx-swift-lm package resolving cleanly.

The README states the licence as MIT. That is permissive for the application code, but it is worth separating two things: the Klee source is MIT, while the model weights you download from mlx-community carry their own licences on HuggingFace. The README does not reproduce those terms, so check the model card before using outputs commercially. This is a description of what the repository says, not legal advice.

Editorial conclusion

Adopt Klee if you have an Apple Silicon Mac with at least 16 GB of RAM and want a local chat app whose tool calls (file_read, file_write, shell_exec) stay on the machine. Do not adopt it if you need Windows or Linux, or if you want a large model on a 16 GB machine, where the README's table caps you at the 8B to 12B class. Before trusting it with real files, verify that the Gatekeeper prompt resolves after the Privacy & Security step and that a downloaded model appears under ~/.klee/models/.

Frequently asked questions

What are the system requirements for signerlabs/Klee?

macOS 15.0 (Sequoia) or later, an Apple Silicon chip (M1 or later), and at least 16 GB of RAM. More RAM raises the ceiling on which models the app will recommend.

Does signerlabs/Klee need an API key or an account?

No account or API key is required for chat, and inference runs locally through MLX. The one exception is web search, which needs a free Jina AI API key that you paste into the sidebar toggle.

Where does signerlabs/Klee store downloaded models?

Models are cached in ~/.klee/models/ and persist across app restarts. The README also states that interrupted downloads resume automatically.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. signerlabs/Klee on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/signerlabs-klee.svg)](https://hysenlabs.com/projects/signerlabs-klee)