# Atomic Chat: a local LLM app and OpenAI-compatible server on port 1337

> Atomic Chat is a Tauri desktop and mobile app that runs open-weight models through three inference engines and exposes them on an OpenAI-compatible endpoint. It is built for people who want agent tooling to point at a local server without giving up the cloud providers they already pay for.

**AtomicBot-ai/Atomic-Chat** — Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V

- Repository: https://github.com/AtomicBot-ai/Atomic-Chat
- Website: https://atomic.chat
- Stars: 1,640 · Forks: 193
- Language: TypeScript
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/atomicbot-ai-atomic-chat

## The problem Atomic Chat solves: agents that need a local endpoint, not just a chat window

Most local chat apps stop at the conversation. You pick a model, type, and read. The interesting part of Atomic Chat is the second half of its own description: an inference engine for agents. The README states that the app runs an OpenAI-compatible server at http://localhost:1337/v1 and calls it a drop-in replacement for the OpenAI SDK. That single claim decides who the project is for.

If you use Claude Code, Codex CLI, Cline, OpenCode, Droid, Goose, OpenHands, Copilot CLI, Kilo Code or Zed, those tools expect a base URL and an API key. Atomic Chat supplies both. The README lists a one-click launch for exactly those agents from an Integrations tab, which is the practical version of the same idea: instead of editing environment variables by hand, the app starts the agent for you with the local endpoint already wired in.

The privacy story follows from the binding, not from a promise. The server binds to 127.0.0.1 by default. Conversations and keys stay on the machine unless you change host to 0.0.0.0. That is a configuration fact you can verify with a port scan, which is a stronger position than a marketing page.

## Three engines behind one port: llama.cpp, the TurboQuant fork and MLX-VLM

Atomic Chat does not implement inference itself. The README names three engines and says all of them are exposed through the same OpenAI-compatible API at http://localhost:1337/v1.

The first is atomic-llama-cpp-turboquant, the project's own llama.cpp fork. Its distinguishing feature is a TurboQuant KV cache with turbo3 and turbo4 modes, described as up to roughly 4.3x smaller KV cache footprint, running on CPU and on GPU through CUDA or Vulkan. The README notes this engine is now a selectable second provider on macOS, Windows and Linux, having previously been macOS-only. KV cache size is what usually decides whether a long-context model fits in VRAM, so this is the engine to try first when you are memory-bound rather than compute-bound.

The second is upstream llama.cpp from ggml-org, and the README says it is the default engine on Windows and Linux. That default is a deliberate trade: upstream tracks the widest hardware coverage and carries MTP support, while the fork carries the quantisation experiments. If a model fails to load on the fork, switching to upstream is the documented escape hatch.

The third is MLX-VLM, which is Apple-silicon specific. The feature list ties EAGLE-3 speculative decoding for Gemma 4, MTP for Qwen 3.5 and 3.6 and DeepSeek V4, and a TurboQuant KV cache with RHT-correct fast paths to MLX. The engine choice is therefore not cosmetic. It is a hardware decision made before you download anything.

## Installing Atomic Chat and pointing a client at localhost:1337

There is no source build required for normal use. The README's Download section links a universal .dmg for macOS, an x64 setup .exe for Windows, and an amd64 AppImage for Linux, plus an iOS App Store listing and an Android Google Play listing. The homepage at atomic.chat is where the README sends you for Getting Started.

If you do build from source, the Makefile shows the shape of it: yarn install, then yarn build:tauri:plugin:api, yarn build:core and yarn build:extensions. The package.json sets the runtime floor at Node.js 20 or later and defines dev as yarn dev:tauri. The workspace name inside package.json is jan-app, which is worth noticing if you go looking for the project under a different name in your node_modules tree.

Once a model is loaded in the app, the server is already listening. This is the README's own curl example, with the model identifier replaced by whatever you loaded:

```bash
curl http://localhost:1337/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-id-loaded-in-atomic-chat>",
    "messages": [{ "role": "user", "content": "Say hello in one word" }]
  }'
```

A successful response is a normal OpenAI chat completion object, and resp.choices[0].message.content holds the text. The Python path is the same endpoint with a different base_url, and the README notes the api_key value is not checked:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:1337/v1", api_key="not-needed")

resp = client.chat.completions.create(
    model="<model-id-loaded-in-atomic-chat>",
    messages=[{"role": "user", "content": "Say hello in one word"}],
)
print(resp.choices[0].message.content)
```

The two things that trip people up are both stated in the README. The model field must match a model loaded in the app, not a name you invented. And the default binding is loopback only, so a container or a second machine on your network cannot reach it until you set host to 0.0.0.0.

## Where Atomic Chat is the wrong tool

The licence is the first obstacle. The repository reports NOASSERTION, and the README does not state licence terms at all. There is a LICENSE file at the top level, so the information exists, but you have to open it. If your organisation requires an OSI-approved licence before a dependency reaches a developer laptop, that check happens before anything else.

Second, the project is a GUI application first. Everything in the README assumes a desktop or mobile app where a model is loaded by hand. The repository layout has a web-app workspace, a core workspace, an mlx-server directory and a foundation-models-server directory, but the README documents no headless mode, no CLI for loading a model, and no server-only deployment. If your target is a Linux box with no display, a bare llama.cpp server is a smaller thing to operate, and Atomic Chat adds a Tauri shell you will never open.

Third, the performance numbers in the feature list are claims about specific model and engine pairings. MTP is described as a 30 to 70 percent throughput boost on supported models and up to 3x on Gemma 4. DFlash block-diffusion decoding is described as up to 6x faster on Qwen 3.6, Gemma 4 and Kimi K2.5. Those figures come with the qualifier "on supported models", and the README does not publish the benchmark setup behind them. Treat them as directions to test, not as numbers to plan capacity around.

Finally, exposing the server on 0.0.0.0 is a one-line change with no documented authentication in front of it. The README presents it as a LAN convenience. On an untrusted network that is an open inference endpoint.

## Atomic Chat versus Ollama and LM Studio

Ollama and LM Studio occupy the same slot, and the differences are structural rather than cosmetic.

Ollama is a daemon with a CLI and its own model registry, and it is the usual pick when the endpoint matters more than the interface. Atomic Chat inverts that: the interface is the product, and the endpoint is a feature of it. It also gives you three engines rather than one, including the TurboQuant KV cache fork and MLX-VLM on Apple silicon, and it ships built-in cloud providers (OpenAI, Anthropic, Mistral, Groq, MiniMax, Qwen, Moonshot) so a single chat can mix a local model with a hosted one. Ollama's model management is more scriptable; Atomic Chat's is more visual.

LM Studio is the closest match, since it is also a desktop app with a local server and a model browser. The distinction the README draws is the agent integration layer: a one-click launch list covering Claude Code, Codex CLI, Cline, OpenCode, Droid, Goose, OpenHands, Copilot CLI, Kilo Code and Zed, plus MCP server connections and an Artifacts preview panel for HTML, CSS and JavaScript. If you only want a chat box with a local model, that layer is weight you carry for nothing.

## Release cadence, upgrade cost and the licence question

The last push to main was on 2026-09-09, and the most recent release, v2.0.35, was tagged the same day. Before that came v2.0.32 on 2026-09-02 and v2.0.23 on 2026-08-21. Three releases in roughly three weeks is a fast cadence, and fast cadences have a cost: the README documents no migration notes, no database schema versioning and no rollback procedure for moving between 2.0.x builds. Conversations, custom assistants, projects and MCP server configuration all live in the app, and nothing in the README says what happens to them when you install the next build over the top.

The practical consequence is that you should decide your upgrade policy before you accumulate state in the app, not after. If your chats matter, export them or keep the installer for the version you are running. If they do not, upgrade freely.

On licensing, the repository reports NOASSERTION. That is not a licence, it is the absence of a machine-readable one. There is a LICENSE file in the repository root, and the package.json copy:assets:tauri script copies LICENSE into src-tauri/resources, so the file ships with the build. Read it before you redistribute anything, and if you need a legal opinion, get one from a lawyer rather than from a README.

## Conclusion

Adopt Atomic Chat if you want a desktop app that doubles as a loopback OpenAI endpoint for local agents, and you are willing to pin a release and check the licence file yourself, because the repository declares NOASSERTION and there is no stated upgrade path between v2.0.23, v2.0.32 and v2.0.35. Skip it if you need a headless Linux server with no GUI, or if you require an OSI-approved licence before you ship. Before installing, open the LICENSE file and confirm which of the three engines your hardware can actually run, since the default engine differs between macOS and Windows or Linux.

## FAQ

### What is Atomic Chat?

It is a local AI app and inference engine for agents, built with Tauri, that runs open-weight LLMs on your own machine and exposes them through an OpenAI-compatible server at http://localhost:1337/v1. It ships desktop builds for macOS, Windows and Linux plus iOS and Android apps.

### Is Atomic Chat safe?

The README states that the local server is bound to 127.0.0.1 by default and that conversations and keys stay on your machine, so nothing is reachable from the network until you set host to 0.0.0.0. The README does not document authentication on that endpoint, so exposing it on a LAN removes the loopback protection without adding anything in its place.

### Is Atomic Chat legit?

It is a public repository under AtomicBot-ai with tagged releases, the most recent being v2.0.35 on 2026-09-09, and the last push to main was on 2026-09-09. The repository reports NOASSERTION for its licence, so the LICENSE file in the repository root is the document to read if licensing matters to you.

### What is Atomic AI?

In this context the name refers to Atomic Chat, the local AI app and inference engine from AtomicBot-ai that runs open-weight LLMs on your machine and serves them at http://localhost:1337/v1. The README describes it as private and offline-capable, with cloud providers available as an optional addition.

## Sources

- [AtomicBot-ai/Atomic-Chat on GitHub](https://github.com/AtomicBot-ai/Atomic-Chat)
- [Issues](https://github.com/AtomicBot-ai/Atomic-Chat/issues)
- [Project website](https://atomic.chat)
- [README](https://github.com/AtomicBot-ai/Atomic-Chat/blob/main/README.md)
- [Releases](https://github.com/AtomicBot-ai/Atomic-Chat/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/atomicbot-ai-atomic-chat
