# Lemonade: Local AI Server with OpenAI, Anthropic, and Ollama API Compatibility

> Lemonade is an open-source local AI server that runs LLMs, image generation, and speech models on a PC's own GPU or NPU, then exposes them through standard OpenAI, Anthropic, and Ollama-compatible APIs. It targets developers and power users on Windows, Linux, and macOS who want to run AI workloads without sending data to a cloud provider.

**lemonade-sdk/lemonade** — Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

- Repository: https://github.com/lemonade-sdk/lemonade
- Website: https://lemonade-server.ai/
- Stars: 5,790 · Forks: 511
- Language: C++
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/lemonade-sdk-lemonade

## What Lemonade Provides and Who It Is For

Lemonade addresses a specific gap: the OpenAI, Anthropic, and Ollama APIs are widely supported by developer tools, but using those APIs sends data to external servers and incurs per-token cost. Lemonade installs a local service that speaks the same API protocols, so any application configured to point at a cloud endpoint can be redirected to the local Lemonade server without code changes.

The README describes two deployment modes. Lemonade Server installs as a background service connected to by any compatible application. Embeddable Lemonade is a portable binary that developers can bundle inside their own applications to give those applications local AI capability that auto-optimizes for the end user's hardware. The project is built by a community with hardware optimizations contributed by AMD engineers, and the README notes explicit optimization for Ryzen AI, Radeon, and Strix Halo PCs.

## CLI Commands: Running Models and Managing the Library

Lemonade ships a command-line interface for all model operations. To run and chat with a model immediately:

```
lemonade run Gemma-4-E2B-it-GGUF
```

To see which models are available and download one:

```
lemonade list

lemonade pull Gemma-4-E2B-it-GGUF
```

Lemonade supports multi-modal operations from the same CLI. Image generation, speech synthesis, and transcription each use the same run command pattern:

```
# image gen
lemonade run SDXL-Turbo

# speech gen
lemonade run kokoro-v1

# transcription
lemonade run Whisper-Large-v3-Turbo
```

To inspect which inference backends are compiled and available on the current machine:

```
lemonade backends
```

Alias management lets developers assign environment-independent names to models, and supports instant failover between models:

```
lemonade alias add production-llm Gemma-4-E2B-it-GGUF
lemonade alias list

# Instant active-standby failover to a different model target
lemonade alias add production-llm Qwen3-0.6B-GGUF
lemonade alias remove production-llm
```

This pattern means application configuration can reference an alias like production-llm rather than a specific model name, and swapping to a different model requires only one CLI command.

## Supported Inference Backends and Hardware

Lemonade's README includes a backend matrix that maps modality, engine, backend, device, and operating system. For text generation, the primary engine is llamacpp, which supports multiple backends: system (x86_64 and ARM64 CPU and GPU on Linux), metal (Apple Silicon GPU on macOS), cuda (NVIDIA GPUs, Turing or newer, on Windows and Linux), and vulkan (GPU-accelerated path on supported platforms). The matrix lists additional engines beyond llamacpp, but the README does not enumerate all supported engines in the install section.

For image generation, speech generation, and transcription, separate engines and backends apply. The README specifically names SDXL-Turbo for image generation, kokoro-v1 for speech, and Whisper-Large-v3-Turbo for transcription as models accessible through the same CLI interface.

Custom GGUF and ONNX models from Hugging Face or ModelScope can be pulled directly. The model library at lemonade-server.ai/models.html lists the included catalog, and the CLI's lemonade pull command fetches models from those sources while retaining the origin URL for future updates.

For hybrid setups, Lemonade can route requests to OpenAI-compatible cloud providers alongside local models. The README describes this as an experimental cloud offload feature.

## Platform Support and Installation Paths

The README lists official installation guides for eight platform configurations: Arch Linux (available in the official Arch repository), Debian, Docker, Fedora, macOS, Snap, Ubuntu, and Windows. Each platform has a corresponding build status indicator in the README.

The Docker image is built from a multi-stage Dockerfile in the repository. The build stage compiles the C++ binaries using CMake with Ninja, and the runtime stage produces a smaller image based on Ubuntu 24.04. The container runs as an unprivileged user (uid 10001) rather than as root, and the application directory doubles as the user's HOME directory so cache and configuration paths resolve without elevated permissions.

Mobile clients are available as separate projects for iOS (on the App Store) and Android (on Google Play), with the source in the lemonade-sdk/lemonade-mobile repository. The mobile apps connect to a running Lemonade Server instance rather than running inference on the device.

The project integrates with a named set of applications including Claude Code, Open WebUI, AnythingLLM, GitHub Copilot, Dify, n8n, OpenHands, and others listed in the marketplace at lemonade-server.ai/marketplace.

## Limitations of Running AI Locally

The core constraint of Lemonade is hardware. Inference speed depends entirely on the GPU or NPU available in the machine. A GGUF model that runs at 50 tokens per second on a high-end discrete GPU may run at 5 tokens per second on an integrated GPU. The README does not document minimum VRAM requirements per model; users must consult the model catalog and match model size to available hardware.

Because the server runs on a single machine, it cannot distribute load across multiple GPUs in different machines. A cloud API can scale horizontally; a local Lemonade instance cannot. For use cases that involve many concurrent users or large batch workloads, local inference imposes a hard throughput ceiling set by the hardware.

The cloud offload feature, which routes some requests to cloud providers, is described in the README as experimental. Organizations that rely on it for overflow capacity should treat it as a development preview rather than a production-ready feature.

Finally, while the API is compatible with OpenAI, Anthropic, and Ollama interfaces, the specific capabilities of each endpoint depend on what the locally running model supports. A model that does not support function calling will return errors for tool-use requests even if the application sends them in the correct OpenAI format.

## Comparison with Ollama

Ollama is the most widely deployed alternative for running local LLMs with an OpenAI-compatible API. The primary difference in approach is scope: Ollama focuses on text-generation models with broad format support, while Lemonade covers text generation, image generation, speech synthesis, and transcription under one server and CLI. Lemonade also exposes an Anthropic-compatible API endpoint in addition to OpenAI and Ollama formats, which means tools configured for Claude's API can connect without a translation layer.

Ollama has broader model format coverage through its Modelfile system. Lemonade specifically supports GGUF, FLM, and ONNX. Applications already invested in Ollama's CLI workflow can migrate by pointing their existing tool configuration at Lemonade's Ollama-compatible endpoint, but the model library and format coverage differ.

The AGENTS.md and CLAUDE.md files in the Lemonade repository suggest the project itself uses AI coding agents during development, which is consistent with its positioning as infrastructure for the same category of tools. This self-referential use case is evidence that the OpenAI and Anthropic API compatibility is not theoretical: the project's own development workflow depends on it working correctly.

## Conclusion

Lemonade is the right fit for developers who need a drop-in local replacement for cloud AI APIs and want to use existing tools like Claude Code, Open WebUI, or n8n without modifying their configuration. It is not the right fit for teams that need centralized usage tracking, fine-grained access control across multiple users, or models that exceed the VRAM available in a single PC. Before deploying it, verify that the target machine's GPU or NPU is listed in Lemonade's supported backend matrix, because inference performance depends on a matching backend being compiled into the release.

## FAQ

### What hardware does Lemonade require to run local AI models?

Lemonade supports CPU inference, NVIDIA CUDA GPUs (Turing or newer), Apple Silicon via Metal, AMD GPUs via Vulkan, and AMD Ryzen AI NPUs. The README notes AMD engineers contributed optimizations for Ryzen AI, Radeon, and Strix Halo hardware. Minimum VRAM requirements depend on the model selected.

### Which API formats does Lemonade Server support for app compatibility?

Lemonade Server exposes OpenAI-compatible, Anthropic-compatible, and Ollama-compatible API endpoints. Applications configured to connect to any of those cloud APIs can be redirected to a local Lemonade instance without changing the application code.

### How do I manage multiple models without changing app configuration in Lemonade?

The lemonade alias command lets you assign a stable name to any model. Applications reference the alias, and swapping to a different model requires only lemonade alias add to point the alias at the new model. The README shows this pattern as an active-standby failover mechanism.

## Sources

- [lemonade-sdk/lemonade on GitHub](https://github.com/lemonade-sdk/lemonade)
- [License: Apache-2.0](https://github.com/lemonade-sdk/lemonade/blob/main/LICENSE)
- [Project website](https://lemonade-server.ai/)
- [README](https://github.com/lemonade-sdk/lemonade/blob/main/README.md)
- [Releases](https://github.com/lemonade-sdk/lemonade/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lemonade-sdk-lemonade
