# Askimo: a native desktop AI client that keeps RAG and chat history on your machine

> Askimo bundles chat, local RAG, MCP tools, agent CLIs and multi-step Plans into one Kotlin desktop app with pluggable providers. It is a good fit if you want one client for both hosted and local models, and a poor fit if you need a server deployment or a permissively licensed codebase.

**askimo-ai/askimo** — AI Client for chat, RAG, Skills, MCP tools, and agents. Support multiple LLMs (Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM, Gemini, OpenRouter)

- Repository: https://github.com/askimo-ai/askimo
- Website: https://askimo.chat
- Stars: 508 · Forks: 108
- Language: Kotlin
- License: AGPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/askimo-ai-askimo

## What Askimo solves for desktop users with more than one model

The problem Askimo targets is fragmentation. If you use OpenAI for one task, Claude for another and a local Ollama model for anything sensitive, you normally live in three browser tabs, copy context between them by hand, and lose the thread of a long session when a tab reloads. Askimo's README frames the pitch directly: "One app. Every AI model. Your files stay local." The supported provider list in the README covers OpenAI, Claude, Gemini, Grok, Ollama, LM Studio, Docker AI, OpenRouter, NVIDIA NIM, Together AI, vLLM Server, and any OpenAI-compatible endpoint.

The intended user is a developer or analyst who works on a laptop or workstation and wants the model choice to be a per-session setting rather than a reason to change tools. The README also claims the app is native rather than a web wrapper, which matters for the long-conversation case it explicitly calls out: "No crashes, no tab reloads, no lost context."

## How the pieces fit: providers, local RAG, MCP tools and agent CLIs

The architecture visible from the README and the repository layout is a Kotlin application split into modules: desktop/, desktop-shared/, shared/, cli/, and tools/, wired together by settings.gradle.kts. That split is the main structural signal. There is a shared core, a desktop front end, and a separate CLI entry point, which suggests the same provider and storage logic is reused across both surfaces rather than duplicated.

On the data side, the README states that files, the RAG index, conversation history and telemetry stay on the machine, with local SQLite for storage and local usage and cost tracking. The RAG pipeline is described as hybrid BM25 plus vector retrieval, with an AI classifier that decides whether retrieval is needed at all and skips it when the query does not require it. That classifier is the interesting design choice: it trades one extra model call for avoiding a retrieval pass on conversational turns where a file search would return noise.

Beyond chat, Askimo connects to MCP-compatible servers, delegates goals to installed agent CLIs (the README names Claude Code, Codex and Antigravity), and runs multi-step Plans from a form UI, with each step building on the previous one and progress shown live. Skills are defined once and then run with whichever agent CLI fits, which the README presents as an alternative to per-agent duplication.

## Installing Askimo and running a first local RAG query

Askimo is distributed as a signed desktop download rather than through a package manager. The README points to https://askimo.chat/download/ for macOS, Windows and Linux, and the system requirements table lists macOS 11+, Windows 10+, and Linux on Ubuntu 20.04+, Debian 11+ or Fedora 35+, with 250 MB of disk and 50 to 300 MB of memory before any model is loaded. There is no Homebrew, winget or apt command in the README, so the install step is the download page and nothing else.

The quick start is three steps: install and open the app, add a provider, start chatting. For a local model, the provider is a running Ollama instance. The README does not print the exact provider configuration keys, so the concrete values below are the ones the README does name, and the rest is left to the setup guide at https://askimo.chat/docs/desktop/ai-providers/.

```bash
# Ollama must already be running locally before you add it as a provider
ollama serve
ollama pull llama3
```

After that, the flow in the app is: open the provider settings, choose Ollama, point it at the local endpoint, then select it for a session. To use RAG, index a local folder or a web URL from the RAG view, then ask a question that references those files. The README states the index is built locally and that hybrid BM25 plus vector retrieval runs on your machine, so the first query after indexing is the one that pays the embedding cost. For hosted providers, the equivalent step is pasting an API key for OpenAI, Claude or Gemini in the same provider screen.

## The AGPL-3.0 licence is the constraint most teams will hit first

Askimo ships under AGPL-3.0, which is stated in the README badge and in the LICENSE file, with a NOTICE file alongside it. For an individual using the desktop app, that is largely a non-issue. For a company that wants to embed Askimo in a product, or run a modified version as a network service, the copyleft terms are the thing to read before writing code against it. The repository also enforces a Developer Certificate of Origin, per CONTRIBUTING.md, which is about contribution provenance rather than usage rights.

This is not legal advice, and the licence text is the authority. The practical point is that AGPL-3.0 is a deliberate choice by the maintainers, and it rules out some adoption paths that an MIT or Apache-2.0 client would allow. If your organisation already has a policy against AGPL dependencies, Askimo is out on that basis alone, regardless of its features.

## Where Askimo is the wrong tool

Askimo is a desktop application. The README describes it as native, and the repository has a cli/ module, but the README does not document a server mode, a headless daemon, or a Docker image for the application itself. If your requirement is a shared RAG service that a team queries over HTTP, or a CI job that calls a model, this is not the shape of the project.

The second limitation is documentation depth in the README itself. It lists Plans as definable "in YAML or gene" (the sentence is truncated in the README as published) and says they can be exported as PDF or Word, but the README does not give the YAML schema, the directive format, or the MCP server configuration keys. Those details live behind the documentation link. Anyone evaluating Askimo from the repository alone will find the feature list longer than the specification.

Third, the provider matrix is broad, but breadth is not the same as parity. The README does not state which features (vision, tool calling, streaming) work identically across every listed provider, so a workflow that depends on image attachments should be verified against your specific provider before you standardise on it.

## How Askimo differs from a chat UI plus a separate RAG stack

The obvious alternative is to assemble the same capability yourself: a chat client such as Open WebUI or LM Studio for conversation, a separate retrieval service for document search, and an MCP-aware agent runner for tool use. That approach gives you more control and lets you swap each component independently, and it is the right call if you already run a shared inference server or need the RAG index to be a service rather than a local file.

The difference in approach is where state lives. In the assembled stack, the index and conversation history typically sit in a server or a shared database, and clients are thin. In Askimo, the index, history and cost tracking are local SQLite on the machine running the app, per the README, and the app is the client. That makes Askimo simpler to run for one person and harder to share across a team. Askimo also folds agent CLI delegation and Skills into the same client, which the assembled approach would leave to separate tools. The trade is centralisation for convenience.

## Maintenance, releases and what upgrading costs you

The repository is not archived, and the last push was on 2026-09-09, the same date as the v1.5.0 release. Before that, v1.4.18 landed on 2026-08-24 and v1.4.17 on 2026-08-20. That is a release cadence measured in weeks, not months, and the version numbers suggest incremental fixes rather than a rewrite in progress.

For upgrade cost, the README does not document a migration process, a database schema version, or a rollback path for the local SQLite store. It is silent on what happens to an existing RAG index or conversation history when the app version changes. That is the practical risk: if you index a large folder and a later release changes the index format, there is no stated procedure for rebuilding or preserving it. Treat the local store as something you can regenerate, and keep the source folders rather than the index as your record.

## Conclusion

Adopt Askimo if you work on a desktop and want one client for hosted and local models, with a RAG index and conversation history that stay in local SQLite. Skip it if you need a headless server deployment, or if AGPL-3.0 is incompatible with how you distribute your own product. Before committing, open the desktop/ and cli/ modules in build.gradle.kts to confirm the CLI surface matches your workflow, and read the full plans YAML schema in the documentation, since the README only shows the form UI and not the file format.

## FAQ

### Does Askimo send my files or chat history to a server?

The README states that files, the RAG index, conversation history and telemetry stay on your machine, with local SQLite storage, and that the only network calls are the ones you configure for your chosen AI provider. If you select a local Ollama model, that provider call stays local too.

### Which AI providers does Askimo support?

The README lists OpenAI, Claude, Gemini, Grok, Ollama, LM Studio, Docker AI, OpenRouter, NVIDIA NIM, Together AI, vLLM Server, and any OpenAI-compatible endpoint, selectable per session. API keys are entered in the provider settings screen.

### What are Askimo Plans and how do I define one?

Plans are multi-step AI workflows run from a form UI, where each step builds on the previous one and progress is shown live, and they can be exported as PDF or Word. The README says plans can be defined in YAML, but it does not print the schema, so the format has to come from the documentation site.

### Can I run Askimo on Linux?

Yes. The system requirements table lists Linux on Ubuntu 20.04+, Debian 11+ or Fedora 35+, alongside macOS 11+ and Windows 10+, with 250 MB of disk and 50 to 300 MB of memory before any model is loaded. Downloads are on the project's download page.

### What licence is Askimo released under?

Askimo is licensed under AGPL-3.0, as shown in the README badge and the LICENSE file, with a NOTICE file in the repository. The repository also enforces a Developer Certificate of Origin for contributions.

## Sources

- [askimo-ai/askimo on GitHub](https://github.com/askimo-ai/askimo)
- [License: AGPL-3.0](https://github.com/askimo-ai/askimo/blob/main/LICENSE)
- [Project website](https://askimo.chat)
- [README](https://github.com/askimo-ai/askimo/blob/main/README.md)
- [Releases](https://github.com/askimo-ai/askimo/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/askimo-ai-askimo
