# Ataru: A Local Search Core for AI Agent Session History

> Ataru indexes Claude and Codex transcripts into a Tantivy and Jieba search layer with optional semantic recall, exposing the same contract to a Tauri desktop app, a JSON CLI and installable Agent Skills. The judgement: the retrieval design is sound, but the CLI path is keyword-only and full rebuilds are not cheap.

**lovstudio/Ataru** — High-performance AI memory retrieval for local agent history — a Rust search core (SDK / API / JSON CLI) plus a desktop GUI. Tantivy + Jieba keyword search, optional semantic recall, stable Turn/Run/Session/Project IDs.

- Repository: https://github.com/lovstudio/Ataru
- Website: https://lovstudio.ai/app/ataru
- Stars: 358 · Forks: 40
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/lovstudio-ataru

## The problem Ataru solves: answers you already produced, buried in transcript files

The README's framing is narrow and worth taking literally. During AI-assisted coding, useful answers have already happened once: a debugging session, a key command, an architecture trade-off, a piece of context discussed a few weeks ago. Those answers live in per-agent local session files, and finding them again depends on memory, directory names and manual scrolling. Ataru's stated goal is not to manage a running agent. It is to make past work searchable again.

That distinction decides who the tool is for. If you run one agent in one repository and rarely revisit old sessions, there is nothing here to retrieve. If you have accumulated Claude CLI, Claude App or Web, and Codex sessions across several projects, the retrieval problem is real and the alternatives are grep over JSONL or nothing. Ataru normalises those differing source formats in an adapter layer, then indexes them locally. Default operation is offline. Semantic retrieval is described as an optional enhancement that degrades clearly to keyword search when it is unconfigured or times out.

## How the retrieval core works: adapters, a Tantivy index, and four aggregation levels

The architecture diagram in the README shows a local-first modular monolith. Source adapters read local transcript files and normalise Claude, Codex and legacy formats. An ingestion and indexing stage owns a manifest, incremental catch-up and reconciliation, and writes into two stores: a Tantivy keyword index with Jieba tokenisation, and an optional SQLite vector store. Embeddings are reached only through explicit opt-in, drawn as a dotted edge in the diagram.

The public boundary is split between an sdk layer and an api layer. The sdk defines SearchRequest, SearchResponse and SearchHit, the SearchLevel enum (turn, run, session, project), the SearchMode enum (auto, keyword, semantic, hybrid), stable entity IDs, warnings and error boundaries. The api layer handles validation, orchestration and fallback. Aggregation sits alongside it so one hit can be rolled up to Turn, Run, Session or Project level, which is what reduces duplicate results when the same exchange was indexed in several chunks.

Jieba is the interesting choice. Tantivy alone tokenises on whitespace and punctuation, which is poor for Chinese text and for the mixed content that agent transcripts actually contain: package names, domains, error strings, code identifiers. Pairing the two is what the README means by Chinese-friendly full-text search. The four-level ID scheme is the other design decision worth noting. Because hits carry stable projectId, sessionId, messageId and lineNumber, a result can be traced back to the original session, message and line rather than a truncated summary. The README explicitly warns against re-deriving entity IDs from titles or display paths.

## Installing Ataru and running a first search

There are two entry points. The desktop application is distributed through GitHub Releases: the README says to download the build for your platform and open the search page. On first use Ataru checks the local index, and if the index is missing, stale, or in an idle state it starts a build, reporting progress through the search-index:build event. Queries only execute once the state is ready. The README notes the index is derived data, the original sessions are never rewritten, and a failed rebuild leaves the previous healthy index in place.

The second entry point is the development build, which the README gives as a clone, install and dev command sequence. The --recursive flag matters because the repository has a .gitmodules entry.

```bash
 git clone --recursive https://github.com/lovstudio/Ataru.git
 cd Ataru
 pnpm install
 pnpm dev:app
```

For a frontend-only loop with hot reload and no Rust restart, the README lists pnpm dev together with pnpm dev:app:no-watch. The package.json scripts confirm dev maps to vite and dev:app maps to tauri dev, with dev:app:no-watch passing --no-watch to the Tauri CLI.

If you are driving Ataru from an agent rather than a person, the install path is the published Skills. The README gives these two commands:

```bash
 npx lovstudio skills add ataru-indexing
 npx lovstudio skills add ataru-search
```

Both Skills are driven by a JSON CLI and do not maintain a second index. The README states the binary is parsed and then gated on --version, requiring 0.41.3 or later, specifically to stop older builds from treating CLI arguments as a request to open a desktop window. After installing, the minimum workflow the README documents is ensure_index, then search, then inspect, then return stable context. Inside a Tauri host, that first step looks like this:

```ts
 const status = await invoke("get_search_index_status");

 if (status.state !== "ready") {
   await invoke("start_search_index_build", { force: false });
   // wait for search-index:build until state === "ready" or state === "error"
 }

 const response = await invoke("ataru_search", {
   request: {
     query: "上次是怎么解决索引没有更新的？",
     level: "turn",
     mode: "auto",
     limit: 20,
   },
 });
```

What you should see is a response carrying hits, the mode that was actually used, elapsed time, warnings, and deep-link locators. If semantic recall was unavailable, the README says to expect an ATARU_*_FALLBACK warning rather than silence.

## The real limits: keyword-only CLI, index build cost, and a refusal to fake zero hits

The most consequential limitation is stated plainly in the README: the CLI has keyword mode only, so a Skill response returns mode as keyword and semanticAvailable as false. If you need semantic or hybrid recall, the README directs you to the desktop app. That splits the product into two capability tiers, and the tier an automation pipeline can reach is the weaker one. Anyone planning to build a retrieval service on the JSON CLI should read that line twice.

The second limit is index cost. The README reports a measured corpus of 2705 sessions, 847526 messages and a 6.6GB index, on which a turn-level query takes roughly 48 seconds. That figure is why the search Skill defaults to a 180-second timeout and the read step to 300 seconds. It also implies that a full rebuild of a corpus that size is not something to trigger casually, which is consistent with incremental catch-up being the default path and with the manifest and reconcile machinery existing at all.

The third is a deliberate refusal that is easy to misread as a bug. When the index is not ready, lov-ataru-search refuses to execute and returns ATARU_INDEX_NOT_READY or ATARU_INDEX_BUILDING. The README's stated reason is that a missing index must never be disguised as zero hits. That is the correct behaviour, but it means an unattended agent loop must handle those two error codes explicitly instead of treating an empty result list as success. Where Ataru is the wrong tool: if you want a hosted service that indexes sessions for a whole team, or if you want to manage agents that are currently running, neither is what this project does.

## How Ataru differs from grep over transcripts and from a general RAG stack

The obvious alternative is recursive grep over the transcript directory. The difference is not just speed. Grep returns lines with no notion of which turn, run, session or project they belong to, so a match on a common token produces a wall of context-free hits, and Chinese text is effectively unsearchable without segmentation. Ataru's adapter layer normalises source formats first, so a query returns hits that carry stable IDs and can be aggregated to a chosen level, and Jieba makes Chinese queries work at all.

A general RAG stack is the other comparison. A typical pipeline chunks documents, embeds everything, and retrieves by vector similarity alone, which is weak at exact matches: a specific error string, a package name, a domain. Ataru inverts that default. Keyword search over Tantivy is the baseline that always works; the vector store is optional, opt-in, and its unavailability is reported as a fallback warning rather than an error. The README describes the ai module as combining intent, semantic recall and RRF, which is the reciprocal rank fusion used to merge keyword and vector result lists when both are available. The trade-off is that Ataru is scoped to agent transcripts specifically, with adapters for Claude and Codex. Point it at a general document corpus and you are outside what its adapters were built for.

## Maintenance, versioning and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-08-24. The release cadence visible in the release list is tight: v0.41.4, v0.41.5 and v0.41.6 all landed on 2026-08-23 and 2026-08-24. The version number in package.json is 0.41.6, matching the latest release tag, and the version script chains changeset version with a Cargo version sync, so the npm and Rust sides are kept aligned by tooling rather than by hand.

That cadence has a practical cost. The Skills gate on a minimum binary version of 0.41.3, which means a Skill installed today can refuse to run against a binary from a few weeks ago. If you pin Ataru, pin the binary and the Skill together. The index schema is also versioned: the startup sequence reads the index manifest and schema version, and treats a stale schema as a reason to rebuild. An upgrade that changes the schema therefore triggers reindexing, which is where the 6.6GB figure becomes a planning number rather than a curiosity.

The licence is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant, with the usual obligations around preserving notices and stating changes. The README does not discuss the licences of the bundled third-party components, and the repository does contain a third-parties directory; if you redistribute a build, check what is in there rather than assuming the top-level licence covers everything. Nothing here is legal advice.

## Conclusion

Adopt Ataru if you already have months of Claude or Codex transcripts and you keep re-answering questions those transcripts already contain, and if you want the desktop GUI, because that is the only client that can run semantic or hybrid recall. Do not adopt it if you expect the CLI or the Agent Skills to do vector search: the README states the CLI is keyword-only, so mode is always keyword and semanticAvailable is always false there. Before committing, verify your index size against the documented figure, since the README reports about 48s for a turn-level query over 2705 sessions and 847526 messages, and confirm that your installed binary reports version 0.41.3 or later, because the Skills gate on --version and older builds can fall through into opening a desktop window.

## FAQ

### What is Ataru?

Ataru is a local search layer for AI agent session history. It indexes transcripts from sources such as Claude and Codex into a Tantivy and Jieba keyword index with an optional semantic store, and ships as a Rust core with a Tauri desktop GUI, a JSON CLI and installable Agent Skills.

### How do you use Ataru?

The README's minimum workflow is ensure_index, then search, then inspect, then return stable context. In practice that means calling get_search_index_status, starting a build with start_search_index_build if the state is not ready, waiting for the search-index:build event, and then calling ataru_search with query, level, mode and limit.

### Does Ataru support semantic search from the command line?

No. The README states the CLI has keyword mode only, so a Skill response always reports mode as keyword and semanticAvailable as false. Semantic or hybrid recall requires the desktop application.

### What happens if the Ataru index has not been built yet?

The lov-ataru-search Skill refuses to execute and returns ATARU_INDEX_NOT_READY or ATARU_INDEX_BUILDING, on the stated principle that a missing index must not be disguised as zero hits. The desktop search page instead starts a build automatically and waits for the state to become ready.

### What licence is Ataru released under?

Apache-2.0, per the repository licence file and the README badge. The repository also contains a third-parties directory whose contents the README does not describe.

## Sources

- [License: Apache-2.0](https://github.com/lovstudio/Ataru/blob/main/LICENSE)
- [lovstudio/Ataru on GitHub](https://github.com/lovstudio/Ataru)
- [Project website](https://lovstudio.ai/app/ataru)
- [README](https://github.com/lovstudio/Ataru/blob/main/README.md)
- [Releases](https://github.com/lovstudio/Ataru/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lovstudio-ataru
