# PrismOS-AI: a local Ollama desktop whose memory is a SQLite graph you can walk through

> The app runs a local model over documents you drop on it and writes what it learns into a SQLite knowledge graph with a 3D and 2D atlas, and the network boundary is documented precisely enough to be checkable. Five agent roles vote before an answer appears, and a wasmtime module approves the operations.

**mkbhardwas12/prismos-ai** — Desktop AI that runs on your laptop. Drop a file, ask a question, it remembers. Nothing leaves the machine.

- Repository: https://github.com/mkbhardwas12/prismos-ai
- Website: https://github.com/mkbhardwas12/prismos-ai
- Stars: 432 · Forks: 219
- Language: Rust
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/mkbhardwas12-prismos-ai

## The network promise is scoped precisely, and the project says where it stops

Most local-first claims are one sentence. This one is a section with a specific boundary, and the specificity is the point.

Private inference requests, which includes document text, images, summaries, and embeddings, go to a fixed local daemon address. The desktop and command line inference clients disable proxies and redirects, so a configured proxy cannot capture them. The configurable daemon address exists for management and status only: the desktop inference path ignores it, and the command line ask path rejects a remote or custom endpoint outright.

Then the disclaimer, which is the most important sentence in the section: this does not attest that the separately managed daemon or the model it is running is offline. The app controls its own clients. It does not control what the daemon does.

The webview is constrained by a content security policy whose connection allow-list contains only the app itself and the two loopback spellings of the daemon port. And again the scope is stated: that constrains webview requests, not Rust network clients, not the daemon, and not your system browser.

What can still reach outside is enumerated: installers, model and voice downloads, configured model management endpoints, and update checks. So the accurate summary is that your documents stay local, while downloads and version checks do not.

## Memory is a SQLite graph, and the atlas is how you check what it learned

The persistence claim is that answers and concepts persist to a local SQLite knowledge graph, and that the next conversation retrieves from it. That is the difference between a chat app and a memory system, and the second half matters as much as the first.

What makes it inspectable is the atlas. The documentation points to a 3D and 2D view of the graph with a set of specific affordances: filter by source, read the notes attached to a node, trace connections between nodes, and scrub through a timeline. So the claim that the graph is yours is not only about the file being on your disk, it is about being able to look at it and correct it.

The timeline in particular is the feature that makes the memory model honest. An agent memory that cannot be reviewed in order is a memory you cannot debug, and being able to see when something entered the graph and what it was attached to is what lets you tell a wrong answer from a wrong source.

The graph rendering is also visible in the dependencies: there is a 2D and a 3D force-directed graph library, a text sprite library for labels, an animation library, and a Markdown renderer. So the atlas is a real component with a real dependency cost, not a placeholder in a screenshot.

The other optional components follow the same pattern of being off unless asked for. Web research is user-directed, meaning only the URLs you type, plus an email keeper over IMAP and a market data keeper. None of them is enabled by default.

## Documents are chunked and retrieved with TF-IDF, which is a deliberate non choice

The document path supports the four formats a working person actually has: PDF, the Microsoft Office word processor format, the presentation format, and the spreadsheet format.

Two decisions in that row are worth pausing on. Text is extracted on the device, which follows from the network boundary and is the reason a PDF never gets uploaded to do the text extraction. And retrieval uses TF-IDF over chunks rather than naively truncating.

Truncation is what most quick implementations do, and it fails in a specific way: a long document gets its beginning embedded and its end ignored. TF-IDF over chunks means every chunk is scored against your question and the relevant ones surface, so a fact on page forty of a report is findable.

It also means the retrieval is lexical rather than semantic. TF-IDF matches on terms, so a question using different words than the document uses will score poorly, and you would expect a vector index to do better there. The project does not claim a vector index, and pairing this with a local model that could embed is an obvious next step rather than a current capability.

That combination, a small local model for generation and a lexical retriever for grounding, is a coherent design for a tool that must run offline. It is also the design that will surprise you on a paraphrased question.

## Five named roles vote, and a wasmtime module gates the operations

The answering path is not a single call. Five roles are named and they vote on the response before it is shown: an orchestrator, a reasoner, a memory keeper, a tool smith, and a sentinel.

A vote rather than a sequence is the interesting choice. It means the answer is agreed rather than produced and then reviewed, which changes the failure mode. A single model with a self-critique tends to agree with itself; five roles with different jobs have more reason to disagree, and a disagreement is information.

What the roles are for is inferable from the names and worth stating as inference rather than documentation. The orchestrator routes, the reasoner works the problem, the memory keeper touches the graph, the tool smith handles the tool calls, and the sentinel is the one whose job is to object.

The enforcement side is separate from the voting side. Operation approvals go through a small wasmtime policy module, meaning the decision about whether an operation is allowed is executed as WebAssembly inside the app rather than being a prompt the model is asked to respect.

That is a meaningful boundary in an app that can write files, and it is worth distinguishing from the hooks-style pattern where a shell command runs and fails open. Here the policy module is the gate. What it decides, and how it fails, is a question for the code rather than for this page.

## Install verifies a checksum, then downloads a model that outweighs the app

The install script does something worth copying. It resolves the right asset for your architecture, verifies a SHA-256 against the digest the forge publishes, and aborts on a mismatch. It then bootstraps the model runtime with one specific small model if you do not already have it. The command itself is the usual shape, and the note above it tells you to read the script first:

```bash
# macOS / Linux x64 — read the script first: scripts/install.sh
curl -fsSL https://raw.githubusercontent.com/mkbhardwas12/prismos-ai/main/scripts/install.sh | sh
```

The Windows form is per-user and needs no administrator rights.

The size accounting is stated bluntly. Before the first answer you have roughly a fifteen megabyte application, the model runtime, and about two and a half gigabytes of model weights. The stated budget is five to fifteen minutes on a decent connection, and about four gigabytes of free memory to run that model comfortably.

So the app is fifteen megabytes and the thing it depends on is two and a half gigabytes, and anyone budgeting disk or thinking about a laptop should know that before starting.

There is a third route for macOS users through a homebrew tap, with a flag on the cask install to skip the quarantine warning. And for anyone who would rather not install a desktop application at all, there is a standalone command line binary built from source that talks to the local daemon directly, with no interface, no signing, and no quarantine flag involved.

The command line path is also pipeable and takes two environment variables to override the model and the daemon address, so it composes with the rest of a shell.

## Unsigned installers and a missing Linux ARM build are the two install frictions

The signing situation is stated at the top of the install section rather than discovered during it. The installers are unsigned, with the reason given as one maintainer and no certificate yet, and getting a certificate is described as being on the roadmap.

What that means concretely is documented for each platform. On macOS you remove the quarantine attribute from the application bundle, or use the privacy and security setting to open it anyway. On Windows you go through the SmartScreen prompt to more info and run anyway. And the documentation offers two ways out that avoid the warning entirely: use the command line binary, or build from source.

The release asset table also has a gap that is worth planning around rather than discovering. Windows has an installer package and an executable, macOS has separate disk images for Apple Silicon and Intel, and Linux x64 has an AppImage and a Debian package. Linux on ARM has nothing published, and the table says so and points at building from source instead.

The stack underneath is a desktop application framework at version 2 with a React interface, built by a Rust backend, with the global shortcut, shell, updater, and window state plugins in the dependency list. The presence of a global shortcut plugin is what makes the summon over any other application feature work, and the window state plugin is what makes the tray minimisation and restoration behave.

Tests run through a single runner with a watch variant, and the build is a type check followed by the bundler step.

## The sample run admits its own error, which is the most useful thing in the docs

The command line section includes an unedited transcript of a real run, and then an honesty note about it. The timing is given as wall-clock seconds including the model's hidden reasoning pass, and a token rate alongside.

The note is the part to remember. The answer contains a factual wobble: the model states that WebAssembly requires code signing, and it does not. The project says so, attributes it to the size of the model, notes that bigger local models sharpen it, and observes that nothing leaves the machine either way.

Most project documentation shows a good output. This one shows a bad one and explains why, with the hardware stated so the reader knows what produced it.

That choice tells you more about the project than the feature list does. A tool that answers questions about your documents with a two and a half gigabyte model will be wrong sometimes, and the useful question is not whether it is accurate but whether you can tell when it is not. The graph, the per-source filter, and the timeline are what make that possible, and the admission in the docs is consistent with having built them.

The release history is consistent with the same posture. The tags are a public launch, a patch shortly after it, and then a gap of months before the current version, with the last push to the repository well after the newest tag.

## Conclusion

This suits someone who wants to ask questions about their own documents without uploading them, has a machine with room for a couple of gigabytes of local weights, and is willing to accept a small model's accuracy. It is a poor fit if you want a system-level guarantee that nothing touches the network, because the project itself says it is not an operating system sandbox, and a poor fit on Linux on ARM since no binary is published there. Before you install, check three things: how much disk and memory you have, since the first run pulls a model of a few gigabytes and wants roughly four gigabytes of free memory; that you read the network boundary section rather than the summary, because the promise is about the app's own clients and not about the model daemon; and whether unsigned installers are acceptable to you, since there is no signing certificate yet and both macOS and Windows will make you override a warning. The newest release is v0.6.0 from 2026-08-12 and the last push was 2026-09-30.

## FAQ

### What is PrismOS-AI and what does it do with my documents?

It is a desktop app that answers questions about documents you drop on it using a local Ollama model, with text extracted on the device, chunked, and retrieved with TF-IDF. It reads PDF, DOCX, PPTX, and XLSX, and what it learns goes into a local SQLite knowledge graph.

### Does PrismOS-AI send anything to the network?

Private inference goes to a fixed loopback daemon address, with proxies and redirects disabled in the inference clients. But the project says this does not attest that the daemon or its model is offline, and that installers, model and voice downloads, model management endpoints, and update checks can reach external services.

### How do I install PrismOS-AI on Linux ARM?

You cannot from a published binary. The release assets cover Windows, macOS on both architectures, and Linux x64, and the table states that Linux ARM is not published and points at building from source. The macOS path also has a homebrew cask, and a command line binary can be built and used with no desktop app at all.

### Why does macOS or Windows warn me when installing PrismOS-AI?

The installers are unsigned, and the stated reason is that there is one maintainer and no signing certificate yet. On macOS you remove the quarantine attribute or open it anyway through the privacy setting, and on Windows you go through the SmartScreen prompt. The command line binary avoids the warning.

### How good are the answers from PrismOS-AI?

The documentation shows an unedited sample run and points out that the answer contains a factual error, stating that WebAssembly does not require code signing. It attributes this to the size of the local model, notes that larger models sharpen it, and says nothing leaves the machine either way.

## Sources

- [License: MIT](https://github.com/mkbhardwas12/prismos-ai/blob/main/LICENSE)
- [mkbhardwas12/prismos-ai on GitHub](https://github.com/mkbhardwas12/prismos-ai)
- [Project website](https://github.com/mkbhardwas12/prismos-ai)
- [README](https://github.com/mkbhardwas12/prismos-ai/blob/main/README.md)
- [Releases](https://github.com/mkbhardwas12/prismos-ai/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mkbhardwas12-prismos-ai
