# Vellum Assistant: a self-hostable AI assistant with eight memory types

> Vellum Assistant is a TypeScript AI assistant that ships a memory model, a proactivity loop and a sandboxed tool runtime. It is aimed at people who would otherwise assemble that stack themselves, and it is honest about being a young project.

**vellum-ai/vellum-assistant** — An AI Assistant that s easy to setup, does your work 24/7, knows your preferences and gets better over time.

- Repository: https://github.com/vellum-ai/vellum-assistant
- Website: https://vellum.ai
- Stars: 1,319 · Forks: 200
- Language: TypeScript
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/vellum-ai-vellum-assistant

## The setup cost Vellum Assistant is trying to remove

The README opens with a direct comparison: if you have set up a personal AI on OpenClaw, Hermes Agent or Claude Code, you already know how long it takes and how many times you have to start over to get it right. Vellum's claim is that it gives you the result out of the box, one download away. That is the whole pitch, and it defines the audience: people who want a personal assistant that remembers them, not a framework they will spend a weekend wiring up.

The target user is someone who has already tried the DIY route and lost patience with it. The README calls out the alternative explicitly, describing the usual approach as "a SQLite + Markdown file you maintain yourself." Vellum replaces that with eight named memory types, each with its own staleness window, hybrid dense plus sparse retrieval, and per-user and per-channel isolation. Whether eight is the right number is a design choice, not a law of nature, but the point is that the categories are decided for you rather than left as an exercise.

## How the memory, identity and proactivity layers fit together

Three subsystems do most of the work. Memory stores structured items (identity, preferences, projects, events) extracted from conversations, with source attribution and deduplication. Embeddings run locally by default on ONNX, with an automatic fallback to cloud providers. Identity lives in files rather than a database: behaviour is defined in SOUL.md, the assistant writes its own personality files during onboarding after observing how you communicate, keeps a per-user journal of reflections, and uses NOW.md as a scratchpad for current focus and active threads. Proactivity is a loop, not an event hook. Every hour the assistant re-reads its notes, looks for anything unfinished or due soon, and messages you if something needs attention. Notifications route to the right channel and are suppressed during an active conversation.

The security model is the part worth reading closely. Actor identity is resolved once into one of three roles (guardian, trusted, unknown) and enforced everywhere, so an unknown actor cannot read memory, trigger tools or escalate. Credentials live in a separate process and never reach the model. Every tool call runs in a sandbox, and the default is to deny. Computer use follows the same rule: the assistant works in its own sandbox and only reaches your actual machine with approval, granted once, for ten minutes, or always. That is a stricter posture than most personal-agent projects adopt, and it is the strongest argument in the README.

## Installing Vellum Assistant from the CLI and hatching an assistant

The README says the CLI works but that the desktop app is the primary focus; the CLI is described as being for advanced users, contributors, and non-macOS environments. Installation is two commands. The first installs the global package with Bun, the second runs the onboarding flow that creates your assistant:

```bash
bun install -g vellum
vellum hatch
```

If you would rather build from source, the repository ships a setup.sh at the top level. Cloning, running it, sourcing your shell profile and hatching looks like this:

```bash
git clone https://github.com/vellum-ai/vellum-assistant.git
cd vellum-assistant
./setup.sh
source ~/.bashrc
vellum hatch
```

Once an assistant exists, the day-to-day commands are short. wake starts services, sleep stops them while keeping data, client gives you a terminal interface, ps lists running assistants, terminal opens a shell inside a managed assistant container, and upgrade moves you to the latest version:

```bash
vellum wake
vellum client
vellum ps
vellum upgrade
```

All commands target the default assistant. If you run more than one, the README says to pass the assistant ID as the second argument. For reaching a self-hosted assistant from a phone or another computer, the documentation covers opening a tunnel and pairing the device. Before you start, copy .env.example to .env; existing environment variables take precedence over values in that file. The proxy allowlist defaults to *.vellum.ai, and localhost is always allowed.

## Where Vellum Assistant is the wrong tool

The README is unusually candid that the CLI is not the main product. If you want a stable command-line contract to script against, you are working on the surface the maintainers describe as secondary. The desktop app is where the attention goes.

Version churn is the second constraint. The releases run v0.11.5, v0.11.6 and v0.11.7 within a single week in August 2026, and the version number is still 0.x. Nothing in the README documents a rollback procedure, a migration guide between minor versions, or a compatibility guarantee for SOUL.md, NOW.md and the memory store. If your assistant accumulates months of episodic and narrative memory, the file formats underneath it are the thing you cannot afford to lose, and the documentation does not yet tell you how to move that data forward safely.

The third limit is scope. This is a personal assistant, not a multi-tenant platform. Memory isolation is per-user and per-channel inside one assistant, and the security model is built around a single guardian. If you need an agent serving many unrelated end users with separate billing and audit boundaries, the actor model here (guardian, trusted, unknown) does not map onto that.

## How Vellum Assistant differs from OpenClaw and Claude Code

The README names OpenClaw, Hermes Agent and Claude Code as the tools people arrive from. The difference is where the work sits. With those, you supply the memory layer: you decide the storage, the retrieval strategy, the staleness rules and the identity files, and you maintain them. Vellum ships those decisions as defaults, including the local ONNX embedding path and the eight memory categories with their individual staleness windows.

Claude Code is also a coding agent first, oriented around a repository and a terminal session. Vellum Assistant is oriented around a person across channels: macOS, iOS, Web, Voice, Email, Telegram, Slack and Twilio, with one assistant and one memory shared between them. That is a different shape of product, and it means the comparison only holds at the layer below, where both need a model provider and a tool runtime. Vellum's multi-provider list covers Anthropic, OpenAI, Google Gemini, Fireworks, OpenRouter, MiniMax, Atlas Cloud, any OpenAI-compatible endpoint, and local models through Ollama.

## Licence, maintenance and what an upgrade actually costs

The repository is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are preserved. That is the permissive end of the spectrum, and it means self-hosting Vellum Assistant does not create source-disclosure obligations. The MIT grant covers the code in this repository; it does not automatically cover the managed Vellum Platform service, the hosted docs, or third-party model providers you connect. Read LICENSE and the terms of whichever provider you point it at rather than assuming one covers the other. This is a description of the licence text, not legal advice.

On maintenance: the repository is not archived, and the last push was on 2026-08-27, the same day as the v0.11.7 tag. That is recent enough to call the project current, but three releases in seven days is also a signal about stability rather than a guarantee of it. The practical upgrade cost sits in the workspace layout. package.json defines a Bun workspace with more than two dozen packages, plus pinned overrides for es-toolkit, lodash and path-to-regexp and three patched dependencies (app-builder-lib, storybook, and the Capacitor camera preview plugin). Patches against upstream packages are the part that tends to break on a dependency bump, because a patch written for one version will not apply to the next. If you self-host, budget for reading those patches when you upgrade, or pin your Bun lockfile and move deliberately.

## Conclusion

Adopt Vellum Assistant if you want a personal assistant whose memory, proactivity and permission model are already designed, and you are willing to run a Bun-based workspace or pay for the managed runtime. Do not adopt it if you need a stable API contract, a documented upgrade path between minor versions, or a project with a long release history; v0.11.7 is the newest tag and the README does not describe rollback. Before committing, run vellum hatch locally, read CONSTITUTION.md and ARCHITECTURE.md, and confirm that your model provider and channel mix are covered by the multi-provider list.

## FAQ

### Is Vellum Assistant an AI tool?

Yes. It is a personal AI assistant that runs as a TypeScript workspace, with a managed runtime on Vellum Platform or a self-hosted mode using the same codebase and data model.

### Is Vellum Assistant free to use?

The repository is MIT licensed, so the code can be self-hosted at no licence cost. The README also describes a managed option that signs in through Vellum Cloud, and the terms of that service are not covered in the repository.

### How do I install Vellum Assistant?

The README gives two paths: bun install -g vellum followed by vellum hatch, or cloning the repository, running ./setup.sh, sourcing your shell profile and then running vellum hatch.

### Which model providers does Vellum Assistant support?

The README lists Anthropic, OpenAI, Google Gemini, Fireworks, OpenRouter, MiniMax and Atlas Cloud, plus any OpenAI-compatible endpoint, with local models running through Ollama. Embeddings default to local ONNX and fall back to cloud providers.

### Can Vellum Assistant run entirely on my own machine?

Yes. The README describes a Local mode in which everything runs on your machine, alongside a Managed mode that signs in through Vellum Cloud and needs no local runtime.

### What are the eight memory types in Vellum Assistant?

The README names episodic, semantic, procedural, emotional, prospective, behavioral, narrative and shared memory, each with its own staleness window and hybrid dense plus sparse retrieval.

## Sources

- [Official documentation](https://vellum.ai)
- [Official README](https://github.com/vellum-ai/vellum-assistant#readme)
- [Project repository](https://github.com/vellum-ai/vellum-assistant)
- [Release notes](https://github.com/vellum-ai/vellum-assistant/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vellum-ai-vellum-assistant
