Memvid: a single-file memory layer for AI agents, without a vector database
Memory layer for AI Agents. Replace complex RAG pipelines with a serverless, single-file memory layer. Give your agents instant retrieval and long-term memory.
At a glance
- What is it?
- Memvid packs content, embeddings, search structures and metadata into one append-only .mv2 file, so an agent can carry its memory instead of querying a server. The design is coherent and the Rust core is the real thing; the surrounding claims and the ecosystem are what you should check before committing.
- Who is it for?
- Adopt Memvid if you want an offline-first, portable memory file that a Rust service or a Python agent can open without running a vector database, and if you accept that the .mv2 format is young. Do not adopt it if you need a managed service with a support contract, if your retrieval must span hundreds of gigabytes, or if you cannot tolerate a format that is still moving at version 2.0.140.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 77 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The problem Memvid targets: retrieval without a retrieval service
A conventional RAG pipeline is a small distributed system. You run an embedding model, a vector index, a metadata store, and a document store, then wire them together and keep them in sync. Memvid's answer is to collapse that into one file. The README states the goal plainly: "Instead of running complex RAG pipelines or server-based vector databases, Memvid enables fast retrieval directly from the file." The file is the database, the index and the payload.
That framing decides the audience. It suits an agent that runs on a laptop, in a container with no network, or on a device where you cannot install Postgres and pgvector. It suits a workflow where memory is an artifact you hand to someone else, the way you would hand over a SQLite file. It does not suit a team that wants a managed endpoint with an SLA, and it does not suit a corpus where a single file becomes an operational liability.
The repository description calls it a "serverless, single-file memory layer" and the README adds that it is model-agnostic and works fully offline. Those two properties are the actual product. Everything else in the README, including the benchmark table, is a claim about how well the file performs rather than what it is.
Smart Frames: an append-only timeline instead of an index rebuild
The mechanism the README describes is borrowed from video encoding. Memory is stored as a sequence of Smart Frames, each an immutable unit holding content plus timestamps, checksums and basic metadata. Frames are grouped so that compression, indexing and parallel reads can operate over groups rather than over one flat table.
Three consequences follow from that design, and the README names them. Writes are append-only, so a new memory does not rewrite existing data. Queries can target past memory states, because earlier frames are still present. And crash safety comes from committed, immutable frames rather than from a write-ahead log you have to manage. The README also mentions time-travel debugging: rewind, replay or branch any memory state. In practice that means the file behaves like a rewindable timeline, which is a different mental model from a mutable table of vectors.
The Cargo.toml backs up the format claim. The core depends on blake3 for hashing, ed25519-dalek for signatures, zstd and lz4_flex for compression, and bincode for serialization. The crate description calls it "crash-safe, deterministic, single-file AI memory." A separate MV2_SPEC.md sits at the repository root, which is where the on-disk layout is actually defined. If you plan to read the file from another language, that spec file is the document you need, not the README.
The trade-off is real. Append-only means deleted or superseded content is not reclaimed in place. Over a long-lived agent, the file grows, and the only way back is to rewrite it into a new capsule. The README does not document a compaction command or a garbage collection step.
Installing Memvid and running a first local memory file
The README offers four entry points. The Rust crate is memvid-core, added with cargo add memvid-core or pinned in Cargo.toml. The Python SDK is memvid-sdk from PyPI. The Node.js SDK is @memvid/sdk. The CLI is memvid-cli from npm. The Rust crate requires Rust 1.85.0 or newer, per the README's requirements section, and Cargo.toml sets rust-version = "1.85.0" with edition 2024.
Start with the crate. The README gives this dependency line:
[dependencies]
memvid-core = "2.0"Memvid is feature-gated, so the default build is not the full system. The README lists the flags you can enable: lex for BM25 full-text search via Tantivy, vec for HNSW vector search with local ONNX embeddings, pdf_extract for pure Rust PDF text extraction, clip for CLIP image embeddings, whisper for audio transcription, and api_embed for OpenAI cloud embeddings. The Makefile's default feature set is narrower than the full list:
cargo check --features lex,pdf_extract
cargo build --features lex,pdf_extractIf you want vector search rather than keyword search, you have to add vec yourself. The Makefile also has a download-models target that fetches SymSpell frequency dictionaries into data/ with curl, which tells you that some spelling or correction path expects those files to be present locally. The README does not document that target.
For a Python-first workflow, the README's install line is short:
pip install memvid-sdkThe repository ships examples you can read before writing anything: examples/basic_usage.rs, examples/pdf_ingestion.rs, examples/text_embedding.rs, examples/openai_embedding.rs, and examples/clip_visual_search.rs. Reading basic_usage.rs is the fastest way to learn the actual API surface, because the README stops at installation and never shows a complete program. That is a genuine gap: there is no end-to-end snippet in the README that opens a capsule, inserts a frame and queries it.
Where Memvid is the wrong tool
The first limitation is scale. A single file that holds content, embeddings and a search index has to be read and memory-mapped by whichever process opens it. The README does not state a maximum capsule size, and no section discusses what happens when the file no longer fits comfortably in the page cache. If your retrieval set is large enough that you would normally shard a vector index, a single file is the wrong shape, and the README offers no sharding story.
The second is concurrency. The core depends on fs2, a file-locking crate, which implies advisory locking around file access. The README does not describe multi-writer behaviour, nor what a second process sees while a first is appending. For a single agent process this is fine. For a fleet of workers all writing memories to one shared capsule, you are in undocumented territory.
The third is the benchmark table. The README claims +35% SOTA on LoCoMo, +76% multi-hop and +56% temporal versus an industry average, plus 0.025ms P50 and 0.075ms P99 latency with 1,372x higher throughput than standard. Those numbers are stated without a link to the harness in the README text, and "industry average" is not defined in the excerpt. The README does say the evaluation is open source and uses LLM-as-Judge on LoCoMo with ten conversations of roughly 26K tokens each. Treat the numbers as a claim to reproduce, not as a property of the library.
The fourth is the packaging. Cargo.toml carries a comment that extractous uses GraalVM native compilation and does not work on Windows ARM or WSL2 on ARM, and that it must be enabled explicitly with --features extractous. If your target is Windows on ARM, the full document extraction path is closed to you, and you fall back to pdf-extract or pdf_oxide.
Memvid compared with Mem0 and with a plain RAG stack
The most common comparison is Memvid versus Mem0, and the difference is architectural rather than a matter of tuning. Mem0 is a memory service: memories are stored and retrieved through a running system, typically backed by a vector store and, in its hosted form, an API. Memvid's README explicitly rejects that shape, describing a memory layer that is "portable, versioned, and portable memory, without databases" and retrieval "directly from the file."
The practical difference shows up in three places. Deployment: Mem0 needs a service to be running; Memvid needs a file to exist. Portability: a .mv2 capsule can be copied, signed (ed25519-dalek is a dependency) and shipped; a Mem0 store lives behind its own interface. Debuggability: because Memvid frames are immutable and ordered, the README can offer rewind and replay of past memory states, which a mutable vector store does not give you for free.
The cost is everything a service provides. There is no HTTP API to point a non-Rust, non-Python client at, no dashboard, no managed backups, no horizontal scaling. Against a hand-rolled RAG pipeline, Memvid removes the vector database and the sync job, but it does not remove the embedding model: the vec feature uses local ONNX embeddings, and api_embed calls OpenAI. You are still choosing and paying for embeddings either way.
Licence, release cadence and what an upgrade costs
Memvid is Apache-2.0, stated in both the README badge and the Cargo.toml license field. That is a permissive licence with an explicit patent grant, and it imposes no copyleft obligation on your own code. It does require that you preserve the licence and notice files when you redistribute the crate or a binary that embeds it. Nothing here is legal advice; if you ship the library inside a product, have counsel read the NOTICE and attribution requirements rather than relying on a summary.
The upgrade picture is the part to weigh carefully. The published releases are v2.0.138 on 2026-03-03, v2.0.139 on 2026-03-13, and v2.0.140 on 2026-05-27, and Cargo.toml pins version = "2.0.140". Between March and May there were three patch releases, then a gap. The last push to the repository was on 2026-07-14, which is two months before today. The project is not archived, but the release history shows a format that is still being corrected at the patch level, and MV2_SPEC.md exists precisely because the layout matters.
The upgrade cost is therefore format risk, not API churn. A .mv2 file written by an earlier 2.0.x build is a binary artifact; if the spec changes, you need a migration path, and the README does not document one. The CHANGELOG.md at the repository root is the file to read before bumping the crate, and you should keep a copy of the capsule alongside the version that wrote it. There is no documented downgrade or rollback procedure.
Editorial conclusion
Adopt Memvid if you want an offline-first, portable memory file that a Rust service or a Python agent can open without running a vector database, and if you accept that the .mv2 format is young. Do not adopt it if you need a managed service with a support contract, if your retrieval must span hundreds of gigabytes, or if you cannot tolerate a format that is still moving at version 2.0.140. Before you commit, verify three things: that the feature flags you need (lex, vec, pdf_extract, clip, whisper) build on your target platform, that the .mv2 file size stays acceptable for your corpus, and that the release cadence since 2026-03-13 matches your upgrade tolerance. The last push to the repository was on 2026-07-14.
Frequently asked questions
What is Memvid?
Memvid is a memory layer for AI agents that packages content, embeddings, search structures and metadata into a single file. The README describes it as a portable, model-agnostic system that retrieves directly from the file instead of from a server-based vector database.
How to use Memvid?
Add memvid-core as a Rust dependency, or install the Python SDK with pip install memvid-sdk, the Node SDK with npm install @memvid/sdk, or the CLI with npm install -g memvid-cli. The README stops at installation, so the examples directory, starting with examples/basic_usage.rs, is where the actual usage pattern is shown.
Is Memvid a legitimate project?
The repository is a real, unarchived Apache-2.0 codebase with a published crates.io package, a Cargo.toml, a test suite and an MV2_SPEC.md format document. Its last push was on 2026-07-14, and the most recent release is v2.0.140 from 2026-05-27.
Is Memvid safe to use?
The core depends on ed25519-dalek for signatures and blake3 for checksums, and the README states that crash safety comes from committed, immutable frames. The README does not document multi-writer locking behaviour, so treat concurrent writes from several processes as unverified.
How does Memvid compare with RAG?
Memvid keeps the retrieval step but removes the pipeline around it: content, embeddings and the search index live in one .mv2 file rather than in a vector database plus a sync job. You still choose an embedding model, either local ONNX through the vec feature or OpenAI through api_embed.
What is a Memvid alternative?
Mem0 is the closest alternative and takes the opposite approach: it is a memory service backed by a running system rather than a single portable file. Choosing between them is mostly a question of whether you want a service endpoint or an artifact you can copy and ship.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/memvid-memvid)
Community notes