# kiwifs: markdown files with git history, six protocols and 62 MCP tools

> KiwiFS is a Go binary that treats a directory of Markdown files as the source of truth and derives everything else from it: a search index, a git history where every write is a commit, and a live event stream. The interesting edge is write identity, because the git commit author comes from a request header that one protocol trusts.

**kiwifs/kiwifs** — Markdown filesystem for agents and teams.

- Repository: https://github.com/kiwifs/kiwifs
- Website: https://docs.kiwifs.com
- Stars: 630 · Forks: 184
- Language: Go
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kiwifs-kiwifs

## Files are the source of truth and every index is disposable

The design argument is short enough to state in one line and everything else follows from it: files are the source of truth, and everything else is a derivative index you can rebuild. The problem it addresses is that Markdown is the shared language of agents and developers, while raw `.md` files are just files, with no search, no versioning and no structure. The alternatives the project names are all worse in a specific way: databases agents cannot read, read-only retrieval layers agents cannot write to, ephemeral sandboxes that vanish, or proprietary SaaS you cannot self-host. The access diagram shows the consequence. An agent works with `cat`, `grep` and `echo` against a mounted directory, and a human works in a web interface with wiki links, a graph view, a block editor, keyboard search, backlinks and a table of contents. Below both sit three derivatives: git versioning as the audit trail, full-text plus vector search as the index, and server-sent events for live updates. Nothing in that list is authoritative, so a corrupted index costs you a rebuild rather than your content.

## The git commit author is a request header

Every write is an atomic commit, and the author recorded on that commit comes from a request header rather than from an authenticated session. Writes take the actor from the `X-Actor` header and use it as the commit author, currently implemented for REST and WebDAV, with the web interface going through REST. The value is normalised before it reaches git: control characters are stripped and the length is capped at 256 characters. When the header is empty or absent the fallback depends on the protocol, with REST falling back to `anonymous` and WebDAV falling back to the server's configured WebDAV actor, which is `webdav` by default. The documentation is candid about the trust model. The header is trusted input, so it only carries the identity you can vouch for. On REST the scoped-token and OIDC middleware overwrite the header with the authenticated identity, so a client cannot forge it. WebDAV authentication is a single shared API key that carries no identity, so the header is taken at face value and anyone holding that key can attribute a write to any actor.

## Extra protocols are build tags, not runtime switches

The deployment files reveal which features are compiled rather than configured. In the compose file the extra protocols appear as commented port mappings with the conditions attached: NFS on 2049 requires a build flag, S3 on 3334 requires a flag, and WebDAV on 3335 requires a flag. The offline embedder works the same way, since building it means compiling with a specific tag before the local model can be used:```bash
kiwifs model download all-minilm-l6-v2
go build -tags onnx -o kiwifs .
```That has a direct consequence for anyone evaluating the project. The feature list advertises six access protocols, but a binary you downloaded rather than compiled may carry fewer of them, and the only reliable way to know is to build it yourself with the tags you need. The same applies to the vector search stack, where an ONNX embedder needs the tag while the hosted providers do not. The recommendation in the documentation for teams running their own infrastructure is to terminate authentication at a gateway that sets the actor header itself and strips any client-supplied value before forwarding, which is the mitigation for the WebDAV trust problem described above.

## One binary, and a dependency list that reaches Firestore and FUSE

The tagline is one binary, zero config, and the module manifest is where the real scope shows. It declares Go 1.26 and requires cloud Firestore, a TOML parser, an HTML-to-Markdown converter, and the AWS SDK v2 with configuration, Bedrock runtime and DynamoDB. From there the list continues through an OpenID Connect library, a filesystem watcher, a go-billy filesystem abstraction, a readability extractor, a MySQL driver, a Graphviz renderer, a feed generator, a FUSE library, a Postgres driver, a fake S3 server for tests, a Notion API client, an Echo web framework, an MCP server library, a BibTeX parser, Parquet, a tokenizer pair, a Redis client, a JSON Schema validator, Cobra, standard-webhooks, Swagger tooling and Testcontainers with an Elasticsearch module. Most of those are optional surfaces rather than what the core does, and several exist for the importers and the protocols rather than for search. Read as a whole, it says the project prefers optional surfaces compiled into one executable over plugins, which is the opposite trade-off to a design where the embedder and the vector store are separate services.

## The runtime image carries Chromium, pandoc, Node and Python

The container build is four stages and the runtime is heavier than a single-binary tool suggests. The interface is built on the host with Node 22 so the output is architecture-independent, the Go binary is cross-compiled for the target architecture with cgo disabled and size-reduced link flags, and a third stage installs runtime dependencies in parallel: git, certificates, a Docker CLI, pandoc, Node with npm, Python with pip, and Chromium, all before creating a non-root user. Two of those installs are allowed to fail silently, with a trailing fallback that ignores the error: a Markdown-to-slides CLI and the MkDocs toolchain. So an image can be missing both of them with nothing in the build log to say so. Chromium is wired up through two environment variables pointing at the same binary path, which tells you browser automation is a supported path rather than an accident. The final stage copies only the compiled binary, creates the data directory owned by the non-root user, declares the port and the volume, and switches to that user. The runtime base is Alpine while the builders are Node 22 and Go 1.26.

## make go-build will happily ship a stale interface

The Makefile encodes the ordering constraint that the embedded interface creates, and it also exposes the shortcut around it. The interface assets are embedded into the binary at compile time, so the full build target builds the interface first and then compiles the binary so the embed picks up the freshest assets. There is a second target that builds only the Go binary, documented as the one to use when the interface has not changed, because the previously built assets are still embedded. That is the right optimisation and the obvious foot-gun: a developer who changes interface code and runs the binary-only target ships the old interface with no error anywhere. The clean target reflects the same layout, since it removes the binary and then deletes specific interface output paths rather than the whole directory. Version stamping comes from a shell call to git describe with a dirty flag, falling back to a placeholder, and the linker flag injects it into a variable in the main package. The rest of the file is conventional: a run target that depends on build, a dev target that runs from source, a Docker development target, a separate interface development server for iterating without recompiling, tests over all packages, and a Swagger generation target.

## The compose file points at pgvector while starting in sqlite search mode

The shipped compose file is worth reading line by line because its defaults are not the defaults the feature list implies. The server service builds from the repository Dockerfile rather than pulling a published image, names the result locally, mounts a knowledge directory onto the data path, and starts with search set to sqlite and versioning set to git. Yet the environment block sets a Postgres vector DSN pointing at the sidecar by default, and the comment above it says to leave those unset to run without semantic search. So the file as shipped is configured for the Postgres path. The sidecar is a pgvector image for Postgres 16 with a health check that runs the readiness probe every five seconds, giving up after ten retries, and the server waits for that health check before starting. The three extra protocol ports are present but commented, as noted. One more distribution detail: the published container image in the quickstart is referenced under a personal Docker Hub namespace rather than the project organisation, while the compose file builds locally, so the two paths can produce different binaries.

## Semantic search needs an embedder configured outside the data root

Vector search is off by default and configured under a vector section of a config file that lives in a hidden directory beside the data, not inside it:```toml
[search.vector]
enabled = true

[search.vector.embedder]
provider = "openai"       # openai | ollama | cohere | onnx | http | ...
model = "text-embedding-3-small"
api_key = "${OPENAI_API_KEY}"

[search.vector.store]
provider = "sqlite-vec"
```The embedder and the store are chosen separately, and the store list is where the architecture shows through, since a local vector extension, an embedded one and four hosted databases all sit behind the same setting. The API key is interpolated from the environment rather than written into the file. The offline route is the interesting one for anyone who cannot send data out. You download a sentence-transformer model through the CLI, compile with the embedder tag, and point the config at the model file with its dimension count, and the tokenizer is optional because it is discovered from the parent directory. A 384-dimension configuration matches the small MiniLM model named in the example. Two operational notes follow from the split between embedder and store: a text-to-embedding run is a batch job over your existing files rather than an on-the-fly call, and switching providers later means re-embedding everything, since the dimension count is part of the index rather than a display preference.

## Conclusion

KiwiFS fits a team that wants agents and humans writing into the same Markdown files and is willing to let git be the audit log, because every write is a commit and everything else is a rebuildable index. Three things to check before you expose it. Write identity is a header, so on WebDAV anyone holding the shared key can attribute a write to any actor, and the documented mitigation is a gateway that sets the header itself. Several protocols and the offline embedder are compiled in behind build tags rather than configured at runtime, which means a binary you did not build may not have them. And the Docker Hub image is published under a personal namespace while the compose file builds from the repository, so confirm which one you are running before you trust its contents.

## FAQ

### How does kiwifs know who wrote a file?

From the X-Actor request header, which becomes the git commit author. On REST the scoped-token and OIDC middleware overwrite it with the authenticated identity, but WebDAV uses a single shared API key with no identity, so the header is taken at face value there.

### Which protocols can kiwifs be reached over?

Six: REST, MCP, NFS, S3, WebDAV and FUSE, all flowing through one storage layer. NFS, S3 and WebDAV are compiled in behind build flags, so a binary you did not build may not include them.

### Does kiwifs need a separate frontend deployment?

No. The web interface, with a block editor, wiki links, backlinks, a knowledge graph and a dark mode, is embedded in the binary and ships with it. The compose file builds from the repository Dockerfile so the embedded interface stays in sync with the backend.

### Can kiwifs run semantic search without sending data to a hosted API?

Yes. You download a local sentence-transformer model with the model command, build with the embedder tag, and point the config at the model file and its dimension count. The hosted embedders include OpenAI, Ollama, Cohere and an HTTP option.

### How does kiwifs version the files it stores?

With git. Every write is an atomic commit, which gives blame, diff and point-in-time restore, and each multi-space workspace has its own git repository alongside its own search index.

### What can kiwifs import data from?

Nineteen importers, including Postgres, MySQL, MongoDB, Notion, CSV and Obsidian. Exports go the other way as JSONL or CSV, with optional embeddings, the link graph and content.

## Sources

- [Issues](https://github.com/kiwifs/kiwifs/issues)
- [kiwifs/kiwifs on GitHub](https://github.com/kiwifs/kiwifs)
- [Project website](https://docs.kiwifs.com)
- [README](https://github.com/kiwifs/kiwifs/blob/main/README.md)
- [Releases](https://github.com/kiwifs/kiwifs/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kiwifs-kiwifs
