Model or dataset
voicetreelab/voicetree avatar
voicetreelab/voicetree

Voicetree ships version 3.0.2 for the desktop and 0.1.0 for its Python package

The spatial IDE for recursive multi-agent orchestration. It's like an Obsidian graph-view that you work directly inside of.

922 stars57 forksTypeScriptNOASSERTION

At a glance

What is it?
A graph workspace where coding agents are terminals living beside markdown files, connected by wikilinks, with a shared context radius instead of a shared transcript. Two language stacks, two version numbers, and a container that mounts a volume over the agent's own config directory.
Who is it for?
Use Voicetree if your problem is that four to ten agent terminals on one screen are overwhelming, and you want the coordination to live in a structure both you and the agents can read. Three things to check first.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 119 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Three identities and two version numbers

Start with what is actually in the repository, because the naming does not line up. The desktop application is released as 3.0.0, 3.0.1 and 3.0.2 within the first week of June 2026, with the last push a day after the newest tag. The root manifest is named differently again, and marked private, and the Python project described in the build configuration is a third name entirely, at version 0.1.0, described as a voice to structured graph system. So a bug report that quotes a version number is ambiguous: 3.0.2 is the desktop build, 0.1.0 is the server side. The Python side also demands 3.13 or newer, while the TypeScript side runs under a pinned package manager version. Two language stacks, two release trains, one product name.

Nodes are markdown files and edges are wikilinks

The data model is deliberately boring, and that is the design. There are two primitives. A node is one of three things: a markdown file, a set of nodes in a folder, or a terminal running a coding agent, with four named as typical. Connections between nodes are wikilinks pointing at the markdown file paths, which means the graph is the filesystem, not a database beside it. That has two consequences worth stating. A human can read and edit the graph with any editor, and an agent editing a file is editing the graph, so there is no import step and no reconciliation. You open a full markdown editor by hovering over a node, and the project also describes a speech-to-graph mode for creating nodes by voice. The stated ambition is that the structure mirrors how you actually think about a problem, so that the canvas rather than your working memory holds the relationships.

Context is a radius around a node, not a transcript

The mechanism that decides what an agent sees is spatial. When you spawn an agent on a node, that node's contents become its task, and it is given everything within an adjustable distance, and it can search semantically against local embeddings. The claim made for this is that it targets retrieval at what is relevant instead of replaying a whole conversation, and the page quantifies the thing being avoided as a thirty to sixty percent performance drop from context rot, cited to a numbered footnote whose text does not appear on the visible page. Search runs twice over: the graph structure is navigable, with one representation for the agent and a graphical one for the person, and the project notes in parentheses that there is a user interface after all.

Subagents are terminals, which is the transparency argument

Decomposition is the reason the graph exists. An agent can break its own task into subgraphs of small connected pieces, and the example given is a plan split by layer: data model, architecture, pure logic, edge logic, interface components, and integration. You can then look at the high level and zoom into the part that matters, which is the point of tracking planning against implementation. From there an agent spawns and orchestrates parallel subagents across its own dependency graph. The differentiator claimed for this is that those subagents are ordinary native terminals rather than something hidden inside another tool, so you can watch them and intervene, and the page contrasts that with other command line agents. The default behaviour of the run operation is spelled out too: it collects the nearby nodes as context and messages them into the agent it just started.

The container mounts a volume over the agent's own config directory

The Docker route is a sandbox, and its volume list is the part to read. Two named volumes are mounted: one over the project directory, and one over the configuration directory for a specific coding agent. So the agent's credentials and settings live in a volume that outlives the container, which is convenient and also means the agent's configuration is not in the image. The environment template explains a consequence: no browser is installed in the container, so that agent's browser sign-in flow is awkward, and you are told to set an API key instead. Other agents' keys are optional and commented out, for the tools an install script can add. The service publishes one port for a browser-based desktop session, allocates a gigabyte of shared memory, and the documentation notes the image is amd64 only for now. The published invocation is short enough to read as a specification of what the image expects:

bash
docker run -d --rm -p 6080:6080 \
    -v voicetree-project:/home/vt/project \
    -v voicetree-claude:/home/vt/.config/claude \
    --shm-size=1g \
    -e ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY" \
    ghcr.io/voicetreelab/voicetree:latest
# then open http://localhost:6080/vnc.html?autoconnect=1&resize=remote

A single published port, a project volume, a volume over one agent's configuration, a shared memory setting and one key passed through from your shell. The compose file next to it declares the same two volumes, the same port, the shared memory size and an environment file, plus a restart policy, and it builds from a Dockerfile in a subdirectory. The last line of the command is the part people miss: the interface is reached through a browser desktop session served from inside the container, not through a web app on that port.

The test harness runs remotely and in tiers

The root manifest is mostly test and measurement plumbing, and it says something about how this is developed. Tests are not invoked directly: a script runs the command remotely and delegates to a local variant, and there are numbered tiers, with a tier-zero run that captures checks up to a threshold and a tier-one run that records a measurement with an identifier, a name and a category before delegating again. The type check filters to the web application package alone. Alongside the tests sit a second vitest configuration for fuzzing, and a set of static gates you rarely see in one repository: an unused-code detector, a duplicate-code detector, a secret scanner, a code scanning configuration, a coverage configuration, and two ignore files for different tools. There is also a spec directory and an evaluations directory at the root, which suggests the project is measured as well as tested.

The Python side keeps a test framework in its runtime dependencies

The Python project is where the packaging habits show. Its runtime dependency list includes a test framework, a requests library, a colour library and the build backend itself, alongside the expected machine learning stack: a pinned sentence-transformers release, a vector database, an inference runtime, a fuzzy matching library, a ranking library, a language model framework and its core, an ASGI server with its standard extras, websockets and a validation library. Two different Google client packages appear side by side, along with an older general-purpose Google package, so the model access layer is not unified. The type checker is configured at close to maximum strictness, and two comments above the settings are worth quoting in spirit: one says there are no exceptions and all code follows the strict rules, the other says no backwards compatibility and a single solution is enforced.

The README states a prediction about 2027 as a claim with a probability

Two passages set the tone and both are unusual for a project README. One is a self-describing claim with a probability attached, stating that markdown hypergraphs have become the de facto programming language for agent cognition swarms in 2027, given a figure of zero point four. Written that way, it is a prediction with a stated confidence rather than a claim about the present, and it is the sort of line that gets quoted out of context. The other is an honest status note between horizontal rules: the project is early beta, powerful but rough, built by a lab that uses it daily as a research and forecasting tool, with development described as spiky and outside contributors explicitly welcome. The section that follows, on how it works in detail, ends partway through a sentence after the point about subagents being native terminals.

Editorial conclusion

Use Voicetree if your problem is that four to ten agent terminals on one screen are overwhelming, and you want the coordination to live in a structure both you and the agents can read. Three things to check first. Decide which surface you are running, because the desktop build, the package manifest and the Python project carry three different identities and two version numbers, and only the desktop one is versioned in releases. Read what the container mounts before you use it, since a named volume sits over the agent's own configuration directory and the OAuth flow is replaced by an API key. And expect roughness: the project calls itself early beta and says its development is spiky.

Frequently asked questions

What is a node in Voicetree?

One of three things: a markdown file, a set of nodes in a folder, or a terminal running a coding agent. Connections between nodes are wikilinks pointing at the markdown file paths, so the graph is the filesystem.

How does Voicetree decide what an agent can see?

Spatially. An agent spawned on a node receives that node's contents as its task plus everything within a configurable radius, and can also search semantically against local embeddings.

What does the Voicetree container mount?

Two named volumes: one over the project directory and one over the configuration directory for a specific coding agent, so the agent's settings persist outside the image. It also passes one provider key through and allocates a gigabyte of shared memory.

What does the Voicetree container need instead of a browser sign-in?

A provider API key. The environment template notes that no browser is preinstalled in the container, so the agent's browser-based sign-in is awkward, and sets the key variable as the required value.

Which versions does Voicetree publish?

Desktop releases 3.0.0, 3.0.1 and 3.0.2 from the first week of June 2026, while the Python project in the build configuration is at version 0.1.0. The root manifest is a third name, marked private, with its own package manager pinned.

How does Voicetree run its tests?

Through a remote runner script that delegates to a local variant, split into numbered tiers. Alongside sit a fuzzing configuration and static gates including an unused-code detector, a duplicate-code detector, a secret scanner, a code scanning configuration and coverage settings.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. voicetreelab/voicetree on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/voicetreelab-voicetree.svg)](https://hysenlabs.com/projects/voicetreelab-voicetree)