# VoiceMem: Streaming Dual-Brain Memory Layer for Real-Time Voice Agents

> A Python library that adds persistent memory to voice assistants using a left-brain fact store and a right-brain emotional graph, with speculative prefetching that keeps retrieval latency below 300 milliseconds during live conversation.

**xzf-thu/VoiceMem** — a real-time and empathetic memory system for voice assistants

- Repository: https://github.com/xzf-thu/VoiceMem
- Stars: 2,277 · Forks: 173
- Language: Python
- License: Apache-2.0
- Published: 2026-08-27 · Updated: 2026-08-27 · Language: en
- Canonical page: https://hysenlabs.com/projects/xzf-thu-voicemem

## The Memory Problem in Live Voice Agents

Existing LLM memory systems were designed for text-based chatbots. They assume a synchronous request-response loop: the user sends a message, the system searches memory, and the model generates a reply. Voice agents break that assumption. A user starts speaking before the sentence is finished, the system must begin retrieving relevant memories mid-utterance, and the reply must arrive with latency measured in milliseconds, not seconds.

VoiceMem addresses this with a speculative prefetch: as soon as the transcribed text reaches a minimum length (six characters in the streaming example), the system begins a background memory search. By the time the user finishes speaking, the retrieval result is already available. The README states a 134 ms response time for VoiceMem, compared to 1,440 ms for Mem0 in the same benchmark.

Beyond speed, the project addresses a second gap: standard memory stores record what users say, but not their emotional state or personality. VoiceMem adds a right-brain component specifically for emotional attribution and interpersonal context.

## Dual-Brain Streaming Architecture

The architecture divides memory into two cooperating stores. The left brain manages factual memory using schema and entity structures. It stores information like dietary restrictions, preferences, and biographical facts, and retrieves them through vector search with a Top-K limit. The README reports 91.2% accuracy on the LoCoMo benchmark at Top-5, compared to 61.68% for Mem0 at the same setting.

The right brain tracks emotional and interpersonal context. It maintains independent nodes for each entity and cross-entity nodes that link relationships and emotional history. The README describes this as managing personality, emotion, and relationship information, handled separately from factual recall so that emotional data does not pollute the precision of fact retrieval.

The full pipeline is streaming. While a user is speaking, VoiceMem runs audio segmentation, speech transcription via funasr, speaker identification via sherpa-onnx, scene detection, and emotion recognition in parallel. Structured information is written to the memory graph incrementally. Queries route through the two brains, rank results, and inject only the Top-K entries into the model context, keeping context length bounded. A single query uses approximately 430 memory tokens, compared to 6,956 for Mem0 and 1,899 for EverMemOS, according to the README benchmark table.

## Installing VoiceMem and Running a First Memory Store

The minimum Python version is 3.10, as specified in pyproject.toml. Install the full package including all built-in audio components:

```bash
git clone https://github.com/xzf-thu/VoiceMem.git
cd VoiceMem
pip install voicemem
```

To use the fine-tuned Qwen reply model:

```bash
pip install "voicemem[slm]"
```

Download the required models from Hugging Face:

```bash
pip install -U huggingface_hub
hf download zhifeixie/VoiceMem_Default_Models_Env --local-dir ./models
```

A minimal offline usage example stores a text fact and queries it back:

```python
from voicemem import VoiceMem

vm = VoiceMem(
    mode="normal",
    openai_key="api_xxx",
    top_k=5,
)
vm.warmup()
vm.ingest(audio="assets/input.wav")
result = vm.search("my dietary restrictions")
print(result.result_leftbrain, result.result_rightbrain)
```

The warmup call loads local models into memory before the first request, avoiding latency on the first ingest or search call. The README notes that writing is slower than searching because ingest extracts facts, assigns tags, and builds the graph, while search uses only vector retrieval.

To launch the interactive web demo:

```bash
python web/run.py
```

The demo runs on port 8787 by default and logs output with timestamps to the results/logs folder.

## Integrating VoiceMem into a Custom Voice Agent

The intended integration pattern is simple: connect VoiceMem's listening loop between the microphone and your model's generation step. The README describes the data flow as: microphone input, VoiceMem listens and pre-fetches relevant memories, your model reads those memories and generates a reply.

The required external API key is an OpenAI key, used only for fact extraction during ingest. Memory retrieval runs entirely locally. To substitute your own generation model:

```python
def my_reply(text, memory_context):
    return my_model.generate(system=memory_context, user=text)

vm = VoiceMem(reply=my_reply)
```

A synchronous function is also accepted; VoiceMem wraps it in a thread automatically. Barge-in handling (interrupting a reply mid-playback) uses a two-stage control mechanism: VAD pauses and buffers audio, and a stable ASR confirmation or explicit stop command clears the queue. The thresholds BARGE_REJECT_SILENCE_MS and BARGE_CANDIDATE_TIMEOUT_MS are configurable.

For fine-tuning the VoiceMem model family adapter, the project provides training code:

```bash
pip install ms-swift==4.5.2 bitsandbytes
python finetune/train.py --data data/train.jsonl
```

The training configuration matches the published checkpoint-3318, according to the README.

## What VoiceMem Does Not Handle

VoiceMem depends on a chain of large local models: funasr for ASR, sherpa-onnx for VAD and speaker identification, transformers for emotion recognition, and sentence-transformers for embedding. The pyproject.toml notes that omitting torchvision causes the emotion attribution module to silently fall back to acoustic-only classification, producing incorrect emotional labels without raising an error. Similarly, omitting accelerate causes the emotion model to silently degrade rather than failing loudly.

The package requires an OpenAI API key for the fact extraction step during ingest. Purely offline operation is not possible if you use the audio ingest path with fact extraction. The README mentions an all-local example in examples/04_all_local_l40s.py, but does not describe its limitations in the main documentation.

The web demo and examples are not included in the pip package; they require cloning the repository. Version v0.0.2, released on 2026-09-01, fixed event date linking and removed redundant right-brain memory categories. The current pip package version is 0.2.3 as listed in pyproject.toml.

## How VoiceMem Differs from Mem0

Mem0 is a text-centric memory layer for LLM applications that stores and retrieves conversational facts through an API. It supports both hosted and self-hosted deployment and works with any language model via an OpenAI-compatible interface. The architecture is a single store with vector and graph retrieval options.

VoiceMem takes a different position: it is built specifically for voice input, processes audio natively through its ASR and speaker identification pipeline, adds an emotional attribution layer, and targets streaming retrieval during speech rather than retrieval after a turn ends. The README comparison puts VoiceMem's accuracy at 91.2% on LoCoMo versus Mem0's 61.68%, and its latency at 134 ms versus 1,440 ms, though both figures come from the project's own evaluation code rather than an independent benchmark.

For text-only agent memory where latency is not a hard constraint, Mem0 is a simpler dependency. For voice agents that need to respond quickly and track who said what with what emotional tone, VoiceMem is the focused choice.

## Releases, License, and Maintenance

The project released v0.0.1 and v0.0.2 on 2026-09-01. The last push to the main branch was on 2026-09-14. The license is Apache-2.0, which permits commercial use, modification, and distribution.

The associated ChatMem-400K dataset and the VoiceMem model family (fine-tuned versions of Qwen2.5-Omni, Qwen3-Omni, and Step-Audio2-Mini) are published on Hugging Face separately. The evaluation code is included in the repository under evaluation/ and is described as fully reproducible, with a small sample dataset for environment verification before running the full benchmark.

## Conclusion

VoiceMem is the right choice for developers building real-time voice agents in Python who need both factual recall and emotional context from past conversations, and who are comfortable with the dependencies it loads: torch, torchaudio, torchvision, funasr, modelscope, sherpa-onnx, and sentence-transformers are all pulled in by a single pip install. Teams that only need text-based memory without voice, or who are running on resource-constrained hardware, should use Mem0 directly instead. The Apache-2.0 license permits commercial use. Check the pyproject.toml for exact Python version requirements before starting, since the dependency versions are specific.

## FAQ

### What Python version and dependencies does VoiceMem require?

VoiceMem requires Python 3.10 or later. A single pip install voicemem pulls in torch, torchaudio, torchvision, transformers, funasr, modelscope, sherpa-onnx, sentence-transformers, and the web demo dependencies. Omitting torchvision causes the emotion module to silently degrade rather than fail.

### How does VoiceMem's dual-brain architecture differ from a single vector database?

The left brain handles factual memory through schema and entity structures with vector retrieval. The right brain handles emotional, personality, and relationship context through independent and cross-entity nodes. Separating the two prevents emotional data from reducing the precision of factual recall during retrieval.

### Can VoiceMem be used without an external OpenAI API key?

The OpenAI API key is used only for fact extraction during audio ingest. The repository includes an all-local example at examples/04_all_local_l40s.py, though the README does not describe its limitations in detail. Memory retrieval itself runs entirely locally.

## Sources

- [Official README](https://github.com/xzf-thu/VoiceMem#readme)
- [Project repository](https://github.com/xzf-thu/VoiceMem)
- [Release notes](https://github.com/xzf-thu/VoiceMem/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xzf-thu-voicemem
