# CyberVerse: a self-hosted digital-human agent platform built around WebRTC

> CyberVerse is a Python framework for voice-first AI agents with persona memory, RAG, tool use and optional talking-head video. It is a heavy install, and the README says very little about what happens when a GPU node dies.

**Lynpoint/CyberVerse** — Self hosted, real-time digital human agent platform. Build voice-first AI agents with WebRTC, persona memory, tools, RAG, and optional digital-human video.

- Repository: https://github.com/Lynpoint/CyberVerse
- Website: https://www.cyberverse.cc
- Stars: 1,685 · Forks: 228
- Language: Python
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/lynpoint-cyberverse

## The problem CyberVerse solves: a voice agent that also has a face

Most agent frameworks treat voice as a text pipeline with a speech layer bolted on. CyberVerse starts from the opposite end. The README describes it as a real-time digital-human Agent framework that uses WebRTC, persona memory, tools, RAG and optional digital-human video to build agents centered on voice interaction. The target user is someone who wants a persistent character, not a stateless assistant: the README's framing is a J.A.R.V.I.S.-style companion, or a character brought to life from a single photo.

That framing matters because it determines the architecture. A character has to remember previous conversations, so history is persisted to local disk and reloaded when you re-enter a conversation. A character has to answer in its own voice, so the stack includes replaceable TTS and ASR modules. And if you want the video-call experience the README advertises, the agent needs a face that lip-syncs in real time. CyberVerse is for teams and individuals who want all three in one self-hosted system rather than stitching a voice SDK, a vector store and an avatar service together themselves.

It is not for someone who wants a hosted API. There is no managed tier described in the README; you run the server, the inference process and the frontend.

## PersonaAgent in the foreground, SubAgents in the background

The multi-agent split is the most concrete design decision in the README. PersonaAgent stays in the foreground. Its job is to keep conversation fluid, respond quickly to interruptions, and handle context switches. Long-running work (search, research, material organization, summarization, HTML report generation) is delegated to background SubAgents asynchronously.

The stated reason is latency. If a research task ran inline, the voice turn would stall while the model worked. By pushing that work to a SubAgent, the user can keep speaking, ask follow-up questions, or change direction, and PersonaAgent returns the result when it is ready. This is a queue-and-notify pattern, not a parallel one: the foreground agent does not block on the background agent, and the README does not describe what happens if a SubAgent never finishes or how results are ordered if two complete at once.

Memory and retrieval sit alongside this. Conversation history is written to local disk per character and loaded on re-entry, which is what makes continuity across sessions possible. Separately, you can import knowledge bases, documents and biographical material for a character; the system indexes them for retrieval-augmented generation so answers align with the character's background. The two stores are distinct: memory is what was said, RAG is what the character is supposed to know.

## Installing CyberVerse and running a first voice session

The repository is a Python project with a Go server component and a Node frontend, and the Makefile is the entry point. The setup target installs the Python package with the dev and inference extras, generates protobuf stubs for both languages, then installs frontend dependencies. Protobuf generation runs after the pip install on purpose, because grpc_tools has to be present first.

```bash
make setup
```

If you only want the agent side, there is a narrower target that installs the agent extra:

```bash
make setup-agent
```

The Makefile looks for Go 1.25 in three places: /usr/lib/go-1.25/bin/go, a user SDK install under $HOME/sdk/go1.25.9/bin/go, then PATH. It looks for Node 22 or newer through nvm under $HOME/.nvm/versions/node. The media stack also expects C libraries such as opus and soxr from a conda environment, defaulting to $HOME/miniconda3/envs/cyberverse. If your toolchain lives elsewhere, override GO, NODE_BIN or CONDA_ENV rather than editing the file.

For development, the inference server is started through a script that reads avatar runtime GPU settings from config/cyberverse.yaml and auto-selects between plain python and torchrun:

```bash
make inference
```

The Makefile documents ad-hoc overrides for that target, for example WORLD_SIZE=2 CUDA_VISIBLE_DEVICES=0,1 make inference. Tests are split by language: make test-py runs python -m pytest tests/unit, and make test-go runs go test inside server. The integration marker is separate and requires CUDA plus local checkpoints.

Runtime behaviour is configured in config/cyberverse.yaml. Provider definitions for omni, LLM, TTS, ASR and embedding live in the built-in infra/config/*_models/ directories, with optional local overrides under config/*_models/. API keys and service endpoints are set in the web UI at /settings, which is where you switch providers and model combinations per scenario. The README does not give a copy-pasteable YAML sample, so expect to read the built-in provider files before you can point CyberVerse at your own endpoints.

## The GPU bill is the real adoption constraint

The README publishes a table of digital-human models with hardware requirements, and it is the most useful page in the project. FlashHead 1.3B at Pro quality needs two RTX 5090 cards for 512x512 at 25+ FPS, or one RTX 5090 for 464x464 at 20 FPS. LiveAct 18B needs two RTX PRO 6000 cards for 320x480 at 20 FPS, or one for 256x417 at 20 FPS. Those are the local options. The cloud options (Vidu S1, Baidu Xiling, Xunfei Digital Human) are marked as requiring no local GPU, with resolution and frame rate determined by the provider.

So the honest reading is that CyberVerse has two deployment profiles with very different costs. If you use a cloud digital-human provider, the local machine only has to run the agent, the server and the frontend. If you want the local models, you are buying datacenter-class GPUs, and the single-card configurations trade resolution for it. The README does not state memory requirements, quantization options, or how many concurrent sessions one card can serve.

There is a second limitation that the README leaves open: failover. Nothing in the README describes what happens when the inference process crashes mid-conversation, whether a session can resume on another GPU, or how the server behaves when a cloud provider returns an error. For a real-time video agent, that is the failure mode users notice first. Treat the absence as a gap to investigate in tests/ before production.

## How CyberVerse differs from a LiveKit-style voice agent stack

LiveKit Agents is the closest comparison, and the difference is in where the abstraction sits. A LiveKit-style stack gives you the real-time transport, room model and agent lifecycle, then leaves the model layer to you: you wire up STT, LLM and TTS plugins and manage the pipeline yourself. CyberVerse bundles the whole vertical. Brain, voice, hearing, tools, memory and face are all replaceable modules, but they ship as one configured system with a web UI for provider selection.

The practical consequence is that CyberVerse is opinionated about the agent loop. PersonaAgent and SubAgent are framework concepts, not patterns you implement. Character memory persisted to local disk is a framework behaviour. If you want to design a different memory strategy, you are working against the grain rather than inside a plugin interface.

The LiteLLM plugin is the escape hatch on the model side. The README states it adds access to 100+ LLM providers (AWS Bedrock, Azure, Vertex AI, Mistral, Cohere and others) through a single unified interface, and pyproject.toml pins it as litellm>=1.80,<1.87. That is a narrow version window, so an upstream LiteLLM release can leave CyberVerse behind until the pin moves.

If your requirement is text-only agents with no avatar, CyberVerse is the wrong tool. You would be paying the WebRTC, media and GPU complexity for a feature you never enable.

## Licence, maintenance and upgrade cost

CyberVerse is GPL-3.0. For self-hosted internal use that is usually unremarkable. If you intend to embed it in a product you distribute, the copyleft terms apply to the combined work, and that is a decision for your own legal review rather than something the repository resolves. The README also notes that the demo characters are examples only, are not bundled with CyberVerse, and are not provided for commercial use, so the persona assets are a separate question from the code licence.

The project is not archived, and the last push was on 2026-08-05. The only release listed is v0.1.0 from 2026-05-16, so the version number in pyproject.toml (0.1.0) matches the tag. Expect the interface to move: a pre-1.0 project with a config file, a web settings page and a provider directory layout has several surfaces that can change between releases.

Upgrade cost is dominated by the dependency groups rather than the core package. The base install is three dependencies (numpy, pydantic, pyyaml). The optional groups are where the weight is: flash_head pins transformers==4.57.3, gradio==5.50.0, xfuser==0.4.5, xformers==0.0.32.post1 and nvidia-nccl-cu12==2.27.3; live_act pins xfuser==0.4.5 as well. Pinned CUDA-adjacent packages like these are the ones that break when a driver or a base image changes, and they are the reason the README offers prebuilt cloud images on Compshare and AutoDL for people who would rather not assemble the environment by hand. The README does not document a rollback procedure for a failed upgrade, so keep the previous environment until a new one is verified.

## Conclusion

Adopt CyberVerse if you need a self-hosted, voice-first agent with a talking face and you already own the GPUs, since the local FlashHead and LiveAct paths assume an RTX 5090 or RTX PRO 6000 and the cloud digital-human providers remove that requirement. Do not adopt it if you only need a text chatbot, or if GPL-3.0 obligations are incompatible with how you ship. Before committing, verify which of the optional dependency groups (rag, milvus, flash_head, live_act) your hardware can actually support, and check whether config/cyberverse.yaml exposes the provider and GPU settings you need, because the README does not document rollback or failover behaviour.

## FAQ

### What is CyberVerse?

CyberVerse is an open-source real-time digital-human Agent framework. It uses WebRTC, persona memory, tools, RAG and optional digital-human video to build AI agents centered on voice interaction, and it is self-hosted rather than offered as a managed service.

### How do I install CyberVerse?

The repository provides a Makefile. Running make setup installs the Python package with the dev and inference extras, generates the protobuf stubs, and installs the frontend dependencies; make setup-agent installs only the agent extra.

### Does CyberVerse need a local GPU?

Only for the local digital-human models. The README's table lists FlashHead 1.3B at two RTX 5090 cards for 512x512 at 25+ FPS, and LiveAct 18B at two RTX PRO 6000 cards for 320x480 at 20 FPS. The cloud options (Vidu S1, Baidu Xiling, Xunfei Digital Human) are listed as requiring no local GPU.

### What licence does CyberVerse use?

GPL-3.0, as stated in the README badge and the LICENSE file. The README also notes that the demo characters are examples only, are not bundled with CyberVerse, and are not provided for commercial use.

### Where do I configure the LLM and TTS providers in CyberVerse?

Runtime behaviour is set in config/cyberverse.yaml, while omni, LLM, TTS, ASR and embedding provider definitions are loaded from infra/config/*_models/ with optional local overrides under config/*_models/. API keys and service endpoints are configured in the web UI at /settings.

## Sources

- [License: GPL-3.0](https://github.com/Lynpoint/CyberVerse/blob/main/LICENSE)
- [Lynpoint/CyberVerse on GitHub](https://github.com/Lynpoint/CyberVerse)
- [Project website](https://www.cyberverse.cc)
- [README](https://github.com/Lynpoint/CyberVerse/blob/main/README.md)
- [Releases](https://github.com/Lynpoint/CyberVerse/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lynpoint-cyberverse
