Model or dataset
YeJe-cpu/talk-to-fengge avatar
YeJe-cpu/talk-to-fengge

talk-to-fengge: Real-Time Voice Conversation with Voice Cloning and Persona Injection

Talk to 峰哥 — 克隆任何人的声音和性格,实时语音对话,工程延迟 < 1 秒 | Clone anyone's voice & personality for real-time conversation. < 1s engineering latency.

468 stars95 forksPythonApache-2.0

At a glance

What is it?
talk-to-fengge is a Python voice agent that combines voice cloning, persona injection, and real-time WebRTC conversation under one second of engineering latency, built on LiveKit Agents with swappable STT, LLM, and TTS providers.
Who is it for?
talk-to-fengge is a good fit for developers who want to build a persona-driven real-time voice agent and are comfortable assembling three separate API accounts plus a GPU TTS service. It is the wrong starting point for a team that needs a hosted, single-credential voice solution without GPU infrastructure.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 83 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What talk-to-fengge Combines That Other Voice Tools Separate

The README observes that most voice-capable tools fall into one of two categories: tools that support real-time conversation but use only the model's default voice, and tools that clone a specific voice but generate audio as a batch process rather than in a live call. talk-to-fengge puts three components together in a single pipeline: voice cloning from 15 to 45 seconds of audio, persona injection that captures speech patterns and characteristic phrases alongside the voice, and real-time conversation delivered through a WebRTC channel.

The default persona built into the repository is 峰哥, a Chinese-language streaming personality with a large online following. The architecture is designed to swap that persona for any other, given suitable voice samples and a persona description. A developer building a custom AI companion, a business building a brand voice assistant, or a researcher studying persona-driven dialogue systems would reach for this repository when they need all three components wired together rather than integrated separately.

The Three-Tier Stack: STT, LLM, and TTS

The README diagrams the data flow as a linear chain: speech enters STT, the recognized text goes to the LLM with persona instructions injected, and the LLM output goes to TTS which generates audio in the cloned voice. An optional memory recall step injects relevant past context into the LLM call. Each tier has a recommended provider and documented alternatives. For speech recognition, the README recommends Cartesia ink-whisper, a model available on the free tier, and lists Deepgram nova-2 as an alternative. For the language model, it recommends MiniMax-M2.7-highspeed, described as having very low time-to-first-byte and usable without a VPN from China. DeepSeek and Gemini are listed as alternatives. For voice synthesis and cloning, it recommends VoxCPM, an open-source model from OpenBMB that requires a GPU. MOSS-TTS, described as CPU-runnable at lower quality, and Cartesia Sonic, a cloud service requiring a paid subscription, are listed as fallback options.

Switching providers requires a change to .env.local. The default configuration sets STT_PROVIDER=cartesia, LLM_PROVIDER=minimax, and TTS_PROVIDER=voxcpm. Each provider uses its own environment variables, documented in .env.example.

Setting Up the Pipeline

The README offers two paths. For the simplest start, clone the repository and hand it to an AI coding agent:

bash
git clone https://github.com/YeJe-cpu/talk-to-fengge.git
cd talk-to-fengge

The agent reads .env.example and prompts for API keys. For manual setup, the dependencies are Python 3.12 or higher, uv for package management, and a locally running LiveKit server. On macOS:

bash
uv sync
brew install livekit
cp .env.example .env.local

Then edit .env.local to fill in at minimum a Cartesia API key, a MiniMax API key, and the TTS provider settings. The VoxCPM service is a separate GPU process. On a RunPod L4 or any NVIDIA machine with at least 8GB of VRAM:

bash
bash runpod_setup.sh

To start the agent manually, three terminal processes run in parallel:

bash
livekit-server --dev --node-ip=127.0.0.1
LLM_PROVIDER=minimax python -m worker.main start
python -m worker.web_server

Opening http://127.0.0.1:8766 in a browser starts the conversation. On macOS a .command launcher file is also provided for a double-click start.

GPU Requirements and the TTS Provider Trade-Off

VoxCPM is the recommended TTS provider because the README describes it as having the best voice cloning quality among the available options. It requires an NVIDIA GPU with at least 8GB of VRAM. The runpod_setup.sh script installs dependencies and downloads the model, which the README notes takes about ten minutes on first run.

For situations where a GPU is not available, MOSS-TTS runs on CPU. The README describes its clone quality and speed as inferior to VoxCPM. Setting TTS_PROVIDER=moss in .env.local switches to it. Cartesia Sonic is a third option: a cloud TTS service that requires a Pro subscription at five dollars per month before voice cloning is available. The README sets TTS_PROVIDER=cartesia for that path.

The trade-off is clear: VoxCPM gives the best voice output at the cost of GPU infrastructure; MOSS-TTS removes the GPU requirement at a quality cost; Cartesia Sonic removes the GPU requirement at a recurring monetary cost.

Practical Latency and Known Limitations

The README states the engineering latency target as under one second, measuring the pipeline from audio input through STT, LLM, and TTS to audio output. The actual conversational latency the user experiences is described as approximately two to three seconds, affected by network conditions and API response times.

The README documents several known limitations directly. The setup requires multiple API accounts and services; there is no single-credential or one-click deployment path at the time of the last push on 2026-07-11. The memory system using OpenViking is limited: because the 峰哥 persona speaks more than the user, the system produces fewer user-side memory anchors to persist. The frontend is a single HTML page with a particle animation effect and no virtual avatar. The project lists animated avatar integration and automated persona generation as planned future work, not yet implemented.

Compared to GPT-4o Voice Mode and Replacing the Persona

The README specifically names GPT-4o Voice as a category example of real-time voice conversation tools that do not support custom voice cloning. GPT-4o Voice Mode is a product by OpenAI that uses the model's own built-in voice. A developer who wants the model's voice without infrastructure overhead uses GPT-4o Voice Mode. A developer who needs a specific cloned voice tied to a specific persona requires a pipeline like talk-to-fengge.

To replace the default 峰哥 persona, the README asks for two things: 15 to 45 seconds of clean voice audio with no background music, and a persona description covering speech style and characteristic phrases. The persona configuration lives in worker/persona.py, with documentation in docs/persona-*.md. The README suggests asking an AI coding agent to generate the new persona configuration by pointing it at the existing 峰哥 implementation as a template. Richer source material such as transcripts or social media content is described as producing better results.

Editorial conclusion

talk-to-fengge is a good fit for developers who want to build a persona-driven real-time voice agent and are comfortable assembling three separate API accounts plus a GPU TTS service. It is the wrong starting point for a team that needs a hosted, single-credential voice solution without GPU infrastructure. The multi-service setup, VoxCPM on an NVIDIA GPU with at least 8GB VRAM plus Cartesia and MiniMax API keys, is the main friction point. Before running it, verify that the required API keys are available and that VoxCPM can reach the VOXCPM_URL set in .env.local.

Frequently asked questions

What API keys does talk-to-fengge require to run?

At minimum the README requires a Cartesia API key for speech recognition and a MiniMax API key for the language model. The TTS provider adds a third dependency: VoxCPM is self-hosted on a GPU, MOSS-TTS requires no external key but runs on your own machine, and Cartesia Sonic requires a Cartesia Pro subscription at five dollars per month for voice cloning.

Can talk-to-fengge run without a GPU?

Yes, but with trade-offs. Switching to MOSS-TTS by setting TTS_PROVIDER=moss in .env.local removes the GPU requirement. The README describes MOSS-TTS as inferior to VoxCPM in both clone quality and speed. Cartesia Sonic is another cloud-based GPU-free option requiring a paid subscription.

What is the difference between engineering latency and practical latency in talk-to-fengge?

The README defines engineering latency as the pipeline's own processing time, which it targets at under one second. Practical conversational latency is described as approximately two to three seconds in real use, because network conditions and the response times of external APIs add to the pipeline's own time.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Project website
  4. README
  5. YeJe-cpu/talk-to-fengge on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yeje-cpu-talk-to-fengge.svg)](https://hysenlabs.com/projects/yeje-cpu-talk-to-fengge)