# talk-to-fengge: a voice clone with a personality file attached

> Three capabilities that usually ship separately are wired into one Python stack here, speech recognition, a language model, and cloned speech synthesis, with persona injection sitting beside the model and a persona file as the only piece you edit to change who you are talking to.

**YeJe-cpu/talk-to-fengge** — Talk to 峰哥 — 克隆任何人的声音和性格，实时语音对话，工程延迟 < 1 秒 | Clone anyone's voice & personality for real-time conversation. < 1s engineering latency.

- Repository: https://github.com/YeJe-cpu/talk-to-fengge
- Website: https://x.com/leaf_sanren/status/2069342335268507976
- Stars: 468 · Forks: 95
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/yeje-cpu-talk-to-fengge

## Three capabilities that usually ship apart

The argument in the README is about a gap in the market rather than a feature list. Plenty of projects clone voices and plenty of projects do real-time voice conversation, but they are usually disjoint: the tools that can hold a live conversation, the project names GPT-4o Voice as an example, do not support custom voice cloning, and the tools that clone voices, Bark and XTTS among them, are text-to-speech systems that cannot converse. This project joins three things. Voice cloning from 15 to 45 seconds of speech. Persona injection, described as speech style, catchphrases and way of thinking rather than mere vocal similarity. And real-time conversation, framed as a phone call rather than a read-out, with the engineering chain under a second. The built-in subject is a specific streaming creator, and the page is explicit that this person is only the first complete example of an architecture intended to be pointed at someone else.

## STT, LLM, TTS, with personality beside the model

The pipeline is drawn as a single line: the user speaks, speech recognition turns it into text, a language model answers, cloned speech synthesis says the answer, and the user hears it. Persona injection and memory recall sit beside the model rather than in the middle of the chain, which is the architecturally interesting choice, since personality is prompt context and memory is retrieval, not audio processing. Each stage has a recommended default. LiveKit handles the real-time audio and video, a WebRTC framework carrying the stream between the browser and the agent. Cartesia ink-whisper does recognition, recommended because the free tier is usable and it handles Chinese with low latency. MiniMax-M2.7-highspeed is the recommended model, chosen for a very low time to first byte and for not needing a VPN. VoxCPM does synthesis and cloning, described as open source with the best cloning results, and it needs a GPU, cloud or local. OpenViking memory is the optional stage.

## Four options per stage, switched by one line

Alternatives are configured rather than coded, with a single line in .env.local deciding the provider. Recognition is Cartesia ink-whisper as the recommendation, Deepgram nova-2, or Gemini. The model is MiniMax as the recommendation, DeepSeek, or Gemini. Synthesis has the widest menu: VoxCPM as the open-source recommendation, MOSS-TTS which runs on CPU as the fallback, Cartesia Sonic in the cloud, and MiniMax TTS. Each carries a cost, and the costs are different in kind. VoxCPM wants an NVIDIA GPU with at least 8 GB of VRAM, and runpod_setup.sh installs the dependencies and downloads the model in roughly ten minutes the first time. Cartesia Sonic needs a Pro subscription at $5 a month before it will clone a voice at all. MOSS-TTS needs nothing but gives up cloning quality and speed. One inconsistency is worth catching early: the prose recommends the highspeed model while .env.example pins MINIMAX_MODEL_NAME to a different MiniMax model name, along with a 300 token cap.

## Three terminals, or one double-clicked script

Running it is the least elegant part of the project. On macOS there is a launcher you double-click:

```bash
./Talk-to-Me-V3.6.command
```

Otherwise three processes go in three terminals:

```bash
livekit-server --dev --node-ip=127.0.0.1
LLM_PROVIDER=minimax python -m worker.main start
python -m worker.web_server
```

The web interface then appears at http://127.0.0.1:8766. Prerequisites are Python 3.12 or newer, uv or pip, and a LiveKit Server, which on macOS installs with brew install livekit. Note the version mismatch with the packaging metadata: the prose says 3.12 and newer while pyproject.toml declares requires-python of at least 3.11 and below 3.15. The LiveKit development server also defaults to devkey and secret in .env.example, which is fine locally and worth changing the moment it is not local. The distribution is also still named deepseek-talk-to-me rather than talk-to-fengge, and uv is configured with package = false, so this is an application you run rather than a library you depend on.

## The persona is a prompt plus a distillation method

Personality is configured, not trained. .env.example carries a persona name and an instruction string that tells the agent who it is and how to talk, with a comment noting that only the shipped persona is currently supported. Behind that prompt sits a method: the persona was distilled with the Nuwa Skill as the underlying methodology and the feng-ge-skill as its output, supplemented with material from livestreams. Swapping it takes two things. Voice material first: 15 to 45 seconds of clear speech with no background music and no noise, which is what VoxCPM clones from. Then a persona description, and the suggested route is to open the repository in a coding assistant and point it at the existing files, the persona notes under docs/ and worker/persona.py, as the template, after which it extracts the defining features through conversation and writes new persona code. The README is candid that richer material helps, listing chat logs, transcribed speech, social media posts and livestream clips.

## The reference audio is committed to the repository

The shipped subject's voice is already in the tree, at assets/voice_samples/fengge_ref.wav, described as compiled from the creator's public livestream content and usable as soon as the VoxCPM service is configured. That convenience is also the ethical question in the project, because it models cloning a real identifiable person from material they broadcast rather than from material they licensed. Anyone using this to clone somebody else should have that person's agreement first, and the technical bar is low enough that the absence of consent is the only real safeguard left. The honest framing is that a voice and a persona are the two halves of the feature, and the second half is cheap to write and expensive to get wrong. The project itself offers no guidance on permission, which is the gap a reader has to fill.

## Four limits, and one of them is about who talks

The known-limitations section is unusually specific. Deployment has many moving parts, because STT, LLM and TTS each need a different key or service and there is no one-click deployment yet. Latency splits in two: the engineering chain is under a second while the felt conversation latency is about two to three seconds, depending on the network and how fast the APIs respond. Memory is limited for a reason that is easy to miss. OpenViking mainly recognises events and entities on the user's side to accumulate memory, and since the AI does most of the talking in this scenario, there is little to accumulate, so the memory system will look broken in exactly the demo it is meant to impress. The frontend is a starfield particle page with no digital human or avatar. Planned next: an animated avatar, one-click deployment with automatic persona distillation from scraped material, and more persona templates. The code is Apache-2.0 and built on LiveKit Agents.

## Conclusion

talk-to-fengge fits someone building a voice agent who needs a cloned voice rather than a stock one, since the TTS stage is swappable between a local GPU model, a CPU fallback and two cloud services. It does not fit a deployment where you cannot hand three separate API keys to a hosted service, and the project says so plainly: there is no one-click path, and the felt latency is two to three seconds even though the engineering chain is under one. Before you swap the built-in persona for someone else, sort out permission rather than technique, because the shipped reference audio was assembled from a named creator's public livestreams, and the fifteen to forty-five seconds of clean speech the README asks for is exactly the material you would need from a person who has agreed to it. And read the memory caveat before you promise long-term recall, since the system only records what the user says, which in a persona-driven conversation is very little.

## FAQ

### What hardware does talk-to-fengge need?

The recommended synthesis path, VoxCPM, needs an NVIDIA GPU with 8 GB of VRAM or more, either local or a cloud GPU such as RunPod L4, and runpod_setup.sh installs the dependencies and downloads the model in about ten minutes. MOSS-TTS is the CPU fallback with weaker cloning, and Cartesia Sonic is the cloud option that needs a $5 per month Pro subscription to clone a voice.

### How do I start the talk-to-fengge server?

Three processes: livekit-server --dev --node-ip=127.0.0.1, then LLM_PROVIDER=minimax python -m worker.main start, then python -m worker.web_server. On macOS you can double-click Talk-to-Me-V3.6.command instead, and you chat at http://127.0.0.1:8766.

### How do I change the persona in talk-to-fengge?

Record 15 to 45 seconds of clear speech with no background music or noise, and describe the persona. The suggested route is to open the repository in a coding assistant, point it at the persona notes under docs/ and worker/persona.py as the template, and let it extract the defining features through conversation and write the new persona code.

### What is the real latency of talk-to-fengge?

The engineering chain is under one second, but the README puts the felt conversation latency at roughly two to three seconds depending on the network and how fast the APIs respond. There is also no one-click deployment, since STT, LLM and TTS each need their own key or service.

## Sources

- [Issues](https://github.com/YeJe-cpu/talk-to-fengge/issues)
- [License: Apache-2.0](https://github.com/YeJe-cpu/talk-to-fengge/blob/main/LICENSE)
- [Project website](https://x.com/leaf_sanren/status/2069342335268507976)
- [README](https://github.com/YeJe-cpu/talk-to-fengge/blob/main/README.md)
- [YeJe-cpu/talk-to-fengge on GitHub](https://github.com/YeJe-cpu/talk-to-fengge)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yeje-cpu-talk-to-fengge
