datascale-ai/opentalking: an orchestration layer for real-time digital humans
OpenTalking: An industrial-grade open-source AI digital human framework that supports real-time conversation, private deployment, and pluggable models.
At a glance
- What is it?
- A Python and React framework that wires LLM replies, speech recognition, speech synthesis and WebRTC into a digital human conversation product, with mock mode for trying it without a GPU.
- Who is it for?
- The useful thing about opentalking is that it treats the digital human conversation path as a product problem rather than a model problem. Session state, interruption control, subtitle events, voice selection, WebRTC playback and the provider plumbing are all first class, and mock mode lets you exercise the whole path without model weights or a GPU.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 33 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Mock mode is the default deployment path
The default compose profile runs four services with no GPU requirement: redis, the API, the worker and the web console, with MOCK synthesis. The comments at the top of `docker-compose.yml` describe the intent plainly, which is to let you click around the UI and exercise the full pipeline except the actual lip sync model.
docker compose upThe GPU profile adds a fourth service, `omnirt`, which is a real inference runtime:
docker compose --profile gpu upThat profile needs an NVIDIA driver and the nvidia container toolkit on the host, and the compose file requests all GPUs for the container. The `omnirt` service is a health checked HTTP service on port 9000, so the API can wait for it to be ready rather than failing requests on a cold start.
Tearing the stack down is the plain compose command:
docker compose downThe API listens on 8000, the worker on 9001, and the environment template sets the web console to 5173 with a separate `VITE_BACKEND_PORT` pointing back at 8000. The API's environment block sets the mock flag with a default of 1 and an empty OmniRT endpoint, so the same compose file serves both profiles once the override applies.
Reading the dependency list as an architecture diagram
The dependency list in `pyproject.toml` is more informative about the design than the prose, because it separates the always-installed orchestration layer from the heavy inference extras.
The base dependencies are the product surface: `fastapi` and `uvicorn` for the API, `pydantic` and `pydantic-settings` for configuration, `redis` for session and job state, `httpx` for outbound calls to model providers, `websockets` and `aiortc` for the media path, `loguru` for logging, and `python-multipart` plus `pillow` for avatar uploads.
Then there are the three provider modules the framework treats as independent: `edge-tts` for speech synthesis, `dashscope` for the Alibaba speech and model APIs, and `openai` for an OpenAI compatible client. `lightrag-hku` and `mem0ai` are there for knowledge base retrieval and conversation memory respectively.
The computer vision stack is notable for how carefully it is pinned. `mediapipe` has two exact pins, one for aarch64 and one for everything else, both restricted to Python below 3.13. `av` is similarly capped below 14.3 on older Pythons, `numpy` is held below 2, and `opencv-python` is held below 4.12. Those are not casual constraints, they are what keeps the install from breaking on a platform where those packages have not published a compatible wheel.
The `engine` extra is where the video model work lives: `torch`, `diffusers`, `xformers`, `xfuser`, `accelerate`, `librosa` and `safetensors` among others. A separate `models` extra exists for people who only want the model backends without the full engine.
The environment template is where the real configuration lives
`.env.example` is unusually detailed for a template file, and reading it is the fastest way to understand the configuration surface. It opens with a warning that matters: LLM, STT and TTS are three independent modules, and unless a vendor explicitly says otherwise you should not reuse one provider key across them. That is a small line that prevents a class of confusing debugging session.
The template also proposes a directory layout for a deployment root, with separate subdirectories for the source checkout, model weights, third party model repositories checked out as source, and sidecar runtime virtual environments. Four environment variables encode that split. Keeping model weights out of the source tree is the point of the whole arrangement, since the weights are large and the source is not.
Service basics are next, with the API and web hosts, the ports, and `VITE_BACKEND_PORT` for the frontend build. Then avatar handling: a directory for avatar assets pointing at `examples/avatars`, an exports directory, and a background removal provider set to `rembg` with a device setting. The template is candid that enabling local background removal requires installing the extra first, and that the provider needs `u2net.onnx` downloaded in advance with the model path filled in explicitly, including the expected MD5.
There is also a default model setting that ships as `mock`, a torch device setting of `auto`, and an ffmpeg binary path that is left empty to auto detect the system ffmpeg and fall back to imageio-ffmpeg. Each of those fallbacks is documented rather than hidden.
What the single v0.1.0 release actually ships
There is one tagged release, v0.1.0, published 2026-06-15, and its notes are worth reading closely because they describe scope accurately rather than broadly.
The release packages five things: the FastAPI service, an async worker, a React web console, a model provider registry, and deployment paths backed by documentation. Python source distribution and wheel artifacts are attached, and three GHCR images are published, one each for the API, the worker and the web console.
Two explicit limitations are stated. Model weights are not bundled in the Python artifacts or the Docker images. And the release images package only the orchestration services, because real talking head inference still requires either a configured local model backend or an OmniRT compatible remote runtime. That is a fair and useful thing for a first release to say outright, since it is exactly the boundary a buyer needs to know about.
Compatibility is Python 3.10 and newer, which matches the `requires-python` in pyproject. The repository was last pushed on 2026-09-04 and is not archived, with 23 open issues against 3039 stars, so the shape of the project is a well received one still early in its release history.
Four workflow categories and a WebUI that configures all of them
The README organises the product around four workflows rather than around components, which tells you what the authors think the hard parts are. There is real time conversation, video creation and cloning, and the demo material is grouped under those headings.
Video creation is split by what drives it: audio driven, text driven, and cloned voice driven. Video cloning is split by input: realtime camera imitation and uploaded video imitation. Both reuse the FasterLivePortrait runtime, so the same engine covers generation from audio and cloning from a camera.
The product scenarios in the demos are worth noting as a statement of intended use: healthcare guidance, live commerce, a tourism guide, e-commerce livestream, a companion character, and a news anchor. That range is the argument for the orchestration layer existing at all, because a companion character and a news anchor want very different latency and interruption behaviour from each other.
The WebUI is where configuration happens, and it is a substantial part of the product rather than a demo page. On one page you can select or create avatars, configure voices and the LLM, TTS and STT providers, pick a digital human driver model, inspect model connection status, and validate real time conversation with subtitles and audio video playback. There are `examples/avatars/` and `examples/personas/` directories in the tree for the asset side of that.
The README is front loaded with demo videos before reaching deployment paths, quickstart, supported models and a roadmap, so a reader looking for install steps should jump rather than scroll.
Development commands and the docs pipeline
The Makefile is three targets, which says the project expects the container path for anything else:
pytest tests -vLinting is Ruff over a specific set of directories rather than the whole tree:
ruff check opentalking/core opentalking/events opentalking/avatar apps testsThat list is a small map of the internal structure. `opentalking/core`, `opentalking/events` and `opentalking/avatar` are the framework packages, `apps` holds the applications including the web console, and `tests` is the test suite. The `events` package name is consistent with the README's emphasis on subtitle events and interruption control as first class concerns.
The web console builds with:
cd apps/web && npm run buildThe rest of the tree is documentation and scaffolding. `docs/` plus `mkdocs.yml` is the documentation site, published at datascale-ai.github.io with English and Chinese variants, and `README.md`, `README.en.md` and `README.zh.md` are three parallel readmes. There is an `AGENT.md` at the root alongside `CONTRIBUTING.md` and `CODE_OF_CONDUCT.md`, a `.pre-commit-config.yaml`, a `conftest.py` at the root for pytest configuration, and a `docker/` directory holding the separate API and worker Dockerfiles the compose file references.
`docker-compose.gpu.yml` and `scripts/` complete the picture: the GPU stack is a separate compose file rather than a profile, and the scripts directory is where helper commands live.
Editorial conclusion
The useful thing about opentalking is that it treats the digital human conversation path as a product problem rather than a model problem. Session state, interruption control, subtitle events, voice selection, WebRTC playback and the provider plumbing are all first class, and mock mode lets you exercise the whole path without model weights or a GPU. That makes it a reasonable shell to build a product on, provided you are willing to supply the inference backend yourself. Start with the default `docker compose up` profile to see the WebUI working with mock synthesis, then look at `docker-compose.gpu.yml` and the `engine` extra in pyproject before committing to a model path. Model weights are deliberately not bundled in either the Python artifacts or the images, so the real work is the OmniRT runtime or a local QuickTalk or Wav2Lip setup. One release tag, v0.1.0, means the API surface may still move.
Frequently asked questions
What is OpenTalking?
It is an open source orchestration framework for real time digital human conversations, covering the full path from frontend interaction and session state through LLM replies, speech recognition, speech synthesis, voice selection, interruption control, subtitle events and WebRTC playback. It is a coordination layer rather than a model, with pluggable local or remote model backends behind it.
Can I run opentalking without a GPU?
Yes, using the default compose profile, which runs redis, the API, the worker and the web console with MOCK synthesis. The comments in docker-compose.yml describe it as a way to exercise the whole pipeline except the actual lip sync model. The GPU profile is separate and needs an NVIDIA driver plus the nvidia container toolkit.
Does the OpenTalking Docker image include the video model weights?
No. The release notes state explicitly that model weights are not bundled in the Python artifacts or the Docker images, and that the images package only the orchestration services. Real talking head inference needs a configured local model backend or an OmniRT compatible remote runtime, which is why the environment template separates model weight directories from the source tree.
Which Python versions does opentalking support?
Python 3.10 and newer. Several pinned dependencies restrict themselves below 3.13, including mediapipe, which has separate exact pins for aarch64 and other platforms, and av below 14.3 on older interpreters. Python 3.12 or lower is the range where those pins line up.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/datascale-ai-opentalking)