# TEN Framework: Real-Time Multimodal Conversational AI

> TEN is an open-source runtime for building real-time multimodal conversational AI agents, connecting speech recognition, language model inference, and text-to-speech through a composable extension graph. It is aimed at engineers who need a coordinated audio pipeline rather than a loose collection of API calls.

**TEN-framework/ten-framework** —  Open-source framework for conversational voice AI agents

- Repository: https://github.com/TEN-framework/ten-framework
- Website: https://agent.theten.ai/
- Stars: 11,143 · Forks: 1,367
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ten-framework-ten-framework

## What TEN Solves and Who It Is For

Building a real-time voice AI product requires more than routing text to an LLM. Audio must be captured, transmitted, transcribed, answered, synthesized, and played back with latency low enough that conversation feels natural. TEN addresses the coordination problem at the framework level. It provides a runtime that manages extension lifecycle, message routing, and real-time communication so that product teams can focus on agent logic rather than rebuilding those layers for each deployment.

The target audience is engineers building conversational agents for consumer or enterprise applications: voice assistants, meeting transcription tools, hardware integrations on embedded microcontrollers, and avatar-based agents. The project homepage is https://agent.theten.ai/ and the README introduces TEN as part of a broader TEN Ecosystem that includes the Framework core, Agent Examples, a VAD component, Turn Detection, and a Portal. Each component ships as a separate repository.

## Extension Graph: How TEN Wires Capabilities Together

TEN structures an agent as a directed graph of extensions. Each extension handles one capability: speech recognition, language model inference, text-to-speech synthesis, real-time communication, or custom business logic. Extensions communicate through a typed message bus managed by the runtime, which handles delivery, ordering, and lifecycle.

This model means that swapping providers is a graph change, not a code rewrite. The README describes a multi-purpose voice assistant that can be extended with Memory, VAD, and Turn Detection extensions independently. The WebSocket-based transport is documented as an alternative to Agora RTC for cases where a different real-time channel is preferred.

The repository layout reflects this modular design. The top-level directories include `ai_agents/` for the agent examples, `core/` for the runtime, `packages/` for extension packages, and `docs/`. The `Taskfile.yml` at the root drives build operations.

## Installing and Running the Agent Examples

The quick-start path runs inside Docker and requires four external API credentials: an Agora App ID and App Certificate for real-time transport, an OpenAI API key for the language model, a Deepgram API key for speech-to-text, and an ElevenLabs TTS key. The minimum system requirement listed in the README is 2 CPU cores and 4 GB of RAM, alongside Docker, Docker Compose, and Node.js LTS v18.

Clone the repository, navigate to the `ai_agents` directory, and create the environment file:

```bash
cd ai_agents
cp ./.env.example ./.env
```

Fill in the required values in `.env`:

```bash
AGORA_APP_ID=
AGORA_APP_CERTIFICATE=
DEEPGRAM_API_KEY=
OPENAI_API_KEY=
ELEVENLABS_TTS_KEY=
```

Start the development containers:

```bash
docker compose up -d
```

Enter the container:

```bash
docker exec -it ten_agent_dev bash
```

Choose an example and navigate to it:

```bash
cd agents/examples/voice-assistant
```

The README states that the first build takes approximately five to eight minutes. Run `task install` before the first launch and after changing Go source code or dependencies; it installs TEN, Python, and frontend dependencies and builds the Go API server.

## The Seven Bundled Agent Examples

The Agent Examples repository ships seven documented examples that illustrate the range the framework covers. The multi-purpose voice assistant supports both RTC and WebSocket connections and can be extended with Memory, VAD, and Turn Detection. The Doodler converts spoken or typed prompts into hand-drawn sketches with a crayon palette and real-time drawing. The speaker diarization example detects and labels speakers in a live stream; the README describes an interactive game as a demonstration use case.

The lip sync avatar example integrates with multiple vendors. The README names Trulience, HeyGen, and Tavus for realistic avatars, and Live2D characters with MotionSync-powered lip sync for anime-style characters. The SIP example powers phone calls through a TEN extension. The transcription example converts audio to text. The ESP32-S3 Korvo V3 example runs a TEN agent on an embedded development board, connecting LLM-powered communication with hardware.

Additional samples beyond the seven highlighted ones are listed in the `agents/examples` folder.

## Real Limits and Cases Where TEN Is the Wrong Tool

The default quick-start pipeline depends on four external paid APIs. A developer who wants to explore TEN without billing exposure on each service must replace each provider through the extension model, which requires understanding how to wire a new extension into the graph. The documentation for doing so is not directly included in the README material.

The License field in the repository metadata is listed as NOASSERTION. This is not a standard SPDX identifier. The README does not explain the licence terms, and the LICENSE file is a separate file in the repository root. Teams with open-source licence policies or commercial deployment requirements should read the actual LICENSE file before building on TEN.

TEN is architected for real-time multimodal interactions. It is not a tool for batch transcription pipelines, offline audio processing, or text-only chatbots where low latency has no value. Using it for those cases adds Docker infrastructure, Agora dependency, and extension graph complexity with no benefit. A simpler direct API integration is a better fit for those workloads.

## TEN Compared to LiveKit Agents

LiveKit Agents is the nearest open-source alternative. Both frameworks tackle real-time voice AI pipelines, but their infrastructure assumptions differ. LiveKit Agents centers on a Python SDK with room abstractions built on top of WebRTC, and it supports self-hosted server deployments with no dependency on a specific commercial RTC platform.

TEN's documented quick-start route uses Agora as the real-time transport. Agora is a commercial service. Developers already using Agora for video calling, or those who want the specific multi-vendor avatar integrations that TEN's examples demonstrate, have a concrete reason to choose TEN. Developers who want a Python-first setup, self-hosted transport, and no commercial RTC dependency will find the LiveKit path lower-friction from a cold start.

## Maintenance State and Licence

The last push to the main repository was on 2026-09-24, and the most recent release was version 0.11.73 on 2026-09-22. The project is not archived. Releases have been frequent in 2026, with patch versions shipping roughly weekly in September. The TEN Ecosystem spans multiple repositories: the framework core, agent examples, VAD, Turn Detection, and Portal each live separately, so tracking changes requires watching each one.

The VAD component handles voice activity detection and the Turn Detection component identifies when a speaker has finished their turn, both designed as pluggable extensions. The homepage for pre-built agent examples and a portal for testing them live is at https://agent.theten.ai/.

The license field shows NOASSERTION in the pack metadata. Until the actual LICENSE file content is reviewed, commercial users cannot determine what obligations apply. This is a practical adoption blocker for teams in regulated environments.

## Conclusion

TEN is the right choice for teams building production voice AI products who need control over the audio pipeline and the ability to swap ASR, LLM, and TTS providers without redesigning the transport layer. It is not suited to developers who want a simple chatbot with a few API calls: the Docker-based setup, four external API keys, and Agora dependency add real overhead. Before wide adoption, confirm that Agora is acceptable as the real-time transport for the target deployment region, and inspect the LICENSE file directly since the LICENSE field in that repository shows NOASSERTION.

## FAQ

### What is TEN Framework?

TEN is an open-source runtime for real-time multimodal conversational AI. It structures agents as a graph of extensions, each handling one capability such as speech recognition, language model inference, or text-to-speech, connected through a typed message bus.

### What external API keys does TEN require to run the bundled examples?

The quick-start path requires four keys: an Agora App ID and App Certificate for real-time transport, an OpenAI API key for the language model, a Deepgram API key for speech-to-text, and an ElevenLabs TTS key. All four are needed before the default voice assistant example will work.

### Does TEN Framework support hardware targets beyond desktop and cloud?

The README documents an example that runs a TEN agent on the Espressif ESP32-S3 Korvo V3 development board, integrating LLM-powered communication with the embedded hardware. A separate integration guide is linked from the README.

## Sources

- [Issues](https://github.com/TEN-framework/ten-framework/issues)
- [Project website](https://agent.theten.ai/)
- [README](https://github.com/TEN-framework/ten-framework/blob/main/README.md)
- [Releases](https://github.com/TEN-framework/ten-framework/releases)
- [TEN-framework/ten-framework on GitHub](https://github.com/TEN-framework/ten-framework)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ten-framework-ten-framework
