livekit/agents: a Python framework for realtime voice agents
Project brief: A framework for building realtime voice AI agents. Use it to create conversational, multi-modal voice agents that can see, hear, and understand.
At a glance
- What is it?
- livekit/agents is a Python framework for building realtime, programmable participants that run on servers. It is a strong fit when your agent has to listen and talk over WebRTC or a phone line, and a poor fit if you only need a text chatbot.
- Who is it for?
- Adopt livekit/agents if you are building a server-side agent that has to hold a spoken conversation over WebRTC or telephony, and you are willing to run a LiveKit server and wire up model plugins. Do not adopt it if you only need a text chat loop, since the audio session machinery, the job scheduler and the deployment surface are all cost you would not use.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem livekit/agents solves: a spoken agent is a session, not a prompt
A text agent is a loop around a model. A voice agent is a session around a model, and the session has to answer questions the model never sees: who is speaking, when they stopped, whether the agent should interrupt, and how the audio gets to the user's device. livekit/agents is aimed at that second problem. The README describes it as a framework for realtime, programmable participants that run on servers, built to create conversational, multi-modal voice agents that can see, hear, and understand.
The audience follows from that framing. It is for engineers who already accept that their agent runs as a long-lived process rather than a request handler, and who need transport, turn detection and scheduling handled by the framework. The repository ships an examples directory with frontdesk, healthcare, hotel_receptionist, drive_thru, telephony and warm-transfer scenarios, which tells you the intended deployments are call centres, reception desks and similar live conversations rather than one-shot assistants.
If your agent only exchanges text, this is more machinery than the problem needs. The core value here is the audio session, and you do not get it for free: you pay for it in plugins, environment variables and a server to run.
How the pieces fit: Agent, AgentSession, entrypoint and AgentServer
The README names four core concepts. An Agent is an LLM-based application with defined instructions. An AgentSession is a container for agents that manages interactions with end users. An entrypoint is the starting point for an interactive session, described as similar to a request handler in a web server. An AgentServer is the main process that coordinates job scheduling and launches agents for user sessions.
The data flow implied by that list is a pipeline of plugins. The README states that any combination of STT, LLM, TTS, or realtime API can be used, and the feature list describes a comprehensive ecosystem to mix and match STT, LLM, TTS and Realtime API providers. Turn detection is not left to silence thresholds alone: the README says the framework uses a transformer model to detect when a user is done with their turn, which it says helps reduce interruptions.
Two details are worth separating from the marketing language. First, dispatch APIs are how end users get connected to agents, so scheduling is part of the framework rather than something you build around it. Second, the repository is a workspace: the top-level pyproject.toml lists livekit-agents and dozens of livekit-plugins-* packages as workspace sources, covering providers such as openai, deepgram, cartesia, elevenlabs, google, groq and others. That layout means plugin versions move with the core package rather than being independently released, which simplifies upgrades but also means a plugin fix arrives on the framework's schedule.
Installing livekit/agents and running a first voice agent
The README gives a single installation command that pulls the core library plus plugins for popular model providers. The bracket syntax is a pip extra, so the provider names are part of the install line and not a separate step.
pip install "livekit-agents[openai,deepgram,cartesia]"The README then states that the simple voice agent example needs three environment variables: LIVEKIT_URL, LIVEKIT_API_KEY and LIVEKIT_API_SECRET. These point the agent at a LiveKit server and authenticate it. The examples directory contains an .env.example file, which is the conventional place to put them.
The example itself defines a function tool, creates an AgentServer, and registers an entrypoint with the server.rtc_session() decorator. The session in the README passes vad=inference.VAD() and notes that any combination of STT, LLM, TTS or a realtime API can be used. The abbreviated snippet in the README does not show the full model configuration, so read the complete file at examples/voice_agents/basic_agent.py before copying it.
from livekit.agents import (
Agent,
AgentServer,
AgentSession,
JobContext,
RunContext,
cli,
function_tool,
inference,
)
@function_tool
async def lookup_weather(
context: RunContext,
location: str,
):
"""Used to look up weather information."""
return {"weather": "sunny", "temperature": 70}For a multi-agent flow, the README points at examples/voice_agents/multi_agent.py and shows the pattern of subclassing Agent, setting instructions in __init__, and calling self.session.generate_reply() inside on_enter to open the conversation. That on_enter hook is where you would put a greeting or an opening question. The README marks that snippet as abbreviated, so treat the file as the reference.
Testing a non-deterministic voice agent with the builtin framework
The README is explicit that automated tests matter here because of the non-deterministic behaviour of LLMs, and it ships a native test integration. The example is a pytest coroutine that builds an AgentSession around a provider LLM, starts an agent, and then runs a scripted user input.
The assertions are event-based rather than string-based. The result object exposes expect, and the README shows skip_next_event_if, next_event().is_function_call(name=...), is_function_call_output(), and a message expectation with a judge. That style is a real design decision: you are asserting on the sequence of events the session produced, not on the exact words the model chose, which is the only stable way to test a system whose output varies between runs.
@pytest.mark.asyncio
async def test_no_availability() -> None:
llm = google.LLM()
async with AgentSession(llm=llm) as sess:
await sess.start(MyAgent())
result = await sess.run(
user_input="Hello, I need to place an order."
)
result.expect.skip_next_event_if(type="message", role="assistant")
result.expect.next_event().is_function_call(name="start_order")
result.expect.next_event().is_function_call_output()The limitation is that these tests exercise the agent logic, not the audio path. A green suite tells you the tool calls and message sequence are right. It does not tell you that a real caller on a noisy phone line will be understood, because the STT and turn-detection behaviour under real audio is outside what this harness asserts.
Where livekit/agents is the wrong tool
The clearest boundary is text-only work. If your agent answers questions in a chat window and never speaks, the AgentSession, the VAD, the semantic turn detector and the dispatch APIs are all overhead. A plain model SDK plus your own request handler will be smaller and easier to debug.
The second boundary is operational. The README presents running the entire stack on your own servers as a feature, and it names LiveKit server, its WebRTC media server, as part of that stack. That is a genuine capability, but it is also a service you now operate. Teams that want a managed voice pipeline and no media server to run should weigh that before starting.
The third is provider coupling in the other direction: the framework mixes and matches providers, but the quality of the conversation depends on the combination you choose. Turn detection uses a transformer model per the README, and the README claims it helps reduce interruptions. It does not claim to eliminate them. If your use case tolerates barge-in poorly, for example a compliance script that must finish each sentence, you will be tuning against the framework's defaults rather than getting the behaviour for free.
Finally, the README does not document rollback or downgrade procedures for the workspace packages. If you need a documented rollback path, that is something to establish yourself rather than something the repository currently provides.
livekit/agents compared with LiveKit AgentsJS and LangChain-style tool agents
The README points readers who want the JavaScript or TypeScript library to AgentsJS at github.com/livekit/agents-js. That is the closest alternative and the difference is mostly language and runtime: the same project family, the same LiveKit transport and dispatch model, but a Node-oriented codebase instead of Python. If your team writes TypeScript and your agent logic lives next to a web backend, AgentsJS avoids running a second runtime for the agent process. If your tooling, tests and model SDKs are Python, the repository you are reading is the one that matches.
The more interesting contrast is with tool-calling frameworks such as the LangChain plugin that appears in this repository's workspace list as livekit-plugins-langchain. That plugin exists precisely because the two approaches are not the same. A general tool-agent framework is built around chains and tool invocation over text; livekit/agents is built around a realtime session with audio in and audio out, turn detection, job dispatch and WebRTC clients. Using the plugin lets you keep LangChain-style logic inside a LiveKit session, which is a reasonable middle path, but it does not remove the session machinery. You still need the LiveKit server, the credentials and the entrypoint.
So the choice is not which framework is better in the abstract. It is whether your agent's primary interface is a live audio session. If it is, the session-level pieces in this repository are the ones you would otherwise have to build.
Maintenance, licensing and the cost of upgrading livekit/agents
The repository is not archived, and the last push was on 2026-08-27. The most recent release listed is [email protected], dated the same day, following 1.7.0 and 1.6.10 earlier in August 2026. That release cadence, with patch versions arriving days apart, is the practical upgrade signal: you should expect to bump the package regularly rather than pinning once a year.
The workspace layout makes the upgrade cost concrete. Because the top-level pyproject.toml declares livekit-agents and the livekit-plugins-* packages as workspace sources, the core library and the provider plugins are versioned together. A single upgrade can move your STT, LLM and TTS integrations at the same time as the session code. The upside is that you will not hit a mismatched plugin against a newer core. The downside is that you cannot upgrade one provider in isolation, and the built-in test framework is the tool the README offers for catching behavioural changes across that jump.
On licensing, the repository is Apache-2.0, which is the licence most teams will evaluate first. Two files in the top-level listing complicate a quick read: MODEL_LICENSE and NOTICE. The presence of a separate model licence file means not everything bundled here necessarily falls under Apache-2.0, and the NOTICE file is where attribution and third-party terms are recorded. Check both against the plugins you actually install. This is a description of what the repository contains, not legal advice; if the distinction matters to your organisation, have counsel read those files rather than relying on the licence label alone.
Editorial conclusion
Adopt livekit/agents if you are building a server-side agent that has to hold a spoken conversation over WebRTC or telephony, and you are willing to run a LiveKit server and wire up model plugins. Do not adopt it if you only need a text chat loop, since the audio session machinery, the job scheduler and the deployment surface are all cost you would not use. Before committing, verify your own STT, LLM and TTS combination end to end with the built-in test framework, because turn-taking quality comes from the plugin mix rather than from the framework alone.
Frequently asked questions
What is livekit/agents?
It is a framework for building realtime, programmable participants that run on servers, used to create conversational, multi-modal voice agents that can see, hear, and understand. The core concepts are Agent, AgentSession, entrypoint and AgentServer.
How do I install livekit/agents?
The README gives one command, pip install "livekit-agents[openai,deepgram,cartesia]", which installs the core library along with plugins for those model providers. The simple voice agent example then needs LIVEKIT_URL, LIVEKIT_API_KEY and LIVEKIT_API_SECRET.
How do I use livekit/agents in a project?
You create an AgentServer, register an entrypoint with the server.rtc_session() decorator, and build an AgentSession inside it with your chosen STT, LLM, TTS or realtime API combination. Tools are added with the function_tool decorator, and the examples directory holds full working files.
What are some examples of livekit/agents use cases?
The examples directory includes frontdesk, healthcare, hotel_receptionist, drive_thru, telephony and warm-transfer scenarios, alongside voice_agents examples such as a starter agent, push-to-talk and background audio.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/livekit-agents)