# LiveKit Agents for Node.js: a realtime voice agent framework that keeps the media server in the loop

> livekit/agents-js is the TypeScript distribution of the LiveKit Agents framework, aimed at teams building conversational voice agents that join WebRTC rooms. It ships a plugin per model provider and a worker model for job scheduling, but the Node.js port is narrower than the Python original.

**livekit/agents-js** — Build realtime multimodal AI agents with Node.js

- Repository: https://github.com/livekit/agents-js
- Website: https://docs.livekit.io/agents
- Stars: 938 · Forks: 379
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/livekit-agents-js

## The problem livekit/agents-js solves

Most voice agent demos are a loop: capture audio, send it to a speech-to-text service, send the transcript to a language model, send the reply to a text-to-speech service, play it back. The hard part is not the loop. It is the interruption handling, the turn boundaries, and the fact that a caller on a phone or in a browser is not a file you can process at your own pace. LiveKit Agents for Node.js targets that second problem. The README describes the framework as being for "realtime, programmable participants that run on servers", and the core concepts list makes the model explicit: an Agent is an LLM-based application with instructions, an AgentSession is the container that manages interaction with end users, an entrypoint starts an interactive session, and a Worker coordinates job scheduling and launches agents for user sessions. That last piece is the tell. This is not a library you call from a request handler. It is a process that registers with a LiveKit server, receives jobs, and spins up a session per participant. The audience is teams already using WebRTC, or willing to, and who want the voice pipeline to be one participant in a room rather than a separate service glued to a client with custom signalling. If your agent never needs to share a room with a human in real time, the worker and session machinery is overhead you are paying for nothing.

## How the worker, session and plugin layers fit together

The architecture visible in the README is a four-layer stack. At the bottom is the LiveKit server, which the project points to as an open-source WebRTC media server you can run yourself. Above that sits the Worker process, which the README calls "the main process that coordinates job scheduling and launches agents for user sessions". The Worker registers with the server and waits. When a user joins, the Worker launches an AgentSession, which holds the Agent and manages the back-and-forth with that user. The Agent itself is where your instructions and tools live. The plugin layer is orthogonal to all of this: each plugin wraps a provider for one or more of STT, LLM, TTS, Realtime or VAD, and you mix them. The README's features list names the pieces that matter for latency: semantic turn detection, described as a transformer model that detects when a user is done with their turn and "helps to reduce interruptions", and MCP support, which the README says lets you integrate tools provided by MCP servers "with one line of code". Data flows both ways over the LiveKit data APIs, and the README calls out RPC as the mechanism for exchanging data with clients. That is the design bet: the agent is a room participant, so client-to-agent calls and agent-to-client calls use the same transport as the audio.

## Installing livekit/agents-js and running a first agent

The README's installation section is short and assumes pnpm. It first tells you to install pnpm globally if you do not have it, then to install the core library plus whichever plugins you need. Note that the core package is installed with pnpm even though the pnpm bootstrap itself uses npm.

```bash
npm install -g pnpm
pnpm install @livekit/agents
```

Plugins are separate packages. The README lists the supported set for Node.js as a table, and it is worth reading before you commit to a provider: @livekit/agents-plugin-openai covers LLM, TTS and STT, @livekit/agents-plugin-google covers LLM and TTS, @livekit/agents-plugin-deepgram covers STT and TTS, @livekit/agents-plugin-elevenlabs covers TTS only, and @livekit/agents-plugin-silero provides VAD. Install the ones matching your pipeline.

The README's usage section points to the quickstart guide rather than repeating it, and then gives a simple voice agent example. The imports show the shape of a real file: the core package exports cli, defineAgent, WorkerOptions, llm, voice and inference, along with the JobContext and JobProcess types, and the example pulls in @livekit/agents-plugin-silero for voice activity detection and zod for tool parameter schemas. A tool is declared with llm.tool, taking a description, a zod parameters object and an execute function that receives the parsed arguments and a context. The snippet is truncated in the README, so treat it as a sketch of the API surface rather than a runnable program.

For a complete starting point, the README recommends the LiveKit Agents Starter for Node.js at livekit-examples/agent-starter-node. It describes that repository as including a ready-made assistant, multilingual turn detection, background noise cancellation, metrics and logging, and a production-ready Dockerfile. For a first real use, that is the faster path than assembling plugins by hand, because the starter already wires the LLM, STT and TTS pipeline together.

## Where the Node.js port is narrower than you might expect

The README states plainly that this is a Node.js distribution of a framework "originally written in Python". That sentence carries the main limitation. The Python framework has a larger plugin catalogue, and the Node.js README does not claim parity. It says "Currently, only the following plugins are supported" and then gives a table. If the provider you want is not in that table, the README gives you no path forward, no adapter interface to implement, and no statement about whether one exists elsewhere. That is a real constraint when you are choosing a stack, because swapping an STT vendor later may mean swapping frameworks.

The second limitation is the deployment shape. A Worker is a long-running process that registers with a LiveKit server and receives jobs. That does not fit a serverless function model, where the process starts per request and dies. The README does not discuss horizontal scaling, worker capacity, or what happens when a Worker dies mid-session, so those are questions to take to the documentation site rather than the repository README. Third, the core concepts assume you accept the room model. If your application is a batch transcription job or a text-only chatbot, the Worker and AgentSession layers are dead weight, and a plain SDK call to your model provider would be simpler.

## LiveKit Agents versus the OpenAI Agents SDK for JavaScript

The OpenAI Agents SDK for JavaScript is the comparison people reach for, and the difference is not the model provider. It is what the agent is attached to. The OpenAI SDK is built around running an agent loop: you give it instructions and tools, it calls a model, it may call tools, it returns a result. Transport is your problem. LiveKit Agents for Node.js assumes the transport is a WebRTC room and builds the session, the turn detection and the job scheduling into the framework. The README's features list reflects that: extensive WebRTC clients across major platforms, RPC and data APIs for exchanging data with clients, semantic turn detection to reduce interruptions, and a VAD plugin in the supported table. None of that is about which model answers the question. It is about keeping a live audio session coherent when a human interrupts.

The practical consequence is that LiveKit Agents is heavier to stand up. You need a LiveKit server, a Worker process, and a client that joins the room. The OpenAI SDK needs an API key. If your product is a voice agent that a user talks to in a browser or on a phone, that weight buys you interruption handling and a media path you would otherwise build. If your product is a tool-calling assistant behind a text box, it buys you nothing.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-10. Recent releases are versioned in lockstep across packages: @livekit/agents@1.8.0, @livekit/agents-plugin-xai@1.8.0 and @livekit/agents-plugin-trugen@1.8.0 all landed on 2026-09-05. The monorepo layout supports that: a pnpm workspace with agents/, plugins/, examples/ and tests/ at the top level, turbo for task running, changesets for publishing, and a ci:publish script that runs the build before changeset publish. Plugins version together with the core, which simplifies upgrades but means a plugin bump can arrive with a core bump you did not ask for.

The README carries a migration warning at the top: it says livekit-agents 1.5.0 was released and links to a migration guide. The presence of a migration guide is the honest signal here. Major-ish upgrades have required code changes in the past, and the README does not document a rollback procedure for a version bump, so pin your versions and read the migration guide before moving.

The licence is Apache-2.0, and the repository layout is consistent with that: a LICENSE file, a LICENSES/ directory, a NOTICE file, a REUSE.toml for REUSE compliance, and SPDX headers in the README itself. There is also a separate MODEL_LICENSE file at the top level, which suggests at least one bundled model carries terms different from the code. Apache-2.0 covers the source; it does not cover the model weights or the third-party provider APIs your plugins call. Check the model licence and each provider's terms separately. This is not legal advice.

## Conclusion

Adopt livekit/agents-js if your voice agent needs to live inside a WebRTC room and you want the media path, RPC and data channels handled by the same stack, with a plugin per model provider. Do not adopt it if you need a Python-only plugin, since the README lists a fixed supported set for Node.js, or if you want a single-vendor realtime API with no room concept. Before writing application code, run pnpm install @livekit/agents in a scratch directory and confirm the plugin you need appears in that supported table, then clone livekit-examples/agent-starter-node and check its Dockerfile against your deployment target.

## FAQ

### What is livekit/agents-js?

It is the Node.js distribution of the LiveKit Agents framework, originally written in Python. The README describes it as a framework for building realtime, programmable participants that run on servers, used to create conversational, multi-modal voice agents that can see, hear and understand.

### How do I install livekit/agents-js?

The README says to install pnpm globally with npm if you do not have it, then run pnpm install @livekit/agents for the core library, adding whichever plugin packages your pipeline needs.

### Which plugins does livekit/agents-js support?

The README states that only a listed set is supported, including @livekit/agents-plugin-openai for LLM, TTS and STT, @livekit/agents-plugin-google for LLM and TTS, @livekit/agents-plugin-deepgram for STT and TTS, @livekit/agents-plugin-silero for VAD, and several avatar and TTS plugins.

### Is there a starter app for LiveKit Agents on Node.js?

Yes. The README recommends livekit-examples/agent-starter-node, which it says includes a ready-made assistant, multilingual turn detection, background noise cancellation, metrics and logging, and a production-ready Dockerfile.

## Sources

- [License: Apache-2.0](https://github.com/livekit/agents-js/blob/main/LICENSE)
- [livekit/agents-js on GitHub](https://github.com/livekit/agents-js)
- [Project website](https://docs.livekit.io/agents)
- [README](https://github.com/livekit/agents-js/blob/main/README.md)
- [Releases](https://github.com/livekit/agents-js/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/livekit-agents-js
