# ElevenLabs JS is the Node SDK, and its play helper needs mpv installed separately

> The official Node library for the ElevenLabs API, generated by Fern, with text to speech, streaming, voice listing and a Speech Engine server that inverts the usual direction and waits for ElevenLabs to connect. Two things to read first: audio playback shells out to mpv, and Speech Engine can be told to skip authentication.

**elevenlabs/elevenlabs-js** — The official JavaScript (Node) library for the ElevenLabs API.

- Repository: https://github.com/elevenlabs/elevenlabs-js
- Website: https://elevenlabs.io
- Stars: 443 · Forks: 86
- Language: TypeScript
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/elevenlabs-elevenlabs-js

## Node is one of three packages, and picking the wrong one costs you mpv

The README opens with a routing note that saves a lot of wasted installs. This is the Node.js library. For the browser there is `@elevenlabs/client`, and for React there is `@elevenlabs/react`.

Installation is the standard pair:

```bash
npm install @elevenlabs/elevenlabs-js
# or
yarn add @elevenlabs/elevenlabs-js
```

The client takes an API key at construction, and the key it is given falls back to an environment variable:

```ts
const elevenlabs = new ElevenLabsClient({
    apiKey: "YOUR_API_KEY", // Defaults to process.env.ELEVENLABS_API_KEY
});
```

So the practical sequence is: install, set `ELEVENLABS_API_KEY`, construct the client once. The generation call is a method on `textToSpeech`, taking a voice identifier and an options object carrying the text and the model id, and the README's own example plays the result back with `play(audio)`.

One housekeeping detail explains a lot about the package. It is generated by Fern, and the repository carries `.fern/`, `.fernignore` and a `reference.md`, so the public surface comes from an API definition rather than being hand maintained.

## play() shells out to mpv, so audio output depends on the machine

The README flags this in a warning line between the example and the next section, and it is the first thing to surprise someone installing the package on a clean server.

`elevenlabs-js` requires MPV and ffmpeg to play audio. Those are system binaries, not npm packages, so `npm install` will succeed on a machine that cannot play a single sample through the helper.

The dependency list explains the mechanism. The runtime dependencies are `node-fetch`, `command-exists` and `ws`, and `command-exists` is how a library asks whether a binary is on the PATH before shelling out to it. The playback path is therefore: write the audio to a temporary file, find mpv, and hand it to a local process.

That is a reasonable choice for a quickstart, and a real constraint for production. A container image that serves an API has no reason to carry a media player, and an air-gapped or minimal image cannot. In that setting the streaming method is the one to use, handing the audio stream to the client application instead of decoding it server side.

The README does not document a configuration flag for disabling or replacing the player, so the choice at install time is the only one available.

## Three models, and the trade is made in the modelId field

The README names three models and positions them against each other on latency, language coverage and price.

Eleven Multilingual v2, as `eleven_multilingual_v2`, is the one recommended for most use cases, described as excelling in stability, language diversity and accent accuracy, and supporting 29 languages.

Eleven Flash v2.5, as `eleven_flash_v2_5`, is the latency option at 32 languages, and the only one with a price claim attached: 50% lower price per character.

Eleven Turbo v2.5, as `eleven_turbo_v2_5`, also covers 32 languages and is positioned as the balance of quality and latency, aimed at developer use cases where speed is the deciding factor.

Two things follow from those numbers. The jump from 29 to 32 languages between the default and the fast options means the recommended model is not the widest one, which is worth knowing if language coverage is your constraint. And the 50% per-character claim applies to Flash, so a latency-driven switch also changes your bill.

The model id is a plain string in the options object, and more models exist than these three, listed in the models documentation rather than the README.

## Speech Engine inverts the direction: ElevenLabs connects to you

Speech Engine is the part of the SDK with an unusual shape. It is for building voice-powered agents with a custom LLM, or adding voice to an existing chat agent. The ElevenLabs API connects to your server over WebSocket, each connection represents one conversation, and the division of labour is that you provide the LLM responses while ElevenLabs handles the speech.

That direction matters operationally. Your server is the one that must be reachable, so it needs a public endpoint and a certificate, and the address you configure on the ElevenLabs side is either a path on a server you already run or a host and port of your own.

There are two setup shapes for it. To attach to an existing HTTP server, used when you already have Express, Next.js or Fastify, you pass the server, a path and a set of callbacks to `speechEngine.attach`, and point the speech engine's server URL at that path. For a dedicated server, `SpeechEngine.Server` takes a port and an engine id, and every connection is treated as a speech engine session, with the URL pointed straight at the host and port:

```ts
const server = new SpeechEngine.Server({
    port: 3001,
    engineId: "seng_123",
    apiKey: process.env.ELEVENLABS_API_KEY,
    async onTranscript(transcript, signal, session) {
        const reply = await generateLLMResponse(transcript, { signal });
        session.sendResponse(reply);
    },
});

server.start();
```

The transcript callback receives the text, an abort signal and the session, which is everything needed to run a model turn and push the answer back.

## Every incoming connection is checked against an ElevenLabs-signed header

The default behaviour is the security story worth understanding. Both `SpeechEngine.Server` and `speechEngine.attach()` automatically verify the `X-Elevenlabs-Speech-Engine-Authorization` header on every incoming connection, rejecting any request that was not signed by ElevenLabs.

Per connection, not once at startup. That is the right shape for it, because a WebSocket endpoint left open is an endpoint anyone can find, and a check at startup protects nothing once the process is running.

The header name is specific enough that you will not collide with it, and the API key used for verification comes from the `apiKey` option or from the `ELEVENLABS_API_KEY` environment variable, following the same pattern as the main client.

So by default, an exposed Speech Engine server accepts traffic only from ElevenLabs. Anything that gets past that would have to forge a header signed by the vendor, and the SDK does the checking rather than leaving it to you.

## disableAuth: true replaces a signed-header check with an open endpoint

There is a documented way to turn that off, and the README is unusually direct about the cost:

```ts
// Standalone — no apiKey required when disableAuth is true
new SpeechEngine.Server({ port: 3001, disableAuth: true, onTranscript }).start();
```

The option exists for a specific deployment. If something in front of your server already restricts incoming traffic to ElevenLabs, such as an IP allowlist scoped to ElevenLabs' egress ranges, then the header check is redundant and the key is not needed.

Without such a layer, the README says it plainly: the server accepts any client that can reach it, and anyone on the internet can open a session and consume your compute and your downstream LLM quota. It also logs a `console.warn` on startup, which is the right level of friction for a setting whose failure mode is someone else's bill.

The two failure paths are worth separating. An IP allowlist misconfigured to allow too much is a silent exposure, and the `console.warn` will not catch it. The setting is reasonable in front of a correct allowlist and dangerous in front of nothing, and the difference between those two states is not visible from inside the process.

## Five session events, and only one of them carries the abort signal

The event table is short and the shape of it matters for anyone building a conversation UI.

`user_transcript` maps to `onTranscript` and is the substantive one: user speech transcribed, including the full conversation history and an abort signal. `init` maps to `onInit` and fires with a conversation id. `close` maps to `onClose` for a clean disconnect from ElevenLabs. `disconnected` maps to `onDisconnect` when the WebSocket drops unexpectedly. `error` maps to `onError` for a protocol or WebSocket error.

Two distinctions are doing real work here. Clean close and unexpected drop are separate events, so a UI can tell a finished conversation from a dropped one and reconnect rather than starting over.

And the abort signal lives on `onTranscript`, not on the connection. It is passed alongside the transcript text and the session, which means a cancelled user turn can be observed at the point the turn actually happens, and the same signal is what the standalone example forwards into `generateLLMResponse` so a superseded model call can be abandoned.

Passing a signal into your own LLM call rather than ignoring it is what keeps a fast-speaking user from stacking up work that nobody will read.

## The build compiles TypeScript into the package root, and the badge points at the Python repo

The packaging metadata explains a few things a user would otherwise trip over. The version is 2.70.0, the licence is MIT, `main` is `./index.js` and `types` is `./index.d.ts`, and a `prepack` script copies `dist/.` into the package root before publishing. That is why the published layout is flat rather than nested under `dist/`.

The toolchain is Biome for format, lint and check, with the same four flags repeated across every script: `--skip-parse-errors --no-errors-on-unmatched --max-diagnostics=none`. Those flags mean generated code never blocks a run, which is the point of a generated SDK but also means a parse error can pass silently.

Tests are Jest split into two projects, `test:unit` and `test:wire`, with msw in the dev dependencies for mocking HTTP. Three runtime dependencies in total: node-fetch, command-exists and ws.

One bug is visible in the README itself. The generator badge at the top points at a query string containing `fern-elevenlabs/elevenlabs-python/readme`, so a build badge for this Node library is reading from the Python repository. It is a copy-paste artefact rather than a functional problem, but it means the one indicator of generation status on the page is reporting on something else.

## Conclusion

Adopt this SDK for a Node service that needs ElevenLabs voices, and read the Speech Engine section before you expose one, because the default is a signed-header check on every connection and turning it off turns a private endpoint into an open one. Reach instead for @elevenlabs/client in a browser and @elevenlabs/react in a React app, since this package is explicitly the Node library and its play helper expects mpv and ffmpeg on the machine. Two checks before you build: install mpv and ffmpeg if you intend to use play() at all, and read the retries section against your own endpoint's behaviour, because the README's description of what counts as retryable stops mid-sentence.

## FAQ

### Does ElevenLabs have an API?

Yes. The HTTP API documentation lives at elevenlabs.io/docs/api-reference, and this repository is the official Node SDK for it, with text to speech conversion, streaming, voice listing and a Speech Engine server.

### Is ElevenLabs open source?

This SDK is MIT licensed and open source, version 2.70.0. The README does not state the licensing of the service behind the API, so it does not answer the question for the platform itself.

### Why does play() fail on my server?

The SDK requires MPV and ffmpeg to play audio, and both are system binaries rather than npm packages. On a minimal container, use the streaming method and hand the audio stream to the client instead.

### What happens if I set disableAuth on a Speech Engine server?

The SDK skips verification of the X-Elevenlabs-Speech-Engine-Authorization header and accepts any client that can reach the server, logging a console.warn on startup. The README says to use it only behind an IP allowlist scoped to ElevenLabs' egress ranges.

## Sources

- [elevenlabs/elevenlabs-js on GitHub](https://github.com/elevenlabs/elevenlabs-js)
- [License: MIT](https://github.com/elevenlabs/elevenlabs-js/blob/main/LICENSE)
- [Project website](https://elevenlabs.io)
- [README](https://github.com/elevenlabs/elevenlabs-js/blob/main/README.md)
- [Releases](https://github.com/elevenlabs/elevenlabs-js/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/elevenlabs-elevenlabs-js
