Library / SDK
elevenlabs/elevenlabs-js avatar
elevenlabs/elevenlabs-js

elevenlabs-js: the official Node SDK for the ElevenLabs API

The official JavaScript (Node) library for the ElevenLabs API.

443 stars86 forksTypeScriptMIT

At a glance

What is it?
A TypeScript client generated with Fern that wraps text to speech, voice listing, streaming and the Speech Engine WebSocket protocol. It is the right layer if you are building on Node and wrong if you are in a browser or React Native.
Who is it for?
Adopt elevenlabs-js if your runtime is Node and you want the official client for text to speech, voice search and the Speech Engine WebSocket server, installed with npm install @elevenlabs/elevenlabs-js. Do not adopt it for browser or React Native work, where the README points to @elevenlabs/client and @elevenlabs/react instead, and do not adopt it if you only need one endpoint and would rather call the HTTP API directly.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What elevenlabs-js actually wraps, and who it is for

The repository is the official JavaScript client for the ElevenLabs API, published to npm as @elevenlabs/elevenlabs-js and written in TypeScript. The README describes it as the Node.js library and explicitly redirects browser and React users elsewhere: @elevenlabs/client for the browser SDK and @elevenlabs/react for the React SDK. That split is the first decision a reader has to make. If your code runs in a browser tab, this package is not the one the vendor points you to.

The audience is narrower than the topic list suggests. This is for server-side Node code that needs to turn text into audio, enumerate the voices available to an account, or run a voice agent loop where your own LLM produces the replies and ElevenLabs handles speech. The README's own example is a multilingual greeting string spanning Latin, Cyrillic, Devanagari, Arabic, Japanese and Korean scripts, which tells you the intended use is content that crosses languages rather than a single-locale demo.

What the package does not do is hide the API surface behind an opinionated abstraction. Methods map closely to endpoints: textToSpeech.convert, textToSpeech.stream, voices.search. If you already know the HTTP API reference, the SDK is mostly a typed transport with the same shape.

The client, the model IDs and the streaming call

The core object is ElevenLabsClient. It takes an apiKey option, and the README notes it defaults to process.env.ELEVENLABS_API_KEY, so the constructor argument is optional in practice. Text to speech is invoked with a voice ID as the first argument and an options object carrying text and modelId as the second.

The README lists three main models. eleven_multilingual_v2 is described as the recommended default for most use cases, with 29 languages and an emphasis on stability and accent accuracy. eleven_flash_v2_5 is the low-latency option at 32 languages and, per the README, 50% lower price per character. eleven_turbo_v2_5 sits between them at 32 languages. The trade-off stated in the README is latency and price against quality, and the model ID is a plain string you pass per call, so switching models does not require a different client.

Streaming uses a separate method, textToSpeech.stream, which the README says returns audio as it is generated. That distinction matters for interactive products: convert waits for a complete buffer, stream hands you something you can begin playing. For a voice agent, the streaming path is the one that keeps perceived latency down.

Voice discovery is a single call, voices.search, with no arguments in the example. The README does not describe the shape of the returned voice objects and instead points at the Get Voices entry in the HTTP API reference, which is a fair indication that the SDK does not add its own voice model on top.

Installing elevenlabs-js and making a first request

Installation is a standard npm or yarn add. The package name is scoped, so copy it exactly.

bash
npm install @elevenlabs/elevenlabs-js
# or
yarn add @elevenlabs/elevenlabs-js

Then construct the client and convert a short string. The README's example uses a specific voice ID and the multilingual model. Run this and you should get back an audio object that the exported play helper can send to your speakers.

ts
import { ElevenLabsClient, play } from "@elevenlabs/elevenlabs-js";

const elevenlabs = new ElevenLabsClient({
    apiKey: "YOUR_API_KEY", // Defaults to process.env.ELEVENLABS_API_KEY
});

const audio = await elevenlabs.textToSpeech.convert("Xb7hH8MSUJpSbSDYk0k2", {
    text: "Hello! 你好! Hola! नमस्ते! Bonjour! こんにちは!",
    modelId: "eleven_multilingual_v2",
});

await play(audio);

That play call is where the first real constraint appears. The README states that elevenlabs-js requires MPV and ffmpeg to play audio, and links to both projects. On a headless server or inside a slim container image, those binaries are usually absent, and the README does not offer a pure-JavaScript fallback for playback. If you are generating files rather than listening to them, drop the play import and handle the returned audio yourself.

Listing voices is the smallest useful second call, and it confirms your key works before you spend characters on synthesis:

ts
const voices = await elevenlabs.voices.search();

The README does not document what that object contains, so log it once and inspect the structure before writing code against it.

Speech Engine: two server shapes and one header check

The part of the README that goes beyond simple synthesis is Speech Engine, described as a way to build voice-powered AI agents with a custom LLM. The API connects to your server over WebSocket and each connection represents one conversation. You supply the LLM responses; ElevenLabs handles the speech.

There are two ways to run it. The attach path takes an existing HTTP server plus a path, for example /api/speech-engine/ws, and registers the handler alongside your current routes. The standalone path, SpeechEngine.Server, takes a port, an engineId and an apiKey, and treats every connection as a session. The README's standalone example uses port 3001 and notes that the engine's server URL points at host and port directly, for example wss://myserver.com:3001.

The callback set is where the design shows. onInit receives a conversationId. onTranscript receives the transcript, a signal and the session object. The signal is documented as auto-aborting if the user interrupts, and the callback passes it straight into the LLM request, so cancellation propagates without you writing interruption logic. session.sendResponse accepts a streaming LLM response, and the README states the SDK extracts text from OpenAI, Anthropic and Gemini stream formats automatically. That is a real convenience and also a coupling: the SDK is parsing provider-specific stream shapes on your behalf.

Both entry points verify the X-Elevenlabs-Speech-Engine-Authorization header on every incoming connection and reject requests not signed by ElevenLabs, reading the key from the apiKey option or ELEVENLABS_API_KEY. The README also documents disabling that authentication when an infrastructure layer already restricts traffic to ElevenLabs. That escape hatch is worth reading carefully before you use it, because it moves the trust boundary to whatever sits in front of your process.

Where elevenlabs-js is the wrong tool

The clearest limitation is stated by the project itself: this is the Node library, and the README sends browser and React users to different packages. Choosing elevenlabs-js for a front-end bundle means fighting the intended split, and the README gives no browser guidance for it.

The second limitation is the playback dependency. Requiring MPV and ffmpeg is fine on a developer laptop and awkward in a container, a serverless function or a CI job. Anywhere you cannot install system packages, the play helper is unusable and you are back to writing your own output handling.

Third, the SDK is generated by Fern, as the badge in the README indicates. Generated clients track the API schema closely, which is good for coverage and less good for ergonomics: methods mirror endpoints, and the README repeatedly defers structural questions to the HTTP API reference rather than documenting return types inline. If you want a curated, hand-shaped interface, this is not that.

Finally, there is a version signal worth noting. The release list includes v3.0.0-alpha.1 alongside the stable 2.x line. The README does not describe migration or compatibility between them, so anyone depending on the pre-release channel should treat the documentation as incomplete rather than assume the 2.x examples carry over.

elevenlabs-js against calling the REST API directly

The obvious alternative is the ElevenLabs REST API, which the README links as the HTTP API documentation. The difference is not capability, since the SDK is a client for that API, but what you inherit. Calling REST directly means you own request construction, auth headers, error handling, retries and the WebSocket handshake for Speech Engine. In exchange you take on no dependency beyond your HTTP client, and you can call the API from any runtime, including browsers and edge workers where a Node-oriented package is a poor fit.

The SDK's advantages are concentrated in two places. Typed method signatures and the generated surface reduce guesswork on parameter names and model IDs. The Speech Engine helpers are the stronger case: the authorization header check, the transcript callback with an interrupt-aware signal, and the automatic extraction of OpenAI, Anthropic and Gemini stream formats are all things you would otherwise write and maintain yourself.

If your usage is a single textToSpeech call in a cron job, the SDK is overhead. If you are building the agent loop, the WebSocket plumbing is exactly the part you do not want to hand-roll. The package also depends on node-fetch and ws, which is a small dependency footprint relative to what the Speech Engine path replaces.

Licence, maintenance and what an upgrade costs

The package is MIT licensed, both in package.json and per the repository metadata. That is permissive and places few obligations on commercial use, though the licence covers the client code and not the ElevenLabs service, which is a separate commercial relationship governed by your account terms. Nothing here is legal advice; read the LICENSE file and your service agreement.

The last push to the repository was on 2026-09-11, and the most recent release listed is v2.68.0 on the same date. The version number and the release cadence suggest frequent publishing, and the presence of a v3.0.0-alpha.1 release from 2026-09-03 means two lines exist at once. That is the main upgrade cost to plan for: pinning a 2.x version protects you from the alpha channel, but the README does not document a migration path between the two, so moving to 3.x will require reading release notes rather than following a guide.

The repository layout shows a Fern directory and a Fern ignore file, which is consistent with the generated-SDK badge. Practically, that means local patches to generated files are likely to be overwritten on regeneration, and contributions are better aimed at the API surface than at hand-editing output. Build tooling is TypeScript with tsc, tests run through Jest with separate unit and wire projects, and formatting and linting go through Biome. None of that affects consumers, but it tells you the maintenance burden sits with the generator plus a thin wrapper layer.

Editorial conclusion

Adopt elevenlabs-js if your runtime is Node and you want the official client for text to speech, voice search and the Speech Engine WebSocket server, installed with npm install @elevenlabs/elevenlabs-js. Do not adopt it for browser or React Native work, where the README points to @elevenlabs/client and @elevenlabs/react instead, and do not adopt it if you only need one endpoint and would rather call the HTTP API directly. Before committing, verify the model IDs your account can use, whether your deployment can install MPV and ffmpeg for local playback, and how the X-Elevenlabs-Speech-Engine-Authorization header check behaves behind your own proxy.

Frequently asked questions

How does ElevenLabs work?

The SDK sends your text and a model ID to the ElevenLabs API and returns generated audio, either as a complete buffer through textToSpeech.convert or incrementally through textToSpeech.stream. For agents, the Speech Engine opens a WebSocket where each connection is one conversation and your own LLM supplies the replies.

Is ElevenLabs open source?

The client library is MIT licensed, as stated in package.json and the repository metadata. That covers the SDK code, not the ElevenLabs service itself, which you reach through an API key.

Does ElevenLabs have an API?

Yes. The README links to the HTTP API documentation at elevenlabs.io/docs/api-reference, and elevenlabs-js is the official Node client for it.

Is there a free API key for ElevenLabs?

The README does not describe free keys or pricing tiers for API access. It only shows the apiKey option and the ELEVENLABS_API_KEY environment variable, so check the ElevenLabs site for account terms.

Official sources

  1. elevenlabs/elevenlabs-js on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/elevenlabs-elevenlabs-js.svg)](https://hysenlabs.com/projects/elevenlabs-elevenlabs-js)