Model or dataset
sauravpanda/BrowserAI avatar
sauravpanda/BrowserAI

BrowserAI: running LLMs, Whisper and Kokoro inside the browser tab

Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser

1,451 stars138 forksTypeScriptMIT

At a glance

What is it?
BrowserAI is an MIT-licensed TypeScript library that loads MLC, Transformers, Flare and Demucs engines in the browser and exposes them behind one API. It fits privacy-sensitive web apps; it does not fit teams that need server-side throughput or documented upgrade paths.
Who is it for?
Adopt BrowserAI if your product needs on-device inference and you can accept WebGPU-only coverage, a fixed model catalogue and a README that does not document upgrade or rollback paths. Do not adopt it for server-side batch inference, for browsers without WebGPU, or where you need to bring an arbitrary checkpoint.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 72 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What BrowserAI actually solves for a web developer

The usual way to put an LLM in a web app is to call a hosted endpoint. That means prompt text leaves the device, you pay per token, and the app stops working when the network or the provider does. BrowserAI takes the opposite route: the model weights are downloaded once and inference runs in the page. The README states the processing is local and that the app is offline capable after the initial download. For a developer building a note-taking tool, a support widget or an internal document assistant where the text must not leave the machine, that is the whole argument.

The intended audience is narrow and stated: web developers building AI-powered applications, companies that need privacy-conscious AI, researchers experimenting with browser inference, and no-code platform builders. The README also points at Browseragent.dev, a no-code agent builder described as powered by BrowserAI, which is the clearest signal of who the maintainers expect to consume the library.

The cost of this approach is paid by the user's machine, not yours. There is no server bill, and there is also no way to hide a slow device behind a bigger instance. That trade is the product.

Four engines behind one loadModel and generateText call

The mechanism is a thin facade over several runtimes. A BrowserAI instance is constructed, a model is loaded by string name, and generateText is called. Internally the library picks an engine: MLC through @mlc-ai/web-llm, Transformers through onnxruntime-web, the Flare engine for GGUF files via WASM, and Demucs for audio source separation. The README describes switching between MLC, Transformers, Flare and Demucs engines as a feature, and the supported model list is grouped by engine, so the engine is effectively chosen by which model name you pass.

That design keeps the call site stable while the runtime underneath changes. The same generateText signature serves a chat completion, a JSON-schema-constrained response, and a GGUF model loaded through Flare. The Demucs path is the exception: it is imported from a separate entry point, @browserai/browserai/demucs, and exposes loadModel and separate rather than generateText. The package.json exports map confirms the split, with a "./demucs" subpath alongside the root export.

Model loading is asynchronous and reports progress through an onProgress callback, which matters because a multi-gigabyte download with no feedback looks like a hung page. The README also mentions Web Worker support for non-blocking UI performance, though the quick-start example does not show a worker being created.

Installing BrowserAI and generating your first response

The package is published as @browserai/browserai. The README gives npm and yarn as the two install paths.

bash
npm install @browserai/browserai

After that, construct a BrowserAI instance, load a model by name, and generate. The README's basic example uses llama-3.2-1b-instruct with the quantization q4f16_1 and logs loading progress as a percentage. The response object follows an OpenAI-like shape, so the text is at response.choices[0].message.content.

javascript
import { BrowserAI } from '@browserai/browserai';

const browserAI = new BrowserAI();

await browserAI.loadModel('llama-3.2-1b-instruct', {
  quantization: 'q4f16_1',
  onProgress: (progress) => console.log('Loading:', progress.progress + '%')
});

const response = await browserAI.generateText('Hello, how are you?');
console.log(response.choices[0].message.content);

Generation options are passed as a second argument. The README shows temperature, max_tokens and system_prompt, and a messages array with role and content can be passed instead of a plain string when you want a system turn.

javascript
const response = await browserAI.generateText('Write a short poem about coding', {
  temperature: 0.8,
  max_tokens: 100,
  system_prompt: "You are a creative poet specialized in technology themes."
});

The repository ships a set of runnable examples under examples/, including chat-demo, voice-demo, tts-demo, demucs-demo, database-demos, realtime-chat-demo, schema-llm and benchmark. Those directories are the practical starting point for a first integration, since the README itself only sketches the API. There is also a live chat demo at chat.browserai.dev if you want to see the result before installing anything.

Structured output, speech and stem separation in the same library

Three capabilities sit outside the plain chat path and are worth weighing separately.

The first is JSON-schema-constrained generation. Passing a json_schema plus response_format of json_object asks the model to emit a shape you define. The README's example builds an object containing a colors array, each entry with a name and a hex string. This is the feature that makes the library usable for tool calls and form filling rather than free text, and it is the one most likely to be sensitive to which engine and model you picked.

The second is speech. Whisper-tiny-en, Whisper-base-all and Whisper-small-all handle transcription, with a built-in recorder through startRecording and stopRecording, and transcribeAudio accepting return_timestamps and language. Kokoro-TTS handles the other direction through textToSpeech, with voice and speed options; the README example uses the voice af_bella at speed 1.0 and decodes the returned buffer with the Web Audio API. The TTS demo is described as powered by Kokoro 82M.

The third is Demucs, which separates an AudioBuffer into drums, bass, other and vocals. The separate call takes shifts and overlap, and the README notes that higher shifts mean better quality and slower processing. That is the clearest performance trade-off stated anywhere in the documentation, and it applies to a model that is doing real signal processing, not text.

Where BrowserAI stops being the right tool

WebGPU is the first boundary. The README presents WebGPU acceleration as the reason inference is fast, and there is no documented CPU fallback path. If your users are on browsers or hardware without WebGPU, the quick-start example is not something you can rely on. The README does not publish a browser support matrix, so that verification is on you before you promise anything to a product team.

Model choice is the second boundary. The supported list is fixed and grouped by engine: MLC covers Llama, Hermes, SmolLM2, Qwen, Gemma, TinyLlama, Phi and the DeepSeek-R1 distills, plus Snowflake Arctic embedding models. Transformers covers Llama-3.2-1b, the Whisper sizes and Kokoro. Flare covers a short GGUF list. If your model is not there, the README says to request it by opening an issue, which is a maintenance dependency rather than a self-service path. The Flare engine does expose loadAdapter for a LoRA adapter from a URL, which is the one documented way to push a model beyond its base weights.

Upgrade cost is the third. The releases list shows v2.0.2 in April 2025, v2.0.4 in May 2025, then v2.2.0 in April 2026, and the last push to the repository was on 2026-07-21. The README does not document migration steps between major versions, nor a rollback procedure. A library that swaps inference runtimes underneath a stable API is exactly the kind where an undocumented breaking change is expensive, and the documentation is silent on it.

Finally, this is not a server-side inference library. If your workload is batch transcription of a thousand files, or a shared model serving many concurrent users, running the model in each user's tab is the wrong shape entirely.

How BrowserAI differs from Transformers.js and from hosted APIs

The closest comparison in the browser-inference space is Transformers.js, which is built on onnxruntime-web and targets Hugging Face models. BrowserAI also depends on onnxruntime-web, and its Transformers engine covers Whisper and Kokoro through that path, so the two overlap directly on speech models. The difference is scope. Transformers.js is a general runtime for ONNX models; BrowserAI is a curated catalogue with a fixed set of pre-configured models and an engine-selection layer on top. If you want to convert and run an arbitrary ONNX model, the general runtime gives you that and BrowserAI does not, except through the Flare GGUF route and LoRA adapters.

The other comparison is a hosted API. A hosted endpoint gives you a model far larger than anything a browser will download, consistent latency, and no dependency on the client's GPU. What it cannot give you is the property BrowserAI is built around: the prompt never leaves the device, and the app keeps working offline after the first download. Those are different products for different constraints, and picking BrowserAI means accepting a smaller model and a client-side performance ceiling in exchange for that property.

Licence, maintenance and what to check before you ship

BrowserAI is MIT licensed, and package.json declares "license": "MIT" with publishConfig access set to public. MIT is permissive: it allows commercial use and modification, and it requires that the copyright notice and permission notice be preserved. The repository also carries an ATTRIBUTION.md file, which is worth reading because the library bundles third-party runtimes, including @mlc-ai/web-llm, @sauravpanda/flare, onnxruntime-web, tesseract.js, phonemizer, pdfjs-dist and mammoth. Those dependencies carry their own licences, and model weights carry their own terms separately from the library. That is a question for your own legal review, not something this article can settle.

On maintenance, the facts are limited. The repository is not archived. The last push was on 2026-07-21, which is recent enough that the project cannot be described as abandoned, but the release cadence shows a long gap between v2.0.4 in May 2025 and v2.2.0 in April 2026. A single maintainer is named as the author. Treat the project as one that moves in bursts rather than continuously, and pin your dependency version accordingly.

Before shipping, verify three things: that the browsers you support have WebGPU, that every model you plan to load appears in the supported list for the engine you expect, and that the docs site at docs.browserai.dev covers the upgrade path the README does not.

Editorial conclusion

Adopt BrowserAI if your product needs on-device inference and you can accept WebGPU-only coverage, a fixed model catalogue and a README that does not document upgrade or rollback paths. Do not adopt it for server-side batch inference, for browsers without WebGPU, or where you need to bring an arbitrary checkpoint. Before committing, verify the WebGPU support matrix of the browsers you target, confirm the model you need is in the supported list or reachable through the Flare engine, and check the docs site at docs.browserai.dev for anything the README omits.

Frequently asked questions

What is BrowserAI and who is it for?

It is a TypeScript library that runs AI models directly in the browser, with WebGPU acceleration and no server component. The README lists web developers, privacy-conscious companies, researchers and no-code platform builders as its intended users.

How do I use BrowserAI in a web app?

Install @browserai/browserai, construct a BrowserAI instance, call loadModel with a model name such as llama-3.2-1b-instruct, then call generateText. The README's basic example reads the output from response.choices[0].message.content.

Is BrowserAI free to use?

The library is MIT licensed and published publicly as @browserai/browserai, so there is no licence fee for the code itself. Model weights are downloaded to the user's browser and may carry their own terms.

Is BrowserAI safe for private data?

The README states that all processing happens locally in the browser and that the library is offline capable after the initial model download, so prompts are not sent to a server by the library. That describes the library's design, not your own application's network behaviour.

Is there a BrowserAI alternative for running models in the browser?

Transformers.js is the closest alternative, built on onnxruntime-web and aimed at running arbitrary ONNX models. BrowserAI also uses onnxruntime-web for its Transformers engine but ships a fixed, pre-configured model catalogue instead of a general conversion path.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. sauravpanda/BrowserAI on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sauravpanda-browserai.svg)](https://hysenlabs.com/projects/sauravpanda-browserai)