WebLLM: running an LLM inside the browser tab with WebGPU
High-performance In-browser LLM Inference Engine
At a glance
- What is it?
- WebLLM is an npm package that runs quantized language models client-side through WebGPU, with an API shaped like OpenAI's chat completions. It removes the inference server, and with it any fallback for browsers that lack WebGPU.
- Who is it for?
- Adopt WebLLM when the model has to live on the user's machine and the audience is on a WebGPU-capable browser; skip it when you need an API key, a server-side audit trail, or support for older devices. Before committing, open https://chat.webllm.ai/ on a representative low-end machine, then run examples/get-started locally and watch the reported download size and time-to-first-token for the model you intend to ship.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The problem WebLLM removes: the inference server
Most chat features are a thin client in front of a GPU somewhere else. WebLLM deletes that tier. The README describes it as a "high-performance in-browser LLM inference engine" that runs "directly onto web browsers with hardware acceleration", with "no server support". The audience is therefore narrow and specific: teams that want the model to execute on the user's device, and are willing to accept the constraints that come with it. The stated motivation is privacy plus GPU acceleration, which is a coherent pairing. If prompts never leave the tab, there is no prompt log to leak, and no per-token bill to meter. The companion project MLC LLM handles the same models on other hardware; WebLLM is the browser branch of that effort. The cost of the design is that the model has to be downloaded to the client and the client has to be capable. Neither of those is negotiable, and the rest of this article is mostly about what follows from them.
WebGPU, WebAssembly and where the weights live
The engine is compiled to WebAssembly and executed on the GPU through WebGPU. The npm package depends on @mlc-ai/web-runtime, @mlc-ai/web-tokenizers and @mlc-ai/web-xgrammar, so tokenization and grammar-constrained decoding are separate compiled pieces rather than JavaScript reimplementations. Structured JSON generation is described as living "in the WebAssembly portion of the model library for optimal performance", which is a deliberate placement: constrained decoding touches the logits on every step, so doing it in JS would dominate the loop. Models are not fetched from a vendor endpoint. They are prebuilt MLC artifacts, listed in prebuiltAppConfig.model_list in src/config.ts, and cached by the browser after the first load. The README names Llama 3, Phi 3, Gemma, Mistral and Qwen among the supported families, and points to mlc.ai/models for the full catalogue. Because the model is a static asset, the first visit is a download and every later visit is a cache hit, which is a very different latency profile from a hosted API.
Installing @mlc-ai/web-llm and creating an MLCEngine
The package is published on npm under the scope @mlc-ai/web-llm. The README gives three package managers, and the version in package.json at the time of writing is 0.2.85.
npm install @mlc-ai/web-llm
yarn add @mlc-ai/web-llm
pnpm install @mlc-ai/web-llmAfter installing, import either the whole module or the single factory you need. The README shows both forms.
import * as webllm from "@mlc-ai/web-llm";
import { CreateMLCEngine } from "@mlc-ai/web-llm";There is also a CDN path for environments without a bundler. The README uses jsdelivr's esm.run endpoint, which it says works on jsfiddle.net, Codepen.io and Scribbler.
import * as webllm from "https://esm.run/@mlc-ai/web-llm";Most operations go through the MLCEngine interface, which you create and which loads the model as part of creation. The README's get-started example lives in examples/get-started, and the minimal browser demo is examples/simple-chat-js. What a reader should expect on first run is a visible download of the model artifact before any token appears; the exact size and the reported progress come from the engine's own loading callback, not from the README. The examples folder is the honest starting point here, because it contains runnable variants for the cases the README only names: examples/get-started-web-worker, examples/service-worker, examples/json-mode, examples/json-schema, examples/function-calling, examples/multi-models, examples/embeddings, examples/abort-reload, examples/resumable-generation, examples/integrity-verification and examples/get-started-latency-breakdown.
An OpenAI-compatible API in a tab that has no server
The claim worth examining is that WebLLM is "fully compatible" with the OpenAI chat completions API. In practice this means the request and response shapes match, so code written against the hosted API can be pointed at a local engine instead. The README lists streaming, JSON mode, logit-level control and seeding as the covered features, and marks function-calling as WIP. That last label matters more than the others. If your application depends on tool calling, the README itself tells you the feature is unfinished, and there is no statement about which models or which call patterns are covered. Treat the compatibility as a shape-level match for the documented features and verify each one you rely on. The advantages are real where they apply: streaming gives the same incremental rendering you would get from a hosted endpoint, seeding makes a generation reproducible, and logit-level control plus the separate xgrammar dependency is what backs the JSON schema example. The disadvantage is that compatibility of interface is not compatibility of behaviour. A 1.5B or 7B model running in a browser tab will not match a large hosted model on reasoning-heavy prompts, and no amount of API symmetry changes that.
The WebGPU requirement is a hard boundary, not a warning
This is the limitation to decide on before anything else. WebLLM accelerates inference with WebGPU, and the README describes the design as running everything inside the browser. There is no documented CPU fallback and no documented server fallback. A user whose browser does not expose WebGPU does not get a slower answer; the engine has nothing to run on. That makes the addressable audience a function of your browser support matrix rather than of your product, and it is the reason a hosted API remains the safer default for a general consumer audience. The second boundary is the model download. Weights are fetched to the client and cached there, so a first visit on a metered connection is a real cost to the user, and the README does not document a way to ship the weights alongside the app to avoid it. The third is the model list. It is a curated subset of the MLC catalogue, not the whole thing, and the README directs anyone who needs more to open an issue or compile a custom model in MLC format. Compiling is a build-pipeline commitment, not a config change.
WebLLM compared with calling a hosted inference API
The obvious alternative is the thing WebLLM imitates: a hosted endpoint such as the OpenAI API itself. The difference is where the compute and the data sit. A hosted call sends the prompt to a provider, returns tokens, and bills per token; the client needs nothing but a network connection and an API key. WebLLM sends nothing, needs WebGPU and a one-time model download, and costs no per-token fee. That trade changes which problems are solvable. Server-side rate limiting, provider-side logging, and central model upgrades are all straightforward with a hosted endpoint and awkward or absent here, because there is no server to hold the key or the audit trail. The reverse is also true: a privacy requirement that forbids sending user text off-device is easy to satisfy with WebLLM and hard to satisfy with a hosted call. A second alternative is the companion MLC LLM project, which deploys the same class of models across other hardware environments rather than inside a browser; if your users are on desktop apps or native clients, that is the sibling to look at first. Neither alternative is strictly better. The choice is about whether the model belongs to the device or to the provider.
Release cadence, Apache-2.0 and what a fork inherits
The repository is not archived, and the last push was on 2026-09-08. The release history in the repository is uneven: v0.2.82 on 2026-03-13, v0.2.83 on 2026-04-24, then a gap to v0.2.85 on 2026-09-08. That is a normal shape for a research-adjacent project, but it means pinning a version is wise if your application depends on a specific behaviour, since the interval between releases has been as long as roughly four and a half months. The licence is Apache-2.0, stated in both the README badge area and package.json. That is a permissive licence with an explicit patent grant, and it is the same licence family used across the MLC projects, which keeps the dependency graph consistent. Two practical notes rather than legal advice: the npm package publishes only the lib directory, so the examples and tests you may want to read are in the repository and not in your node_modules; and the runtime is split across @mlc-ai/web-runtime, @mlc-ai/web-tokenizers and @mlc-ai/web-xgrammar, so a fork inherits those dependencies as well as this one. Check licenses/ in the repository if your legal review needs the full set.
Editorial conclusion
Adopt WebLLM when the model has to live on the user's machine and the audience is on a WebGPU-capable browser; skip it when you need an API key, a server-side audit trail, or support for older devices. Before committing, open https://chat.webllm.ai/ on a representative low-end machine, then run examples/get-started locally and watch the reported download size and time-to-first-token for the model you intend to ship.
Frequently asked questions
What is a web LLM?
In this project's terms, it is a language model that runs inside the browser rather than on a server. WebLLM describes itself as an in-browser inference engine accelerated with WebGPU, so the model executes on the user's machine with no server support.
Is ChatGPT an LLM or NLP?
The README does not discuss ChatGPT's internals. It only notes that WebLLM is compatible with the OpenAI chat completions API and lists chatgpt among the repository topics, so nothing here settles how ChatGPT is classified.
Which LLM is best at web search?
The README makes no claim about web search quality for any model. It lists supported families such as Llama 3, Phi 3, Gemma, Mistral and Qwen, but the engine runs locally and the documentation does not describe a search integration.
Can I run an LLM in the browser with WebLLM?
Yes, provided the browser exposes WebGPU, since that is the acceleration path the README describes and no CPU or server fallback is documented. You also need to accept that the model artifact is downloaded and cached on the client.
What are the alternatives to WebLLM?
The direct alternative is a hosted API such as the OpenAI API, which WebLLM is designed to be compatible with but which keeps the compute and the prompt on a provider's side. The companion project MLC LLM covers the same models across other hardware environments instead of the browser.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mlc-ai-web-llm)
Community notes