wllama ships no wasm binaries, so the in-browser path starts with a docker build
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference
At a glance
- What is it?
- wllama wraps llama.cpp in WebAssembly so inference runs in a browser worker, and the manifest and the release line agree on version 3.8.1. What the repository does not hand you is the compiled binary: llama.cpp arrives as a submodule, the build runs shell scripts under docker, and the glue layer between the TypeScript API and the C++ engine is generated at build time.
- Who is it for?
- Taken as a binding, wllama is unusually explicit about its own limits: the 2GB file ceiling, the two headers that unlock threads, the model split recipe, and the single GPU dial. The gap worth planning around is the build, because the npm package carries the compiled binaries while the git route compiles llama.cpp under docker and generates its own glue code, which is a heavier setup than the feature list implies.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Inference runs in a worker, but multi-thread still needs isolation headers
The feature list says inference happens inside a worker so the UI keeps rendering, and that it needs no backend and no GPU, running through WebAssembly SIMD. One requirement sits underneath that. Multi-thread mode only works when the page is served with `Cross-Origin-Embedder-Policy` and `Cross-Origin-Opener-Policy` headers, which is what makes the shared memory the threads need available to the page. The page links a discussion on ffmpeg.wasm for the details. Without those two headers the library falls back on its own, since it switches between the single-thread and multi-thread build based on browser support, and the single-thread path can be forced by adding `{ "n_threads": 1 }` to the load config. The repo's own static server carries the matching switch: `npm run serve:mt` runs the same script with `MULTITHREAD=1` in the environment. So the only server in this design is one that sets two headers.
A 2GB ArrayBuffer ceiling is why llama-gguf-split shows up in the docs
A single model file is capped at 2GB, and the reason given is not a project limit but the size restriction on ArrayBuffer. The escape hatch is chunking, done with `llama-gguf-split` from the llama.cpp release page:
# Split the model into chunks of 512 Megabytes
./llama-gguf-split --split-max-size 512M ./my_model.gguf ./my_modelThe output names follow a fixed pattern: `-00001-of-00003.gguf`, `-00002-of-00003.gguf`, and so on. Give the first chunk to `loadModelFromUrl` or `loadModelFromHF` and the rest are picked up automatically:
const wllama = new Wllama(CONFIG_PATHS, {
parallelDownloads: 5, // optional: maximum files to download in parallel (default: 3)
});
await wllama.loadModelFromHF({
repo: 'ngxson/tinyllama_split_test',
file: 'stories15M-q8_0-00001-of-00003.gguf',
});The recommended chunk size is 512MB, and splitting is offered as a download speed win for small models too, since chunks arrive in parallel. Three files at a time is the default; the example raises it to five.
Where those bytes come from at runtime is a second decision. The plain usage example points the config at a path inside the package, `./esm/wasm/wllama.wasm`, and the constructor comment explains that the single-thread or multi-thread build is chosen from browser support at that moment. There is a CDN variant as well, `wasm-from-cdn.js` from the same package, and the comment attached to it says it is not recommended and should be used only when wasm files cannot be embedded in the project. Progress during load is reported through a callback that receives `loaded` and `total` and rounds the ratio into a percentage.
No wasm binaries in the repo: the build needs docker and a submodule
The alternative install path is a git submodule, and it opens with a warning: wasm binaries do not come pre-built with this repo, and docker has to be installed on the machine.
# recommend to clone as git submodule
git submodule add https://github.com/ngxson/wllama.git wllama
git submodule update --init --recursive
# run the build
cd wllama
npm ci
npm run build:wasm && npm run buildThe `llama.cpp` entry at the top level, together with `.gitmodules`, explains why that recursive submodule update is in the sequence: the C++ engine this TypeScript package binds to is not vendored source but a checked out dependency. `build:wasm` runs `./scripts/build_wasm.sh` and then `build:glue`, which is `node ./cpp/generate_glue_prototype.js`. The glue between the TypeScript API and the C++ functions is generated rather than committed, which makes `CMakeLists.txt` and `cpp/` build inputs instead of reference material.
n_gpu_layers is the only WebGPU dial, and the default is every layer
WebGPU arrived through pull request 215 and switched on automatically at V3.1. The default is to offload all layers to the GPU, and the single control is the `n_gpu_layers` parameter of the load options:
// (optionally) will allow running WebGPU on Firefox via compat mode; performance will be significantly degraded
wllama.setCompat('default', 'firefox_safari');
await wllama.loadModel(files, {
n_gpu_layers: 4, // meaning 4 layers are offloaded to GPU; set to 0 to disable GPU inference
});Set it to 0 and GPU inference is off; set it to 4 and four layers go across. That one number is the answer to a model that does not fit in VRAM, and the page says so directly. The compat call in the same snippet is the other half of the story: compatibility problems are pointed at a separate package, `@wllama/wllama-compat`, and the browser that needs it is named in the code comment, with the cost stated there as well.
npm test runs vitest without run, and four more entries set env vars
The test entry is `vitest`, not `vitest run`, so the plain command stays in watch mode. Four sibling scripts vary the environment instead:
"test": "vitest",
"test:auto": "AUTO=1 vitest",
"test:wgpu": "WEBGPU=1 vitest"`test:firefox` sets `BROWSER=firefox` and `test:safari` sets `BROWSER=safari`, and `vitest.config.ts` at the top level is what reads those variables. One further script rebuilds the wasm with `WLLAMA_TEST_BACKEND=1` before testing, and the tree holds a matching `examples/test-backend-ops/` directory. That directory is missing from the demo list on the page, which walks five examples: basic completions and embeddings, embeddings with cosine distance, multimodal vision, tool calling and decision models.
The publish script formats the tree and ships two packages
Publishing is one script, and it does more than publish:
"upload": "npm run format && npm run build && node scripts/check_package_size.js && npm publish --access public && (cd compat && npm publish --access public)"`format` is `prettier --write .`, so the upload path rewrites the working tree before anything is packed. `build` then runs `clean`, `build:worker`, `build:tsup`, `build:minified`, `build:typedef`, `./scripts/post_build.sh` and `docs` in that order, and `clean` deletes `./esm`, `./docs` and `./wasm`. `build:tsup` bundles `src/index.ts` to cjs and esm even though the manifest entry is `index.js` with `"type": "module"`, and `build:minified` runs terser over `esm/index.js`. The size gate sits directly in front of publish, and the `compat` directory goes out as its own package in a second publish step.
Development runs through the same family of shell scripts: `npm run serve` starts `node ./scripts/http_server.js`, and `build:worker` is a shell script of its own, while `build:glue` is a node script that reads the C++ side. `tsup.config.ts` sits next to `vitest.config.ts` at the top level, and the tree also carries `dev/`, `AGENTS.md` and `README-dev.md` for work that is not part of the published package.
The banner still points at the V3 guide while the manifest reads 3.8.1
The callout at the top of the page announces V3 with WebGPU, multimodal and tool calling, and sends readers to `guides/intro-v3.md`. The manifest in the same repository reads version 3.8.1, and the recorded releases run 3.8.1 on 2026-10-02, then 3.6.1 on 2026-08-27 and 3.6.0 on 2026-08-16. The last push to the repository is dated 2026-10-02. So the package line and the release line agree with each other, and it is the introduction banner alone that stays pinned to a much older line. Compatibility questions are redirected one level down, to `@wllama/wllama-compat` in the `compat/` directory, which is also the directory published as the second package. The licence file at the top level is spelled `LICENCE`, and the recorded licence is MIT.
Two places a reader would check first end mid sentence
The section on custom loggers sits under a heading about suppressing debug messages, begins `When initial`, and the sentence does not finish. The manifest's author field gives the name Xuan Son NGUYEN and then stops partway through the contact address that follows. Those are the two spots a reader heads to first, one for logging control and one for who to contact, and neither completes its thought in the files themselves. Elsewhere the page is unusually concrete: the feature list claims an OpenAI-compatible typed API, no runtime dependency, automatic single and multi-thread switching, worker based inference, image and audio input, tool calling, and quantized Q4, Q5 or Q6 recommended over IQ with imatrix, which is called out as slow and low quality. A CDN path exists for the wasm files and is marked as not recommended, kept for projects that cannot embed them.
Editorial conclusion
Taken as a binding, wllama is unusually explicit about its own limits: the 2GB file ceiling, the two headers that unlock threads, the model split recipe, and the single GPU dial. The gap worth planning around is the build, because the npm package carries the compiled binaries while the git route compiles llama.cpp under docker and generates its own glue code, which is a heavier setup than the feature list implies. Anyone planning browser inference should first decide whether embedding the wasm in the bundle fits their deployment, then read the compatibility package and the unfinished logger section before relying on either.
Frequently asked questions
Does wllama need a server to run inference in the browser?
Not for inference. It runs on WebAssembly SIMD inside a worker with no backend and no GPU, so nothing blocks UI render. Multi-thread mode does need Cross-Origin-Embedder-Policy and Cross-Origin-Opener-Policy headers served over HTTP, which is what npm run serve:mt sets up locally.
How large a model can wllama load in one piece?
A single file is limited to 2GB by the size restriction on ArrayBuffer. Larger models are split into 512MB chunks with llama-gguf-split, then loadModelFromUrl or loadModelFromHF is given the URL of the first chunk and the rest load automatically.
Can wllama run WebGPU on Firefox?
Through compat mode, called as setCompat('default', 'firefox_safari'), and the code comment for that line states performance will be significantly degraded. Compatibility issues are directed to the separate @wllama/wllama-compat package, and the test scripts include a BROWSER=firefox entry.
Which quantization should wllama users pick?
The page recommends quantized Q4, Q5 or Q6 as a balance among performance, file size and quality. IQ with imatrix is explicitly not recommended, since it may bring slow inference and low quality.
What does wllama add on top of llama.cpp itself?
It exposes llama.cpp through WebAssembly with an OpenAI-compatible typed API, WebGPU offloading, image and audio input, tool calling, parallel loading of split model chunks, automatic single and multi-thread switching based on browser support, and inference inside a worker. The C++ engine is checked out as a submodule rather than vendored.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ngxson-wllama)