Library / SDK
huggingface/transformers.js avatar
huggingface/transformers.js

Transformers.js: Running Hugging Face Models in the Browser with WebGPU

State-of-the-art Machine Learning for the web. Run 🤗 Transformers directly in your browser, with no need for a server!

16,327 stars1,195 forksJavaScriptApache-2.0

At a glance

What is it?
Transformers.js ports the Python transformers pipeline API to JavaScript and runs ONNX models locally via WASM or WebGPU. It fits client-side inference and static sites, but it is not a server-side training or batch inference stack.
Who is it for?
Adopt Transformers.js when inference has to run on the user's device, the model is already published in ONNX, and WebGPU or WASM coverage matches your browser matrix. Do not adopt it for training, fine-tuning, or large server-side batch jobs; the library is an inference runtime for the web.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Transformers.js Is For

The README states the goal plainly: run Transformers directly in the browser, with no need for a server. The library is designed to be functionally equivalent to Hugging Face's Python transformers library, so the same pretrained models run behind a very similar API. That equivalence is the selling point. If you already know pipeline('sentiment-analysis') in Python, you know the JavaScript version.

The audience is web developers who need machine learning output without standing up inference infrastructure. Text classification, named entity recognition, question answering, summarization, translation, text generation, image classification, object detection, segmentation, depth estimation, automatic speech recognition, audio classification, text-to-speech, and multimodal tasks such as zero-shot image classification are all listed as supported modalities. The common thread is that each of these is a forward pass over a pretrained model. Nothing in the README describes training or fine-tuning, and that is the correct boundary to keep in mind: this is an inference library.

The practical consequence is a different cost model. Model weights download to the client, inference runs on the client's hardware, and the server only serves static files. That is attractive for privacy-sensitive input and for sites that cannot afford per-request GPU bills. It is unattractive when the client is a low-end phone, because the same weights still have to be fetched and executed.

How the ONNX Runtime and Pipeline Layer Fit Together

Transformers.js uses ONNX Runtime to run models in the browser. That single sentence explains most of the architecture. The library does not implement its own tensor kernels for the web; it delegates execution to ONNX Runtime's web builds, which the README describes as precompiled WASM binaries served from a CDN by default.

The pipeline API is the layer above. According to the README, pipelines group a pretrained model together with preprocessing of inputs and postprocessing of outputs. So pipeline('sentiment-analysis') resolves a model, tokenizes the input text, runs the ONNX graph, and maps the raw logits back to labels and scores. The README's own example returns [{'label': 'POSITIVE', 'score': 0.999817686}], which is the same shape the Python library produces.

Model conversion is the other half of the data flow. Pretrained PyTorch, TensorFlow, or JAX models are converted to ONNX using Optimum, and the resulting weights live on the Hugging Face Hub. At runtime the library fetches them from there. Two execution backends are documented: WASM, which the README says is the default in the browser and runs on the CPU, and WebGPU, selected with device: 'webgpu'. The dtype option controls quantization, with fp32 described as the default for WebGPU and q8 as the default for WASM. That pairing matters. A quantized model on WASM and a full-precision model on WebGPU are different bandwidth and memory profiles, and the defaults already reflect the constraints of each backend.

Installing Transformers.js and Running a First Pipeline

The README gives two installation paths. The npm package is @huggingface/transformers, and the install command is a single line. Run it in your project root:

bash
npm i @huggingface/transformers

If you do not want a bundler at all, the README shows a vanilla ES module import from jsDelivr, pinned to the 4.3.0 release. Paste this into an HTML file and open it in a browser:

html
<script type="module">
    import { pipeline } from 'https://cdn.jsdelivr.net/npm/@huggingface/[email protected]';
</script>

The first real use is the sentiment pipeline. The README's example allocates the pipeline, then calls it with a string. Note that both calls are awaited, because model loading is asynchronous and the weights are fetched over the network on first use:

javascript
import { pipeline } from '@huggingface/transformers';

const pipe = await pipeline('sentiment-analysis');
const out = await pipe('I love transformers!');
// [{'label': 'POSITIVE', 'score': 0.999817686}]

What you should see is an array with a label and a numeric score. To pick a specific model instead of the pipeline default, pass a model id as the second argument. The README's example is a multilingual sentiment model:

javascript
const pipe = await pipeline('sentiment-analysis', 'Xenova/bert-base-multilingual-uncased-sentiment');

For GPU execution, add the device option. The README uses a DistilBERT sentiment model here, and warns that the WebGPU API is still experimental in many browsers, with a bug report template specifically for WebGPU errors:

javascript
const pipe = await pipeline('sentiment-analysis', 'Xenova/distilbert-base-uncased-finetuned-sst-2-english', {
  device: 'webgpu',
});

Quantization is a separate option. The README notes that in resource-constrained environments such as browsers it is advisable to use a quantized model, and that typical dtype choices include fp32, fp16, q8, and q4:

javascript
const pipe = await pipeline('sentiment-analysis', 'Xenova/distilbert-base-uncased-finetuned-sst-2-english', {
  dtype: 'q4',
});

That is the whole first-run loop: install, import, allocate a pipeline, call it, and optionally move execution to the GPU or shrink the weights.

WebGPU Is Experimental, and the Defaults Are Not Interchangeable

The README carries a warning block directly under the WebGPU example: the WebGPU API is still experimental in many browsers, and users who hit problems are directed to a bug report template titled for WebGPU errors. That is the library's own framing, and it should shape how you plan. A WebGPU path is not a drop-in replacement for the WASM path across every browser your users have installed. If your support matrix includes browsers where WebGPU is unavailable or behind a flag, you need a WASM fallback, and the two paths have different default dtypes.

Quantization is the second constraint. The README says the available dtype options may vary depending on the specific model. So dtype: 'q4' is not a universal switch you can apply to any model id. You have to check what the model actually publishes before assuming a quantized variant exists.

There is a third, quieter limitation in the repository layout. The top-level package.json is named @huggingface/transformers-monorepo, is marked private, and pins [email protected] as the package manager, with build and test scripts that run recursively across packages. That tells you the repository is a monorepo for the library plus framework integrations, not a single flat package. If you intend to build from source rather than install from npm, you are working with a pnpm workspace, not a plain npm project, and the README's install instructions do not cover that workflow.

Finally, model loading is a network operation on the client. The README describes hosted pretrained models and precompiled WASM binaries that work out of the box, and a Custom usage section for overriding both, but the section is truncated in the README text available here. What is documented is that defaults point at the Hub and a CDN; what is not documented is offline or air-gapped operation.

Transformers.js Compared with TensorFlow.js

The most natural alternative for browser machine learning is TensorFlow.js, and the difference is not cosmetic. TensorFlow.js is a general numerical and neural network library: you define layers, compile a model, train it in the browser or in Node, and load weights you produced yourself. Transformers.js is scoped to running pretrained Transformer models. It does not expose a layer API, and the README describes no training loop.

The second difference is the model source. Transformers.js consumes models converted to ONNX with Optimum and hosted on the Hugging Face Hub, and its API is deliberately aligned with the Python transformers library. TensorFlow.js consumes TensorFlow SavedModel, Keras, and GraphDef formats from the TensorFlow ecosystem. If your models already exist as ONNX exports of Transformer checkpoints, Transformers.js matches that pipeline directly. If your models are custom architectures you trained in TensorFlow, TensorFlow.js is the shorter path, and converting them to ONNX just to use Transformers.js would add a step with no benefit.

The third difference is task coverage. Transformers.js ships pipelines for named tasks, so a sentiment classifier is one call. In TensorFlow.js you would assemble tokenization and postprocessing yourself, or find a model package that already does it. That is a real convenience gap, and it is the main reason to pick Transformers.js when the task is one of the listed pipelines.

Licence and Versioning

The repository is licensed Apache-2.0, stated in the README badge, the LICENSE file at the top level, and the license field of the root package.json. Apache-2.0 is a permissive licence that includes an explicit patent grant, which matters for a library that may end up in commercial web applications. That is a description of the licence text, not legal advice; if your organisation has specific obligations around attribution or notice files, have counsel review how you distribute the bundled library.

The licence covers the library code. It does not automatically cover the model weights you load. Those come from the Hugging Face Hub, and each model repository carries its own licence, which the README does not discuss. If you ship a product that downloads a specific checkpoint, the checkpoint's terms are a separate question from the Apache-2.0 grant on the runtime.

On releases, the repository lists 4.3.0 on 2026-09-16, 4.2.0 and 4.1.0 both on 2026-04-23, and the last push to the repository on 2026-09-18. The CDN example in the README pins the version explicitly, which is the right pattern for a browser import: an unpinned CDN URL means your users get whatever ships next. The npm install path gives you the usual lockfile control. The README does not document a deprecation or migration policy between major versions, so if you are moving from an older major line, check the release notes for that version rather than assuming API stability.

Where Transformers.js Is the Wrong Tool

It is the wrong tool when the model does not fit the client. The README's own advice to use quantized models in resource-constrained environments is an acknowledgement that full-precision weights are heavy for browsers. If your task needs a large generative model that a phone cannot hold in memory, moving inference to a server is the correct answer, and Transformers.js has nothing to offer there.

It is also the wrong tool for training and fine-tuning. The README describes converting pretrained PyTorch, TensorFlow, or JAX models to ONNX, and running them. There is no optimizer, no gradient API, and no dataset abstraction in what is documented. If your workflow involves adapting a model to your own data, that work happens outside this library, in Python, and the result arrives as an ONNX artifact.

A third case is a browser matrix that cannot guarantee WebGPU. Because the README labels WebGPU experimental in many browsers and defaults to WASM on the CPU, a compute-heavy model on WASM may be slower than your users will tolerate. That is not a defect in the library; it is the shape of client-side inference. But it means the decision to run in the browser should be made per task, not as a blanket policy.

Editorial conclusion

Adopt Transformers.js when inference has to run on the user's device, the model is already published in ONNX, and WebGPU or WASM coverage matches your browser matrix. Do not adopt it for training, fine-tuning, or large server-side batch jobs; the library is an inference runtime for the web. Before shipping, verify that your chosen model id resolves on the Hub, that your target browsers enable WebGPU, and that the dtype you select is available for that model.

Frequently asked questions

How do I use Transformers.js?

Install the package with npm i @huggingface/transformers, then import the pipeline function and await a pipeline for your task, such as pipeline('sentiment-analysis'). Call the resulting function with your input and await the result. The README also shows a vanilla ES module import from jsDelivr if you do not want a bundler.

What is Transformers.js?

It is a JavaScript library that runs pretrained Transformers models directly in the browser with no server, using ONNX Runtime. The README states it is designed to be functionally equivalent to the Python transformers library, so the same models run behind a very similar API.

Is Transformers.js free?

The repository is licensed Apache-2.0, which is a permissive open source licence. The library code is covered by that licence; the model weights you load from the Hugging Face Hub carry their own separate licences, which the README does not address.

Transformers.js vs TensorFlow.js: what is the difference?

Transformers.js is scoped to running pretrained Transformer models converted to ONNX, with built-in pipelines that handle preprocessing and postprocessing for named tasks. TensorFlow.js is a general neural network library in the TensorFlow ecosystem, which the README does not discuss; Transformers.js documents no layer API and no training loop.

Official sources

  1. huggingface/transformers.js on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huggingface-transformers-js.svg)](https://hysenlabs.com/projects/huggingface-transformers-js)