gpu.js: transpiling JavaScript kernels to GLSL, WGSL and the CPU fallback
GPU Accelerated JavaScript
At a glance
- What is it?
- gpu.js turns a plain JavaScript function into a GPU kernel and runs it on WebGL, WebGPU or, when no GPU is present, in ordinary JavaScript. The library works, the async migration does not, and the v3 breaking change should decide whether you start here.
- Who is it for?
- Adopt gpu.js if you are shipping a browser visualization or a Node pipeline where a single 512x512 elementwise or matrix kernel is the bottleneck and you can tolerate the readback stall. Do not adopt it for new code that will live past the v3 release unless you write every kernel against mode: 'async' from the first commit, and do not adopt it if you need WebGPU today, since the README treats that backend as the reason for the breaking change rather than a finished option.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem gpu.js solves, and the JavaScript developers it targets
The library exists for one narrow situation: you have a numeric loop in JavaScript, it is the slow part of the page or the process, and rewriting it in CUDA, OpenCL or a native addon is more work than the speedup justifies. gpu.js takes a JavaScript function that computes a single element of an output array, transpiles it into shader language, and runs the whole grid of invocations on the GPU in parallel. The README describes the result as "1-15x faster depending on your hardware", which is a wide band and an honest one: the gain depends on how parallel the loop is and how much data has to come back.
The audience is the browser first. The browser build ships as dist/gpu-browser.min.js and is available from unpkg and jsdelivr, so a page can load it with a script tag and start issuing kernels without a build step. Node is the second audience, and there the same kernel source runs through the same API. TypeScript users get type definitions with the package, and the examples directory carries both simple-javascript.js and advanced-typescript.ts alongside visual demos such as mandelbrot-set.html and raytracer.html.
What it is not is a tensor framework. There is no autograd, no model format, no operator library. You supply the arithmetic; gpu.js supplies the execution.
How a JavaScript function becomes a shader: createKernel, setOutput and this.thread
The mechanism is a transpiler plus a runtime. When you call gpu.createKernel(fn), gpu.js parses the function source (the package depends on acorn 8.14.0, and the source tree lives under src/) and emits shader code. The function you write is not called per element by the CPU; it is the body of the shader that the GPU executes once per output element. Inside it, this.thread.y and this.thread.x give the coordinates of the element currently being computed, which is why the README's matrix multiplication reads a[this.thread.y][i] * b[i][this.thread.x].
setOutput([512, 512]) declares the shape of the result, and the shape determines how many GPU invocations are launched. The transpiler only understands a subset of JavaScript: loops, arithmetic, array indexing, conditionals and a few math functions. Anything outside that subset either fails at compile time or silently falls back, which is the first thing to watch when a kernel misbehaves.
The execution path after compilation depends on the backend. On WebGL, kernels run as fragment shaders and results come back through readPixels. The README is unusually blunt about the cost of that call, stating that it "silently freezes the page until the GPU catches up" and that in one measurement a readback-heavy loop spent roughly 96% of its time frozen. That is the architectural fact that shapes everything else in the project, including the v3 plan.
Installing gpu.js in Node and running a first matrix kernel
In Node the package installs from npm under the name gpu.js, and the CommonJS entry point is src/index.js. The package declares an engine requirement of Node >=8.0.0, and gl 8.1.6 is listed as an optional dependency, which is how the Node build reaches a GPU context on a headless machine. If that optional dependency does not build on your platform, the library still loads and the kernels run in JavaScript instead.
Install it and require the GPU constructor:
npm install gpu.jsThe README's Node example builds a 512x512 matrix multiplication kernel. The function body computes one output element by walking a row of a and a column of b, and setOutput declares the 512x512 result shape:
const { GPU } = require('gpu.js');
const gpu = new GPU();
const multiplyMatrix = gpu.createKernel(function(a, b) {
let sum = 0;
for (let i = 0; i < 512; i++) {
sum += a[this.thread.y][i] * b[i][this.thread.x];
}
return sum;
}).setOutput([512, 512]);
const c = multiplyMatrix(a, b);c comes back as a 512x512 array. The same kernel is shown for the browser, where you load dist/gpu-browser.min.js and construct new GPU() from the global namespace, and for TypeScript, where the kernel signature is annotated as (a: number[][], b: number[][]) and the result is cast to number[][]. The README also points to the examples directory for further TypeScript cases, including examples/advanced-typescript.ts.
One detail worth knowing before you benchmark anything: if no GPU is available, the documentation states the functions still run in regular JavaScript. The API does not change and no error is raised, so a kernel that looks fast on your laptop may be running on the CPU on a colleague's machine.
The synchronous readback stall, and why v3 makes every kernel call return a Promise
The most consequential thing in the README is not a feature. It is a warning: the next major version makes every kernel call return a Promise, and synchronous kernel calls written today will not survive the upgrade unchanged. The project's own explanation is that the synchronous design was a mistake, that WebGL let it pretend a GPU is a synchronous device, and that the pretense is paid for on every readback.
That has two practical effects. First, if you are starting a project now, the synchronous examples in the README are the ones that will break. Second, the migration path already exists: mode: 'async' was added in 2.20.0, and the README states that code written against it will run on v3 unchanged. The migration guide in the README is described as six steps. If you are writing new kernels, the honest advice is to write them async from the start rather than migrating later.
The escape hatch is setAsyncMode(false), which the WebGL backends keep through the migration. The README states plainly that WebGPU can never offer a synchronous mode. So the compatibility story is asymmetric: WebGL users can stay synchronous for a while, WebGPU users cannot, and any code that mixes the two has to be async.
WebGPU accuracy claims versus the WebGL path you are probably using today
The README makes a strong case that WebGPU is not just faster but more correct, and the accuracy argument is more interesting than the speed numbers. Three specific claims stand out.
In precision: 'unsigned' mode, the README states that every float entering or leaving a WebGL kernel is encoded into an 8-bit-per-channel RGBA pixel and decoded on the far side. That is a quantizing round trip, and it means the values you get back are not the values the kernel computed. WebGPU kernels read and write raw IEEE-754 f32 storage buffers instead, so there is no encode step. If your results look slightly wrong in a way that tracks precision settings, this is the mechanism to check.
The second claim concerns integers. GLSL fragment shaders emulate integers with floats, and the README says this is why the fixIntegerDivisionAccuracy setting exists: some GPUs return values like 2.999 for 9/3, and the library patches around them card by card. WGSL has real 32-bit integers. A per-card workaround list is a maintenance liability, and moving off it is a real gain.
The third is addressing. Fragment-shader kernels locate data through float texture-coordinate arithmetic, which the README identifies as the source of a family of large-array off-by-one bugs. A WGSL kernel indexes its buffer with an integer. If you have ever had a kernel that is correct at 512x512 and wrong at 4096x4096, that is the class of bug being described.
The performance numbers in the README come from the project's own measurement on an Apple M1 Max against its own WebGL2 backend: 1024x1024 matrix multiplication at 6.3 ms versus 18.5 ms, and a three-kernel pipeline at 9.0 ms versus 36.6 ms. Those are the maintainers' figures on one machine. They are not a general benchmark, and the README does not present them as one.
Where gpu.js is the wrong tool, and what to reach for instead
gpu.js is the wrong choice when your computation is not a dense parallel grid. Sparse graph work, branch-heavy code, string processing and anything that needs dynamic memory allocation inside the kernel do not map onto the model. The transpiler accepts a subset of JavaScript, and the further your code drifts from arithmetic over typed arrays, the more likely you are to hit that boundary.
It is also the wrong choice if your data is small. The readback cost is fixed per kernel call, so a 64x64 kernel can easily be slower than the equivalent loop. The README's own speed range, 1-15x, is a reminder that the low end is close to no gain at all.
The obvious alternative for numeric work in JavaScript is TensorFlow.js, which approaches the same hardware from the opposite direction. TensorFlow.js gives you a fixed operator library (matmul, conv2d, reductions) and a graph or eager execution model, with backends for WebGL, WebGPU and Node. You do not write the kernel; you compose operations the library already implements. gpu.js is the inverse: you write arbitrary element functions and the library compiles them, but you get no operator library, no autograd, and no model tooling. If your problem is already expressible as tensor operations, TensorFlow.js will get you there with less code. If your problem is a custom elementwise formula that no operator library exposes, gpu.js is the one that lets you write it directly.
A second alternative is to skip JavaScript entirely and write a native kernel, which removes the transpiler's subset restrictions and the readback design at the cost of a build toolchain and a much larger change to your deployment.
Licence, maintenance and the upgrade cost of a v3 breaking change
The repository is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and licence text are retained. That is the permissive end of the spectrum, and it means the main licence question is not whether you may use it but whether you are willing to depend on a package whose API contract is about to change. This is not legal advice; check the LICENSE file in the repository for the exact terms.
On maintenance, the last push to the develop branch was on 2026-08-05, and version 2.24.0 was released the same day, with 2.23.0 and 2.22.0 landing two days earlier. The repository is not archived. The default branch is develop, not main, which is worth knowing if you pin to a branch rather than a release.
The upgrade cost is the real budget item. v3 makes every kernel call asynchronous, so every call site that consumes a kernel result has to become await-aware, and any code that assumed the result was ready on the next line has to be restructured. The README frames this as a six-step migration and provides mode: 'async' as a way to validate the change on the current version. The practical sequence is to switch your kernels to mode: 'async' on 2.24.0, confirm your tests pass, and only then take the v3 upgrade. Doing it the other way round means debugging two changes at once.
The dependency footprint is small but not zero: acorn, gpu-mock.js and webgpu are direct dependencies, with gl as an optional one. The webgpu package is listed as 0.2.9, a pre-1.0 version, which is worth noting if your organisation has policies about pre-1.0 dependencies.
Editorial conclusion
Adopt gpu.js if you are shipping a browser visualization or a Node pipeline where a single 512x512 elementwise or matrix kernel is the bottleneck and you can tolerate the readback stall. Do not adopt it for new code that will live past the v3 release unless you write every kernel against mode: 'async' from the first commit, and do not adopt it if you need WebGPU today, since the README treats that backend as the reason for the breaking change rather than a finished option. Before committing, verify which backend your target machines actually select, and run your own kernel under mode: 'async' to see whether the Promise contract changes your call sites.
Frequently asked questions
How do I install gpu.js and run a kernel in Node?
Install it from npm as gpu.js, then require the GPU constructor and call gpu.createKernel with a function that computes one output element, followed by setOutput to declare the result shape. The README's Node example builds a 512x512 matrix multiplication kernel and calls it with two input arrays.
Does gpu.js still work if the machine has no GPU?
Yes. The README states that in case a GPU is not available, the functions will still run in regular JavaScript, so the API is unchanged but the execution is on the CPU.
What does gpu.js use for the WebGPU backend?
The package lists webgpu 0.2.9 as a direct dependency, and the README describes WebGPU as a first-class backend that reads and writes raw IEEE-754 f32 storage buffers instead of the packed RGBA encoding used in WebGL unsigned precision mode.
Why is the next major version of gpu.js asynchronous?
The README states that v3 will make every kernel call return a Promise because a GPU is an asynchronous device and the synchronous WebGL readback path blocks the main thread. Code written against mode: 'async', added in 2.20.0, will run on v3 unchanged.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gpujs-gpu-js)