Framework
mattn/tensai avatar
mattn/tensai

tensai: a neural-network framework where the fast path is still pure Go

A tiny neural-network framework in pure Go with AVX2 SIMD kernels (GOEXPERIMENT=simd)

111 stars5 forksGoMIT

At a glance

What is it?
A small machine-learning framework with no external dependencies in its default build, and three optional accelerations layered on top: experimental SIMD kernels, a WebGPU backend reached through dlopen, and int8 or int4 quantization. The dependency-free claim survives the SIMD path, which is written with Go's experimental simd package rather than assembly.
Who is it for?
tensai fits a Go developer who wants a model to train and serve in one binary with no Python in the pipeline, and who will read the docs because the API surface is broad and the build tags matter.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

No dependencies by default, and the module retracts its own 0.1 tags

The dependency claim is specific. The default build has no external dependencies at all. The optional wgpu build tag adds exactly one, ebitengine/purego, and it is cgo-free.

That is unusual enough to be worth checking in go.mod rather than taking on trust, and the file is small. The module is github.com/mattn/tensai, it declares go 1.22, and it requires a single package, ebitengine/purego at v0.10.2, which is pulled in for the GPU path.

The retraction block is the more interesting part, and it explains the version numbering. Two versions are retracted with a comment saying the first was tagged in error before the package split and that releases follow v0.0.x. So v0.1.0 and v0.1.1 exist in the history and are unusable, and the current line is v0.0.x. That is a project at v0.0.33 in practice, whatever the earlier tags suggest.

The repository layout matches the package structure rather than a single blob. Directories cover autograd, dataset, encoding, gpu, internal, knn, layer, loss, metrics, model, optim, quant, rnn and tokenizer, with the n-dimensional tensor and the two-dimensional Matrix at the root beside accel.go and separate dot kernels for NEON, SIMD and the generic case.

The SIMD path needs a specific Go version, and falls back silently

The acceleration story is three layers, and the first one has a version gate that is easy to miss.

AVX2 kernels on amd64 and NEON kernels on arm64 are written with Go's experimental simd/archsimd package, which keeps the promise of pure Go intact: no cgo, no assembly files. Matmul, ReLU and LeakyReLU, Sigmoid, Tanh and Softmax through a vectorized polynomial exp, GELU through a vectorized erf, LayerNorm, and the Adam update are all eight-lane vectorized on AVX2. On arm64 the coverage is narrower, and the description is honest that it covers the decode path so far.

The build requirement follows from that. On amd64 you need GOEXPERIMENT=simd with Go 1.26 or 1.27. On arm64 you need Go 1.27, described as the first release whose simd/archsimd has an arm64 half.

What happens otherwise is the part that makes this safe to use. Every other build uses the portable fallbacks automatically, and the results are identical. So the same code runs on a Raspberry Pi and on a workstation, and the only thing you lose is throughput rather than correctness. A per-kernel and per-OS breakdown lives in the platform guide in the docs.

A third path is the WebGPU backend, built with -tags wgpu on Linux, macOS or Windows, which runs batched MatMul as a WGSL compute shader on anything wgpu-native reaches, including Vulkan, Metal and D3D12. The bindings go through purego, so the wgpu-native shared library is dlopen-ed at runtime and there is still no C compiler involved.

Scratch buffers are reused, and a full MLP step is about 29 allocations

The allocation strategy is the quiet reason a pure-Go framework can be competitive. Layers reuse their forward and backward scratch buffers across training steps, and a full multilayer perceptron step runs in roughly 29 allocations. The stated effect is that the garbage collector stays out of the training loop.

That is the right target. In a training loop, allocation rate is not a memory concern, it is a latency one, because every collection interrupts a step that is supposed to be predictable. Predict is the deliberate exception, since it always returns freshly allocated results rather than reusing anything, which is the correct trade for an inference call that outlives the buffer.

The layer set is conventional rather than exhaustive. Embedding, Dense, Conv2D, MaxPool2D, BatchNorm, LayerNorm and Dropout, with ReLU, LeakyReLU, GELU, Sigmoid, Tanh and Softmax as activations. Loss functions are MeanSquaredError for regression, SoftmaxCrossEntropy for multi-class classification and BinaryCrossEntropy for binary targets. Optimizers are momentum SGD, Adam and AdamW with decoupled weight decay.

For a baseline, there is a knn.Classifier whose distance matrix runs on the same SIMD matmul kernel, which makes it a fair comparison against the networks rather than a slow one.

Two ways to build the same model, and the GPU one is the interesting path

There are two front ends over the same weights, and picking between them is the main API decision.

The Sequential form stacks layers and runs Compile, then Fit or FitStep, then Predict. The alternative is Graph(), which builds the same model as an autograd graph instead. Every layer has a graph form, so the two paths train the very same weights, and the graph form is what can run on a GPU.

The autograd engine underneath is described as micrograd-style reverse mode over n-dimensional tensors, with Param, Input and Backward as the vocabulary, for models that do not fit the Sequential shape. Values are Tensors, so the same operations run on batch, sequence and model shaped activations, and the list of gradient-carrying ops is specific: broadcasting element-wise arithmetic, batched MatMul, Transpose and Reshape, last-axis Softmax, axis reductions, LayerNorm, embedding Embed, and CrossEntropy. A Matrix is still accepted anywhere a leaf is built. ToDot renders the graph.

The device story builds on that. UseAccelerator moves every product above 4e8 multiply-accumulates onto the GPU, including both transposed products a backward pass needs, and tape.UseDevice keeps values, gradients and the Adam update resident on the device. The numbers given are for a 2048-wide block: one step takes 1382ms on the CPU, 442ms with products offloaded one at a time, and 193ms fully resident. For a whole transformer block at width 512 it is 209ms against 54ms.

int4 is the difference between a 7B model fitting in RAM and not

Quantization is the feature with the most concrete payoff stated, and the framing is about memory rather than speed.

quant.Quantize and quant.Quantize4 build weight-only quantized twins. The int4 path is group-wise with float32 accumulation. The int8 path is a full integer path: weights held in interleaved row quads, activations dynamically quantized to 7 bits, and the whole dot product running on the 256-bit u8 by s8 pairwise multiply-add plus a widening pair-add, which works out to two instructions per column, four rows deep. The stated result is that it reaches memory bandwidth, quoted as roughly 31GB/s of weights on 16 cores.

The reason that number matters is bandwidth, not arithmetic. Once a dot product is memory-bound, making the arithmetic cheaper changes nothing, and shrinking the weights is the only lever left. Going to int4 halves the weights again, and the README states the consequence in one line: it is the difference between a 7B model fitting in RAM or not.

Also present is a dataset package with Shuffle, a copy-free train and test Split, buffer-reusing mini-batch iteration through Batches, and Standardize or StandardizeWith for feature scaling.

Image generation and audio understanding are in the same binary

The framework ships two model-running commands that are larger than the training library in scope.

tensai image runs Qwen-Image-2.1 end to end in pure Go. The prompt goes through the language half of a Qwen3-VL 8B encoder, then twenty flow-matching steps through a 32-block denoising transformer, then a 2D autoencoder that turns each latent position into a sixteen-pixel square with an alpha channel. The weights quantize as they load, which is 7GB at eight bits and 3.5GB at four, and the result is cached beside the checkpoint. A 256 by 256 picture takes under four minutes on a laptop. The correctness claim is checked against diffusers and transformers, agreeing to between 6.5e-8 and 2.4e-4.

tensai audio runs Qwen2-Audio-7B-Instruct. WAV samples become Whisper's log-mel spectrogram, Whisper-large-v3's encoder turns each 40ms slice into a vector in the language model's embedding space, and those vectors stand in for the audio placeholder in the prompt while the Qwen2 language model answers. The front end matches transformers to 1e-4 and the encoder to a relative 8e-6.

Both are the same bet as the rest of the project: a single static binary that does inference without a Python environment, at the cost of reimplementing pipelines that the reference implementations already do.

The container is scratch, and it deliberately leaves out the GPU path

The Dockerfile is short and its comments explain the decisions better than the file does.

The build stage is golang:1.27 with CGO_ENABLED=0 and GOEXPERIMENT=simd, producing a stripped static binary. The runtime stage is scratch, containing the CA certificate bundle, the binary, and nothing else. Two environment variables matter: XDG_CACHE_HOME points at /cache because models download there, and TENSAI_ADDR is set to :8080 because serve binds loopback by default and a container has to answer on the bridge for a port mapping to work. There is a declared volume on /cache for exactly the model-download reason, with the documentation suggesting mounting the host cache instead to keep models across runs.

The notable omission is the wgpu tag, and the reason is in a comment: purego's dlopen layer needs a dynamic libc, and scratch has neither that nor Vulkan. The container is therefore the CPU path, and the untagged build is genuinely static.

The Makefile shows how careful the build matrix is. It defaults to the wgpu24 binding, and the comment explains the alternative: wgpu24 binds the v29 C API, which can see non-conformant Vulkan drivers, and that is what reaches a real GPU through Dozen inside WSL2 where a plain wgpu build falls back to a software rasterizer. Both bindings ship in the cross-built archives so a user whose driver disagrees with one is not left with no binary and a Go toolchain to install. Versioning runs through gobump, and cross-compilation through goxz for linux, darwin and windows on amd64 and arm64.

Editorial conclusion

tensai fits a Go developer who wants a model to train and serve in one binary with no Python in the pipeline, and who will read the docs because the API surface is broad and the build tags matter. Leave it if you need the ecosystem around a mainstream framework, since the layer list is small and the project is at version 0.0.x. Verify first which Go toolchain you have, because the SIMD path needs Go 1.26 or 1.27 on amd64 and 1.27 on arm64, and check that a plain build still works, because every other configuration is supposed to fall back to portable kernels with identical results.

Frequently asked questions

What does the tensai framework do?

It implements forward passes, backpropagation and optimization in pure Go for learning and experiments. Layers include Embedding, Dense, Conv2D, MaxPool2D, BatchNorm, LayerNorm and Dropout, with MeanSquaredError, SoftmaxCrossEntropy and BinaryCrossEntropy losses and SGD, Adam and AdamW optimizers.

How do I enable SIMD acceleration in tensai?

Build with GOEXPERIMENT=simd, which needs Go 1.26 or 1.27 on amd64, or Go 1.27 on arm64 since that is the first release whose simd/archsimd has an arm64 half. Any other build uses the portable fallbacks automatically and produces identical results.

Does tensai have any external Go dependencies?

The default build has none. The optional wgpu build tag adds exactly one, ebitengine/purego, and it is cgo-free because the wgpu-native shared library is dlopen-ed at runtime. go.mod declares go 1.22 and requires only that package.

How do I run a tensai model on a GPU?

Build with -tags wgpu, which works on Linux, macOS and Windows, and reaches anything wgpu-native supports including Vulkan, Metal and D3D12. UseAccelerator moves products above 4e8 multiply-accumulates onto the device, and tape.UseDevice keeps values, gradients and the Adam update resident there.

Can tensai run a Docker image model or understand audio?

Yes. tensai image runs Qwen-Image-2.1 end to end, and tensai audio runs Qwen2-Audio-7B-Instruct, both in pure Go. The image path quantizes weights as they load, giving 7GB at eight bits and 3.5GB at four, and a 256x256 picture takes under four minutes on a laptop.

What version of tensai should I depend on?

The releases follow the v0.0.x line, currently around v0.0.33. go.mod retracts v0.1.0 and v0.1.1, with a comment saying v0.1.0 was tagged in error before the package split, so those two versions are not usable.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mattn-tensai.svg)](https://hysenlabs.com/projects/mattn-tensai)