go-attention: three examples cut off mid-statement, printed weights that do not match their inputs, and a benchmark that argues with its own consistency claim
A full attention mechanism and transformer in pure go.
At a glance
- What is it?
- go-attention is a pure Go library with two packages, attention and transformer, and a go.mod that requires nothing at all. Read the documentation closely before trusting it: all three quick examples are cut off part way through a statement, the published example output cannot be reproduced from the query and keys printed above it, and the same README promises a single code path, assembly fallbacks, and runtime auto-tuning at once.
- Who is it for?
- The underlying library looks like a real, dependency-free implementation, and the zero-dependency claim is the one thing in this repository you can verify from a single file: go.mod carries a module line and a go directive and nothing else. What you should not do is trust the README as a specification.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 109 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
All three quick examples stop part way through a statement
The README offers three numbered examples and none of them is complete. The basic dot-product example builds a query, a key matrix, and a value matrix, calls `attention.DotProductAttention`, and then stops at `if err !=`. That is not valid Go and not a complete thought. The multi-head example configures `MultiHeadConfig`, constructs the module with `NewMultiHeadAttention`, handles the error, and then stops at `batchSiz`. The transformer example reaches `seqLen := 3` and stops at `inpu`. Three code blocks, three mid-token cuts, all of them in the places where the interesting part begins. The core type declarations above them are intact:
type Vector []float64 // Represents a 1D vector of float64 values
type Matrix []Vector // Represents a 2D matrix of float64 valuesso the API surface the examples are built on is clear even where the examples are not. `API.md` sits in the tree and is pointed at for complete documentation, which is where the missing lines have to come from.
go run api_examples.go names a file the repository root does not have
The Quick Start is two commands:
# Get the module
go get github.com/takara-ai/go-attention
# Run the examples
go run api_examples.goThe top level of the repository holds `.gitignore`, `API.md`, `LICENSE`, `README.md`, `go.mod`, and four directories: `attention/`, `cmd/`, `examples/`, and `transformer/`. There is no `api_examples.go` at the root. The examples live under `examples/`, which holds four subdirectories named `advanced`, `basic`, `data`, and `performance`, so the file the command names is somewhere below that directory or absent from the tree. The command is written as if run from the module root, and `go get` on the module's own path does not place a file in the working directory, so the second line cannot resolve against anything the first line fetched.
The printed attention weights do not follow from the printed query and keys
The example input is small enough to check by hand. The query is `[1.0, 0.0, 1.0, 0.0]`. The three keys are the same vector, the vector `[0.0, 1.0, 0.0, 1.0]`, and the all-halves vector. The dot products are therefore 2, 0, and 1. Softmax over those three scores gives about 0.665, 0.090, and 0.245. Scaling by the usual factor of one over the square root of four first, so the scores become 1, 0, and 0.5, gives about 0.506, 0.186, and 0.307. The README prints neither set:
Query: [1 0 1 0]
Attention Weights: [0.523 0.174 0.302] // Shows focus on similar patterns
Output: [2.558 3.558] // Weighted combination of valuesThe weights and the output are consistent with each other, since the output is the weighted sum of the three value rows, but neither matches the softmax of the printed inputs under either convention, and the three printed weights sum to 0.999 rather than 1. Treat the sample output as illustrative rather than reproducible.
A single code path, assembly fallbacks, and auto-tuning cannot all hold
Three claims in this README are mutually exclusive. The Performance Features section says a single optimized code path ensures repeatable results across all input sizes, and the Performance Characteristics section repeats that the implementation uses a single, highly optimized code path. The section on Go's limitations in AI says SIMD support works through assembly-optimized critical paths with automatic fallbacks. The production monitoring block then offers a switch named `EnableAutoTuning`:
// Enable performance monitoring and auto-tuning
config := attention.DefaultPerformanceConfig()
config.EnableMonitoring = true
config.EnableAutoTuning = true
attention.SetPerformanceConfig(config)
// Use memory pools to reduce allocations
v1 := attention.GetVectorFromPool(size)
defer attention.PutVectorToPool(v1)
// Get performance statistics
stats := attention.GetAllPerformanceStats()A fallback path is by definition a second code path, and a global auto-tuning switch can select among paths at runtime. Alongside the promise that the same input always produces the same output, that is three mechanisms where the README describes one. Whether the toggles are on by default is not stated anywhere in the visible documentation.
The consistency claim sits next to numbers that grow more than twentyfold
The benchmark figures are given for one machine, an Apple M1, measured with `go test -bench=. ./attention`. Small vectors in the 64 to 256 dimension range are reported at about 17 to 62 nanoseconds per dot product. Medium vectors from 512 to 1024 are reported at about 250 to 290 nanoseconds. Large vectors at 4096 and above are reported at about 970 to 1230 nanoseconds. Those bands are the opposite of the headline. A single figure for every input size would support the claim of no performance cliffs; instead the worst case in the largest band is more than seventy times the best case in the smallest. Even inside one band the spread is wide, since the small-vector range covers nearly a factor of four. A fourth bullet then generalises the single machine's numbers to any Go-compatible platform.
MultiHeadConfig names the head dimensions and TransformerConfig does not
The two configuration structs describe the same idea at different levels of explicitness. `MultiHeadConfig` asks for five fields: `NumHeads`, `DModel`, `DKey`, `DValue`, and `DropoutRate`. The comments on `DKey` and `DValue` both read as the size per head, expressed as DModel divided by NumHeads, and the example sets NumHeads to 4, DModel to 64, and both DKey and DValue to 16, which is self-consistent. `TransformerConfig` asks for four fields: `DModel`, `NumHeads`, `DHidden`, and `DropoutRate`. The per-head dimensions are gone, and a feed-forward width of 256 against an embedding size of 64 appears in its place. So the same layer can be configured two ways, one of which silently derives what the other demands you state, and no example shows a transformer config where 64 does not divide evenly by the head count.
cmd/ ships an entry point the README never tells you to install
The module declares no requirements at all. `go.mod` is two lines of content: the module path `github.com/takara-ai/go-attention` and a go directive of `1.23.4`. There is no require block, so the claim of zero external dependencies is checkable by reading a single file rather than taking it on trust. Two conventions sit oddly beside that cleanliness. The go directive pins a full patch version rather than a language version, which requires a recent toolchain to parse at all. And the tree contains a `cmd/` directory that no instruction in the README mentions: the Quick Start offers `go get` plus `go run`, and never a `go install` of any path. A framework that ships a command directory and documents only library usage leaves the binary's name and flags undefined.
A frontier research claim with no homepage, no release, and no tag
The README opens by attributing the project to the Frontier Research Team at takara.ai and calls it the first pure Go implementation of attention mechanisms and transformer layers. Neither claim can be checked from anything in the repository, and the surrounding metadata offers nothing to check them against: no homepage is recorded for the project, and the repository has no GitHub releases at all. There is also no tag, so the module you pull with `go get` has no version identity, and the only version-like number anywhere is whatever the build recorded. The last push is dated 2026-06-15. The MIT licence and a LICENSE file at the root are the one part of the packaging that is unambiguous, so a reader can tell what they are allowed to do with the code even when they cannot tell which build they have.
Editorial conclusion
The underlying library looks like a real, dependency-free implementation, and the zero-dependency claim is the one thing in this repository you can verify from a single file: go.mod carries a module line and a go directive and nothing else. What you should not do is trust the README as a specification. Fix the example arithmetic before you write a test against those numbers, read API.md instead of the truncated blocks for the actual signatures, and run the benchmark yourself on your own hardware rather than adopting the Apple M1 figures. The absence of any release or tag means there is no version to pin and no changelog to check what broke.
Frequently asked questions
What Go version and dependencies does go-attention require?
`go.mod` declares the module as `github.com/takara-ai/go-attention` with a go directive of `1.23.4` and no require block, so the module pulls in no external packages.
How do I run the examples that come with go-attention?
The Quick Start shows `go get github.com/takara-ai/go-attention` followed by `go run api_examples.go`. The repository root does not contain that file; the examples are organised under `examples/` into `advanced`, `basic`, `data`, and `performance` subdirectories.
Which attention functions does the go-attention package expose?
`attention.DotProduct` for a plain dot product, `attention.DotProductAttention` taking a query vector plus key and value matrices and returning output, weights, and an error, and `attention.NewMultiHeadAttention` built from a `MultiHeadConfig`. The transformer package adds `NewTransformerLayer` from a `TransformerConfig`.
Where is the full API reference for go-attention?
In an `API.md` file at the repository root, which the README points to for complete documentation. The three quick examples in the README itself are each cut off part way through a statement.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/takara-ai-go-attention)