# gotreesitter: a pure-Go tree-sitter runtime without CGo

> gotreesitter reimplements the tree-sitter runtime in Go so grammars become binary blobs instead of C sources. It removes the C toolchain from cross-compilation, at the cost of a young, single-maintainer codebase.

**odvcencio/gotreesitter** — Pure Go tree-sitter runtime. It cross-compiles to any GOOS/GOARCH target Go supports, including wasip1.

- Repository: https://github.com/odvcencio/gotreesitter
- Website: https://gotreesitter.m31labs.dev
- Stars: 568 · Forks: 47
- Language: Go
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/odvcencio-gotreesitter

## The CGo problem gotreesitter exists to remove

Every Go binding to tree-sitter in the ecosystem links against the C runtime. That single dependency has three consequences the README calls out directly. Cross-compilation needs a C cross-toolchain per target, so a build with GOOS=wasip1 or GOARCH=arm64 from a Linux host, or any Windows target without MSYS2/MinGW, will not link. CI images must carry gcc and the grammar's C sources, which means go install fails for downstream users who have no C compiler. And the Go race detector, coverage instrumentation, and fuzzer cannot see across the CGo boundary, so bugs in the C runtime or in FFI marshaling stay invisible to go test -race.

gotreesitter is aimed at Go teams that hit one of those three walls. If you are writing an editor, a linter, a code indexer, or a browser-side parser and you want a single static binary that builds the same way on every platform, this is the project's whole reason for existing. It is not aimed at people who need the C runtime's exact semantics, or who want to write grammars in C and link them dynamically.

## How the parser, lexer and grammar blobs fit together

The architecture is a full reimplementation, not a wrapper. The README states that the parser, lexer, query engine, incremental reparsing, arena allocator, external scanners, and tree cursor are all implemented in Go. The only input is the grammar blob.

Those blobs come from ts2go, a tool that extracts grammar tables from upstream parser.c files, compresses them into binary blobs, and deserializes them on first use. The format is the same parse-table format the C runtime uses, so the tables are not a new invention; they are a repackaging. The registry ships 206 grammars. That number is the practical boundary of the project: if your language is in the registry, you get it without a C toolchain; if it is not, you are in regeneration territory.

On top of the parser sit several layers. The query engine supports the S-expression pattern language with structural quantifiers, alternation, field constraints, negated fields, anchors, and standard predicates. A FactProgram compiles once and extracts definitions, calls, heritage edges, and imports in a single traversal, which the README positions for hot indexing paths that need common symbols but not arbitrary tags-query semantics. A taproot package is a small harness for grammargen-backed DSLs, caching generated or blob-loaded languages and returning a Walker with CST helpers. The WebAssembly targets expose parsing, queries, and highlighting, and return both UTF-8 byte offsets and JavaScript UTF-16 code-unit offsets.

## Installing gotreesitter and parsing a file

The module installs with a single go get. There is no C toolchain step and no cgo build tag to set.

```bash
go get github.com/odvcencio/gotreesitter
```

The README's quick start builds a parser from the bundled Go grammar and prints the root node. Note that the example discards the error from Parse.

```go
import (
    "fmt"

    "github.com/odvcencio/gotreesitter"
    "github.com/odvcencio/gotreesitter/grammars"
)

func main() {
    src := []byte(`package main

func main() {}
`)

    lang := grammars.GoLanguage()
    parser := gotreesitter.NewParser(lang)

    tree, _ := parser.Parse(src)
    fmt.Println(tree.RootNode())
}
```

If you are parsing files from disk rather than a string, grammars.DetectLanguage("main.go") resolves a filename to the matching LangEntry, so you do not have to map extensions to languages yourself.

The first real decision comes with error handling. The default parse methods preserve tree-sitter's partial-tree behavior: if a timeout, cancellation flag, token-source EOF, or parser safety limit stops the parse early, the returned tree records the stop reason and the error stays nil. That is the right default for an editor, where a partial tree still has value. For a batch job it is the wrong default, because you can silently index half a file. The strict variants exist for that case:

```go
parser.SetTimeoutMicros(50_000)

tree, err := parser.ParseStrict(src)
if errors.Is(err, gotreesitter.ErrParseStoppedEarly) {
    fmt.Println(tree.ParseStopReason())
}
```

Strict variants are available for full parse, incremental parse, token-source parse, factory parse, ParseWith, and ParserPool, so the choice is not limited to the simplest entry point.

## Partial trees are the default, and that is a real failure mode

The partial-tree default is the sharpest edge in the library. A parse that hits a safety limit returns a tree with a stop reason and a nil error. Code that checks only err will treat a truncated tree as a successful parse. For an editor this is correct behaviour and the README says so. For anything that writes results to a database, a cache, or a code index, it means you can persist an incomplete symbol table with no signal that anything went wrong.

The strict surface is the answer, but it is opt-in per call site. grammars.ParseFilePooledStrict returns a pooled *BoundTree only after a complete parse, and the caller must call Release on a returned tree; it returns nil and an error for parser setup failures, parse failures, and early stops. Tagger.TagStrict returns tags only after a complete parse and releases its internal tree before it returns. The strict incremental methods, HighlightIncrementalStrict and TagIncrementalStrict, behave differently again: the README says they return the partial tree together with ErrParseStoppedEarly and skip running queries after an early stop. So the failure contract is not uniform across the API, and you need to read each method's behaviour rather than assume one rule.

The second limitation is grammar coverage. 206 grammars is a lot, but it is not every grammar upstream tree-sitter has, and the README does not document what happens when you need one that is missing. The code-understanding helpers also skip unsupported languages or ambiguous shapes, which means a fact extraction that returns nothing may mean "no definitions here" or "this shape is not handled." Those two cases are indistinguishable from the return value alone.

## Alternatives: upstream tree-sitter and its CGo bindings

The obvious alternative is upstream tree-sitter itself, used from Go through a CGo binding. The difference in approach is where the parser lives. Upstream tree-sitter is a C library, and the Go binding marshals across the FFI boundary. gotreesitter is a Go program that reads the same parse tables. That single change is what buys cross-compilation to any GOOS/GOARCH target Go supports, including wasip1, and what lets go test -race and the Go fuzzer observe the parser.

The trade-off runs the other way too. Upstream tree-sitter has a C runtime that has been exercised by a much wider set of embedders, and its grammar ecosystem is the reference. With gotreesitter you are relying on ts2go to have extracted the tables correctly for your grammar, and on the Go reimplementation to match the C runtime's behaviour on ambiguous input. The repository's test tree is dominated by admission and route-equality tests with names like admission_route_equality_leaf_tiling_test.go and admission_switch_csharp_materiality_test.go, which suggests the maintainers are actively chasing grammar-specific divergences. That is a good sign for diligence and a warning about how much surface area there is to match.

If your project already builds with CGo and you have no cross-compilation requirement, the upstream route has fewer moving parts. If you are shipping a single binary to multiple platforms, or you want the race detector to see parser bugs, the calculus flips.

## Licence, maintenance and the cost of upgrading

gotreesitter is MIT licensed. The module depends on golang.org/x/sync and gopkg.in/yaml.v3, both permissive, so there is no copyleft obligation introduced by the runtime itself. The grammar blobs are a separate question: the README says ts2go extracts tables from upstream parser.c files, and those upstream grammars carry their own licences. The repository does not document a per-grammar licence inventory, so if you redistribute a built binary containing grammar blobs, that is the thing to check with your own counsel. Nothing here is legal advice.

The version history is dense and recent. v0.50.0 and v0.50.1 both landed on 2026-08-14, and v0.51.0 followed on 2026-08-16. Three releases in three days, with a minor version bump between v0.50.1 and v0.51.0, is the shape of a project that is still settling its API. The last push to the default branch was on 2026-08-16. The repository is not archived.

For an adopter, that means pinning matters more than usual. There is a CHANGELOG.md at the repository root, and the strict-versus-partial split means an upgrade can change error behaviour without changing a signature. Read the changelog between the version you pin and the version you move to, and re-run your own corpus through ParseStrict rather than trusting that a patch bump is inert.

## Conclusion

Adopt gotreesitter if you are shipping a Go binary that must cross-compile to wasip1, arm64, or Windows from a Linux host without a C cross-toolchain, and you can live with a runtime that is still moving. Do not adopt it if you need the upstream C runtime's exact behaviour on every grammar, or if you depend on a grammar that is not in the registry and cannot regenerate a blob. Before committing, verify that your target grammar parses your real corpus with ParseStrict, and check the CHANGELOG between v0.50.0 and v0.51.0 for API churn.

## FAQ

### What is the point of gotreesitter compared with the C tree-sitter runtime?

It removes the CGo dependency from Go programs that parse source code. The README states that the parser, lexer, query engine, incremental reparsing, arena allocator, external scanners, and tree cursor are all implemented in Go, so the build cross-compiles to any GOOS/GOARCH target Go supports, including wasip1, and go test -race can observe the parser.

### Is gotreesitter an alternative to nvim-treesitter?

No. gotreesitter is a Go library for embedding a parser in a Go program, not a Neovim plugin. It does expose queries and highlighting through its WebAssembly targets, but its documented use is calling NewParser and Parse from Go code.

### How do I install gotreesitter?

Fetch the module with go get github.com/odvcencio/gotreesitter and import it alongside github.com/odvcencio/gotreesitter/grammars. No C toolchain or cgo build step is involved.

### What happens if a gotreesitter parse stops early?

By default the returned tree records the stop reason and the error stays nil, which the README describes as useful for editors and diagnostics. Use ParseStrict and check for gotreesitter.ErrParseStoppedEarly when partial output should fail the request.

### How many grammars does gotreesitter ship?

The README states that 206 grammars ship in the registry, extracted from upstream parser.c files by ts2go and compressed into binary blobs that are deserialized on first use.

## Sources

- [Official documentation](https://gotreesitter.m31labs.dev)
- [Official README](https://github.com/odvcencio/gotreesitter#readme)
- [Project repository](https://github.com/odvcencio/gotreesitter)
- [Release notes](https://github.com/odvcencio/gotreesitter/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/odvcencio-gotreesitter
