Open-source project
kythe/kythe avatar
kythe/kythe

Kythe: a schema-first attempt to make code knowledge portable across languages

Kythe is a pluggable, (mostly) language-agnostic ecosystem for building tools that work with code.

2,161 stars275 forksGoApache-2.0

At a glance

What is it?
Google's open source system for representing facts about code as graph nodes, with indexers for C++, Go and Java and a verification layer to check them.
Who is it for?
Kythe is a schema and a set of indexers rather than a product you switch to, which is the right frame for judging it. What the repository ships is well specified: a protobuf schema, extractors for the common build systems, a verifier, and a sample cross-reference service that demonstrates what the data is good for.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A graph schema, not a search engine

The project's one-line description calls Kythe a pluggable, mostly language-agnostic ecosystem for building tools that work with code. That phrasing matters. Kythe is not a search product, and it is not an IDE. It is a way of writing down facts about code in a form that any tool can read.

Those facts are graph nodes and edges. A node can be a variable, a function, a file or a reference. An edge can say that a variable is declared here, that this call resolves to that function, or that this definition comes from that file. Because the representation is a schema rather than an index, two unrelated tools can produce it and a third can consume it without either knowing about the other.

The README's feature list is short enough to quote in full, and each item maps to a layer of that idea. Extensive documentation of the Kythe schema comes first, because the schema is the contract. Indexer implementations for C++, Go and Java follow. Compilation extractors cover the build systems that front those compilers: javac, Maven, cmake, Go and Bazel. Then there is a generic verifier for indexers, a sample cross-reference service, and a set of utility commands for working with the resulting artifacts.

That ordering tells you how the project thinks about itself. The schema is the durable part. Indexers and extractors are replaceable, and the verifier exists precisely because replaceable pieces drift.

The repository is Apache 2.0 licensed, written mostly in Go, and hosted at kythe.io. It has 2,159 stars, 276 forks and 273 open issues, with the last push on 2026-09-18.

Getting the snapshot, and what is actually in it

The README's getting started section is three shell commands. You download a release tarball, unpack it, and replace any previous installation at `/opt/kythe`:

bash
tar xzf kythe-v*.tar.gz
rm -rf /opt/kythe
mv kythe-v*/ /opt/kythe

Then the packaged README in that directory describes the tools and their usages, which is where the real reference lives rather than in the GitHub README. That split is typical of a project whose audience is engineers who already know what they want to index.

The repository tree is consistent with a large Bazel monorepo that has absorbed several languages. At the top level there is `WORKSPACE`, `BUILD`, `external.bzl`, `setup.bzl` and `visibility.bzl`, which is Bazel's older workspace style rather than the newer bzlmod layout. There is `go.mod` and `go.sum` for the Go components, `package.json` and `pnpm-lock.yaml` for the TypeScript indexer, and `maven_install.json` for Java dependencies. Alongside them sit `third_party/`, `buildenv/`, `tools/` and `kythe/` itself, plus `.bazelci/` and a `.pre-commit-config.yaml`.

Two things in that tree are easy to overlook. `README.langserver.md` sits next to the main README, which tells you language-server integration is treated as a first-class consumer. And `RELEASES.md` sits in the root, separate from the GitHub releases, which suggests version history is curated in the repository.

One dating detail is worth noting. The README is AsciiDoc and carries a date line of 19-May-2015 under the title. The schema and the basic architecture date from that era. What moves now is the extractor maintenance.

The TypeScript indexer as the smallest complete example

The `package.json` at the repository root is small and shows how an extractor is wired together. It is named `kythe-typescript-indexer`, its binary is `./indexer.js`, and its scripts map directly onto the pieces of the pipeline:

json
"scripts": {
  "build": "tsc",
  "watch": "tsc -w",
  "browse": "node indexer.js tsconfig.json | ./browse",
  "test": "bazel test //kythe/typescript:indexer_test",
  "unit_test": "bazel test //kythe/typescript:utf8_test",
  "fmt": "clang-format -i *.ts"
}

Read those scripts and the structure of the project becomes concrete. `indexer.js` takes a `tsconfig.json`, which means the TypeScript indexer is driven by a real compiler configuration rather than by a heuristic scan. Piping its output into a `browse` binary gives you the sample cross-reference service the README mentions, which is how you look at what an indexer produced. The tests run through Bazel even though the indexer is TypeScript, so the test story is shared with the rest of the repository rather than bolted on.

The dependencies are also revealing. `typescript` is pinned to exactly `5.3.2` rather than a caret range, which is what you would expect when the extractor's output depends on compiler internals. Alongside it sit `source-map` and `source-map-support`, needed to map compiled positions back to original source, plus `balanced-match`, `brace-expansion` and `minimatch` for source layout analysis. That is a recognisable profile for a tool that has to reason about where a token came from.

Note what is absent: no LangChain, no LLM dependency, no model weights. Kythe is deterministic program analysis. If you are evaluating it, that is a feature rather than a limitation, though it does mean it will not answer semantic questions about your code that require a model.

Why the Go dependency list explains the scope

The root `go.mod` declares its module as `kythe.io`, and the dependency list reads like an inventory of what the system does. Language serving and language-server support are there with `github.com/sourcegraph/go-langserver` and `github.com/sourcegraph/jsonrpc2`. Storage and compression show up as `github.com/syndtr/goleveldb`, `github.com/golang/snappy` and `github.com/DataDog/zstd`. Tree-structured data is handled by `github.com/beevik/etree`, which is there for the XML that comes out of javac and cmake extractors.

Two entries stand out because they explain features the README does not mention. `github.com/apache/beam` at a version marked incompatible is the Beam SDK, which is how the project handles indexing at a scale larger than one machine. And `github.com/hanwen/go-fuse` is a FUSE implementation, which is how the sample cross-reference service can present an index as a browsable filesystem rather than only over HTTP.

`github.com/google/codesearch` is another one. It comes from Google Code Search, and its presence suggests Kythe carries over some indexing machinery from that lineage. `github.com/minio/highwayhash` gives stable hashing, which you need if you want to identify equivalent nodes across independent index runs.

The `go.mod` also pins the language-side tooling: `golang.org/x/tools` at v0.48.0 for the Go extractor, `google.golang.org/grpc` at v1.83.2 for the service boundary, and `google.golang.org/protobuf` at v1.36.11, which is the serialization format the whole schema is built on. If you want to know whether a given release of Kythe handles a language well, the extractor dependency versions are more telling than the README.

The verifier, and why indexers need one

The generic verifier for indexers is listed in the README features and is the least glamorous item there, which is a shame. It exists because the failure mode of this kind of system is silent.

If an indexer emits facts about code and one of those facts is wrong, nothing crashes. A tool consuming the graph gets a confidently wrong answer, such as a rename that misses a caller because the indexer failed to resolve a reference. There is no compiler error to notice. The verifier is how you check that an indexer's output obeys the schema's invariants before you trust it, and having it be generic means you can point it at any indexer, including one you wrote.

This also explains why the schema documentation is the first item in the feature list. The schema is the contract between producers and consumers, so its edge cases and required properties are the thing everything else depends on. The docs live at kythe.io/docs, and the schema documentation is extensive enough that the project lists it as a feature rather than assuming readers will infer it from code.

The utility commands round this out. Working with graph artifacts is not something you want to do with ad hoc scripts, and commands for inspecting and transforming them are part of the deal.

If your interest is cross-language navigation rather than extraction, the sample cross-reference service is the piece to look at first. It is the reference answer to the question of what Kythe data can power once you have it, and the TypeScript indexer's `browse` script wires an indexer straight into it.

Release history as a picture of where the work goes

The versions are numbered 0.0.x, which is a deliberate signal that the snapshot format and tooling may still move. The three most recent releases are a good guide to what actually consumes maintenance time.

v0.0.76, on 2026-07-16, touched only the Rust extractor: placing the root module at the start of `source_file`, and setting an `is_workspace_member` flag to indicate which crate is being indexed. Both are small correctness and layout fixes, the kind that only surface once someone indexes a real multi-crate Rust project.

v0.0.75, on 2026-03-12, is the opposite: a Go extractor Docker image updated to Go 1.25, an upgrade to the Bazel extractor, and BUILD and bazelrc changes for Bazel 9 based clients. Two contributors made their first contributions here, which is a reasonable signal that outside contributions do land.

v0.0.74, from 2025-11-07, is the messier one. It fixed JDK installation notes for Apple Silicon, bumped libffi and zlib to fix macOS builds, and added missing `@CanIgnoreReturnValue` annotations. It also bumped `rules_go` to 0.43.0 and then reverted that bump, both in the same release.

Reading those three together gives a clear picture. Upstream toolchain churn is a recurring tax, language extractors get targeted fixes when someone hits a real case, and the version number has stayed below 0.1 for a decade. If you plan to depend on the snapshot format, that is the thing to plan around.

Editorial conclusion

Kythe is a schema and a set of indexers rather than a product you switch to, which is the right frame for judging it. What the repository ships is well specified: a protobuf schema, extractors for the common build systems, a verifier, and a sample cross-reference service that demonstrates what the data is good for. What it does not ship is an answer to the question every consumer actually has, which is how to keep extractors working as compilers and build tools move underneath them. Release v0.0.76 changed only the Rust extractor, and v0.0.75 was mostly dependency and Bazel version churn, which tells you where the effort goes. If you are building code search, a code knowledge graph, or a cross-language refactoring tool, read the schema documentation at kythe.io first, then look at the TypeScript indexer's package.json as the smallest complete example of how an extractor is wired together.

Frequently asked questions

What does LALR stand for in the context of parser generators?

That question is about parser theory rather than this repository, so it has no Kythe-specific answer. Kythe does not generate parsers. It consumes compilation data from javac, Maven, cmake, Go and Bazel, and the parsing work happens inside those tools and inside Google's internal toolchain rather than in this project.

What is Kythe actually used for?

It is an interchange format for code facts. An indexer emits nodes and edges describing declarations, references and definitions, and another tool consumes them. The project ships a sample cross-reference service that demonstrates one use, and a language-server integration path that demonstrates another. The schema is the part other projects depend on.

Which languages does Kythe index out of the box?

The README lists indexer implementations for C++, Go and Java, and the root package.json is for a TypeScript indexer, so the TypeScript path is in the repository too. Releases through v0.0.76 have been adding Rust extractor fixes, which indicates Rust is being brought up to a comparable level.

Do I need to write my own extractor to use Kythe?

Not to read data. If you want to index a language the project does not cover, you would write an extractor that emits the schema, and you would run the project's generic verifier against its output before trusting it. The schema documentation at kythe.io is the contract to code against, and it is listed as the project's first feature.

Why is the version number still 0.0.x?

Because the snapshot format and the surrounding tooling have not been declared stable. The releases show the pattern: v0.0.74 spent most of its effort on macOS build fixes and a Bazel dependency bump that was later reverted, while v0.0.75 and v0.0.76 were mostly toolchain and extractor updates. Anyone depending on the format long term should expect it to move.

Official sources

  1. kythe/kythe on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kythe-kythe.svg)](https://hysenlabs.com/projects/kythe-kythe)