Tabby: a pinned Tantivy rev, an OpenAPI doc scrubbed by jq, and a UI copied into the binary
Self-hosted AI coding assistant
At a glance
- What is it?
- Tabby is a self-hosted AI coding assistant written in Rust, positioned as an open-source and on-premises alternative to GitHub Copilot, with inference through llama.cpp or Ollama and support for consumer-grade GPUs. The workspace pins its search engine to a single commit of a Tantivy fork, the release version is bumped from main with one cargo command, and the web interface is a pnpm and Turborepo build that gets copied into the Rust webserver crate.
- Who is it for?
- Use Tabby when your code cannot leave the network and a hosted assistant is not an option, and when you can accept owning a SQLite file, an embedded model and an index that the project updates by hand. Do not use it as a drop-in for GitHub Copilot without reading the roadmap and the version you are actually installing, because the newest tagged release is v0.32.0 from 2026-01-25 while the workspace version in Cargo.toml reads 0.33.0-dev.0.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 93 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Self-contained means an embedded SQLite file, not the absence of a database
The feature list leads with three claims: self-contained with no need for a DBMS or cloud service, an OpenAPI interface for integrating with existing infrastructure such as a cloud IDE, and support for consumer-grade GPUs.
The first claim is true in the sense that matters operationally, and the workspace shows how. There is an ee/tabby-db crate, an ee/tabby-db-macros crate, an ee/tabby-schema crate, and a dedicated crate called sqlx-migrate-validate whose whole job is validating migrations. The Makefile's schema target dumps from a file called schema.sqlite using the sqlite3 command line, then visualises it into an SVG.
So there is a relational database, it is SQLite, and it is embedded in the process rather than a server you install and secure. That is a better trade than the claim sounds like: nothing to connect to, nothing listening, and the whole state of Tabby is a file you can copy to back it up. It also means the migration story is yours to own, and a schema change that a server database would have handled with a connection string is a file you have to move.
Two inference backends in the workspace, and they are not interchangeable
The crate list is where the architecture is legible. There is crates/tabby-inference, crates/tabby-download, crates/tabby-index, crates/tabby-crawler, crates/tabby-git, and then two backend crates with very different names: crates/llama-cpp-server and crates/ollama-api-bindings.
One is a llama.cpp server compiled in, the other is a set of bindings that speak to an Ollama instance. Those are different operational models. The first means the model runs inside Tabby's process and you ship the weights; the second means something else is already serving the model and Tabby is a client of it. A model that one can load and the other cannot is a normal outcome, so which backend you run decides which models you can use.
The history supports the range. Metal inference on Apple's M1 and M2 landed in v0.1.1, Llamafile deployment integration arrived in v0.21.0, and the model directory has grown its own registry with CodeGemma, CodeQwen and a Codestral integration announced in July 2024. The consumer GPU claim is the one that makes this deployable on a single workstation rather than a datacentre, and the support for it is in the two backend crates.
The index is one commit of a Tantivy fork, not the published crate
One line in the workspace dependencies is worth stopping on: tantivy, declared as a git dependency on quickwit-oss/tantivy at revision 4143d31. Tantivy is the search library, and it is what a code assistant queries when it looks for relevant context across a repository, so this is load bearing rather than incidental.
The dependency is pinned to a commit hash rather than a version. That means the build fetches that exact tree, that upgrading it is a deliberate edit to the manifest, and that the version of the search engine in your Tabby is not something a routine cargo update will move you off. It is also the kind of choice that needs a reason, and the manifest does not give one, so if you are evaluating the project this is a question to ask rather than a detail to accept.
The rest of the dependency list is more conventional and tells you the shape of the service. Axum and hyper for HTTP, utoipa for the OpenAPI surface that matches the README's integration claim, juniper alongside it for GraphQL, tokio and tokio-cron-scheduler for the async runtime and scheduled jobs, tracing for logs, git2 for repository access, and reqwest with reqwest-eventsource for outbound calls.
The OpenAPI document is generated from a running server and then scrubbed
The Makefile contains one target that tells you more about the API surface than any documentation would. update-openapi-doc curls http://localhost:8080/api-docs/openapi.json, pipes it through jq, and deletes a specific list of paths and schema properties before the result is committed. The deletions include the /v1beta/chat/completions, /v1beta/search and /v1beta/server_setting paths, and inside the schemas it strips the prompt and debug_options properties of CompletionRequest, the debug_data property of CompletionResponse, and the DebugData schema itself.
Three things follow. The document is generated, not hand written, so it cannot drift from the implementation without someone regenerating it. It is scrubbed, so what you read in the repository is the public contract with the debug surface removed, and the debugging endpoints still exist in the running server. And the version prefix is visible, /v1beta, which is how the project marks an API that may still move.
If you are integrating against Tabby, that target is the procedure: run a server on port 8080, run the target, read the result. Hand-editing the committed document is how the two fall apart.
The web interface is a pnpm and Turborepo build that gets copied into the Rust crate
The JavaScript side of the repository is a monorepo, and the root package.json makes the shape explicit. It is private, it declares turbo as its only dev dependency, and every script delegates: build runs turbo build, lint runs turbo lint, test runs turbo test. The engines are Node 18 or newer and pnpm 9 or newer. Alongside it sit turbo.json, pnpm-workspace.yaml and a pnpm lockfile.
Then the Makefile shows how the two halves meet. update-ui runs the pnpm build and then copies the output of ee/tabby-ui into ee/tabby-webserver/ui, replacing the directory first, and copies the output of ee/tabby-email into ee/tabby-webserver/email_templates. The fix-ui target is just pnpm lint:fix.
So the server binary embeds a built web interface and a set of email templates, and getting from a change in the frontend to a running server is a build plus a copy, not a package link. That has an obvious upside, one artefact to deploy, and an obvious cost: building Tabby properly means running both toolchains, and the Rust side alone will not give you the current interface.
Versions are bumped from main in one command, and the manifest is ahead of the tags
Releases are cut mechanically. The Makefile has a bump-version target that runs cargo ws version with --force on every member, --no-individual-tags, and --allow-branch main, and a separate bump-release-version target for the release branch pattern. The workspace version in Cargo.toml currently reads 0.33.0-dev.0.
Compare that with the published tags: v0.32.0 from 2026-01-25, a next-alpha tag from 2026-02-09, and a nightly tag last touched on 2023-09-08. So the tree is a development version ahead of the newest release, the alpha channel moved in February 2026, and the channel literally named nightly has not been updated since 2023. The last push to the repository was 2026-06-30.
None of that is a problem in itself, but it means channel names are not a reliable guide to what is current. If you are pinning, pin an explicit version and read CHANGELOG.md, which is fed by the .changes/ directory and the .changie.yaml configuration at the root, rather than trusting that a tag called nightly means last night's build.
Development needs tmuxinator, and the completion behaviour is a rule catalogue
The dev target is one line: tmuxinator start with the project file at .tmuxinator/tabby.yml. The repository ships that directory, so a contributor is expected to have tmuxinator and a terminal multiplexer available, and the server plus its workers are expected to run in split panes rather than one process in the foreground. That is a small thing that tells you the intended development shape, and it is a real barrier for someone who has not used tmux.
The more interesting part for an adopter is that the completion behaviour is inspectable. At the root there is MODEL_SPEC.md and a rules/ directory, and the change history shows this being treated as a first-class concern: locally relevant snippets pulled from local LSP declarations and recently modified code were added for code completion in April 2024, a blog post explains the rank fusion behind enhanced code context understanding, and repository level context arrived as RAG-based code completion in v0.3.0 in October 2023.
That is worth more than a model card. A completion assistant is mostly a policy about what to send to the model, and here that policy is a versioned set of rules in the repository rather than something you have to infer from output. Read the rules directory before you conclude the completions are wrong, because you may simply be looking at a rule you can change.
Editorial conclusion
Use Tabby when your code cannot leave the network and a hosted assistant is not an option, and when you can accept owning a SQLite file, an embedded model and an index that the project updates by hand. Do not use it as a drop-in for GitHub Copilot without reading the roadmap and the version you are actually installing, because the newest tagged release is v0.32.0 from 2026-01-25 while the workspace version in Cargo.toml reads 0.33.0-dev.0. Before you deploy: pick an inference backend deliberately, llama.cpp or Ollama, since they support different models; read MODEL_SPEC.md and the rules directory, because the completion behaviour is a catalogue rather than a black box; regenerate the OpenAPI document from a running server on port 8080 with the Makefile target rather than editing it; and remember that development needs tmuxinator and the release process needs main.
Frequently asked questions
What is Tabby?
Tabby is a self-hosted AI coding assistant, presented as an open-source and on-premises alternative to GitHub Copilot. Its three stated features are being self-contained with no need for a DBMS or cloud service, offering an OpenAPI interface for integrating with existing infrastructure such as a cloud IDE, and supporting consumer-grade GPUs.
What hardware does Tabby need to run?
The README states that Tabby supports consumer-grade GPUs, and the workspace shows inference through either a bundled llama.cpp server or Ollama bindings. Metal inference on Apple's M1 and M2 arrived in v0.1.1 and Llamafile deployment integration in v0.21.0, and the project maintains a model directory listing what it supports.
How do I install Tabby?
There is an official image on Docker Hub under tabbyml/tabby, and the documentation site is tabby.tabbyml.com, which also carries the installation pages, the model directory and a deployment guide for running Tabby on any cloud with SkyServe from SkyPilot. The README itself contains no command to copy, so the documentation is the place to start.
What is the Tabby Answer Engine?
It was introduced in v0.13.0 as a central knowledge engine for internal engineering teams, described as integrating the team's internal data to deliver reliable answers to developers. Later releases added switching between backend chat models in v0.20.0, turning Answer Engine messages into persistent shareable Pages in v0.28, and indexing your own documentation through REST APIs in v0.29.
How current is Tabby, and which version should I install?
The last push to the repository was 2026-06-30, the newest tagged release is v0.32.0 from 2026-01-25, and the workspace version in Cargo.toml reads 0.33.0-dev.0. There is also a next-alpha tag from 2026-02-09 and a tag called nightly that was last published on 2023-09-08, so channel names are not a guide to freshness. CHANGELOG.md is fed from the .changes/ directory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tabbyml-tabby)