# SeaGOAT: local semantic code search that runs its own server

> SeaGOAT is a Python code search engine that combines ChromaDB vector embeddings with ripgrep, keeps everything on your machine, and answers queries like "Where are the numbers rounded". This review covers the install path, the server architecture, and where it falls short.

**kantord/SeaGOAT** — local-first semantic code search engine

- Repository: https://github.com/kantord/SeaGOAT
- Website: https://kantord.github.io/SeaGOAT/
- Stars: 1,310 · Forks: 93
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kantord-seagoat

## What SeaGOAT solves, and who it is for

Grep answers the question you already know how to phrase. It fails when you know what the code does but not what it is called. SeaGOAT targets that gap: the README describes it as "a local search tool that leverages vector embeddings to enable you to search your codebase semantically", and the example query it gives is gt "Where are the numbers rounded". That is a natural-language description of intent, not a symbol name.

The intended user is a developer working on a mid-sized repository in one of the supported languages who wants semantic search without a remote service. The README's FAQ states plainly that SeaGOAT "does not rely on 3rd party APIs or any remote APIs and executes all functionality locally". For teams under code-confidentiality constraints, that is the whole pitch. It is also not a code generator: the FAQ notes that SeaGOAT is "not a code generator but a code search engine", so it will not write the rounding fix for you, only point at where rounding happens.

## How the server, ChromaDB and ripgrep fit together

SeaGOAT is not a single binary that scans on demand. It splits into a long-running server and a thin client. seagoat-server start /path/to/your/repo brings up an indexing process for one repository; gt and seagoat are the client commands that query it. Because the server holds state, the README's FAQ explains the reason directly: embeddings and vector databases "at the moment cannot be replace with an architecture that processes files on the fly".

Inside the server, two retrieval paths run side by side. ChromaDB stores vector embeddings produced by a local embedding engine, with telemetry disabled by default, and ripgrep provides "regular expression/keyword based matches in addition to the AI-based matches". Queries can mix both: the README shows gt "function calc_.* that deals with taxes", a regex embedded in an otherwise natural-language string.

Indexing is deliberately slow and non-blocking. The FAQ says SeaGOAT is "designed to allow you to use your computer while processing files" and calls this "an intentional design choice to avoid blocking/slowing down your computer". The practical consequence is that you can query during the first indexing pass. You get a warning with an accuracy estimate, and ripgrep-backed results appear immediately while the embedding side catches up. The client also has a display layer: bat renders results when colour is on, pygments is the fallback when bat is absent, and a grep-line format takes over inside a pipeline.

## Installing SeaGOAT and running a first query

SeaGOAT needs Python 3.11 or newer and ripgrep already on the machine. bat is optional but the README calls it "highly recommended" because it is used to render results whenever colour is enabled. The documented install path is pipx, which keeps the CLI isolated from your project environments.

```bash
pipx install seagoat
```

After that, point the server at a repository. The README uses this exact form, and the path is the only thing you change:

```bash
seagoat-server start /path/to/your/repo
```

With the server up, query from another shell. Expect a short wait on the first query of a large repository, since indexing continues in the background, and expect a warning telling you how accurate the current results are:

```bash
gt "Where are the numbers rounded"
```

The gt and seagoat commands are interchangeable; pyproject.toml maps both entry points to the same CLI function. When you are done, shut the server down with the mirrored command:

```bash
seagoat-server stop /path/to/your/repo
```

Project-specific settings live in a .seagoat.yml at the repository root, merged over a global config. The README's example changes the server port, which matters if 31134 is already taken:

```yaml
# .seagoat.yml

server:
  port: 31134  # Specify server port
```

For contributors, the development path is Poetry rather than pipx. The README lists Poetry, Python 3.11 or newer and ripgrep as requirements, then poetry install, with poetry run ptw as the recommended watch-mode test loop and poetry run seagoat-server start ~/path/an/example/repository for manual testing against a development build.

## The hard coded language list is the real ceiling

The most consequential limitation is not performance, it is scope. The FAQ states that SeaGOAT is "hard coded to only process files in the following formats": .txt, .md, .py, .c and .h, .cpp/.cc/.cxx/.hpp, .ts/.tsx, .js/.jsx, .html, .go, .java, .php and .rb. If your repository is mostly Rust, Kotlin, Swift, C#, Scala or shell, the semantic half of the tool has almost nothing to index. You would be left with the ripgrep path, which is a worse grep than grep.

Binary files are ignored outright, and the FAQ says the preferred encoding is UTF-8 while noting that "most other character encodings should also work". That is a soft guarantee, not a tested one.

Platform support is uneven and the project says so. The README marks Linux as tested, macOS as "partly tested" with a link to issue 178, and Windows as "help needed" with a link to issue 179. On macOS and Windows, treat SeaGOAT as something you evaluate rather than something you depend on.

There is a second, quieter cost. Because the server maintains a vector index per repository, every repository you want to search semantically needs its own server lifecycle and its own indexing pass. SeaGOAT is a poor fit for one-off searches across many small checkouts; plain ripgrep wins there on setup time alone.

## How SeaGOAT differs from plain ripgrep and from hosted code search

The closest comparison is ripgrep itself, which SeaGOAT already depends on. ripgrep is stateless and instant: no server, no index, no background CPU, and it works on any file type because it does not care what the file is. Its queries are patterns, so you must already know the token you are looking for. SeaGOAT inverts that trade: a server and an indexing pass buy you intent-based queries, and the README's own example, asking where numbers are rounded, is precisely the query ripgrep cannot express.

The other comparison is hosted semantic code search, which typically indexes your repository on someone else's infrastructure. SeaGOAT's difference is architectural rather than qualitative: the README states the server runs entirely locally and "works even if you don't have an internet connection", and the FAQ says there are no third party or remote APIs in the current version. It does leave a door open, noting that future optional features might send data remotely "if any further improvement can be gained from that", so the local-only guarantee is a statement about the current version, not a permanent contract. The FAQ frames its own answers the same way, calling them "indications of how SeaGOAT works, but are not a legal contract".

A smaller but telling detail sits in pyproject.toml: the dependency list includes ollama and mcp, and the entry points define a seagoat-mcp command. The README does not document what either does. If you are evaluating SeaGOAT specifically as an MCP tool for an editor or agent, the repository layout suggests the surface exists, but the README gives you nothing to install against.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-09, which is recent. That said, the most recent release listed is v0.54.17 from 2025-05-14, so the release cadence and the commit cadence are not the same thing, and you should read the CHANGELOG.md in the repository rather than assume a tagged release matches main.

Upgrade cost is dominated by reindexing rather than by dependency churn. ChromaDB is pinned as ^1.0.0 in pyproject.toml, and the embedding model comes from ChromaDB's default, so a ChromaDB major bump is the event most likely to invalidate an existing index or change how results rank. The FAQ does not describe any migration or index-versioning procedure, and the README does not document rollback, so plan on deleting and rebuilding the index after such an upgrade.

On licensing: pyproject.toml declares license = "MIT" and the repository carries a LICENSE file, so the code itself is permissively licensed. That says nothing about the models or data your dependencies pull in. If the embedding model's provenance matters to your organisation, the README points you at the source code and at the issue tracker rather than answering the question, and the FAQ explicitly invites you to examine the source or raise a concern. This is not legal advice; check your own dependency chain.

## Conclusion

Adopt SeaGOAT if you work on a repository written in one of the hard coded languages, want semantic queries without sending code to a remote API, and can accept a background server plus a one-time indexing pass. Skip it if your codebase is mostly Rust, Kotlin, Swift, C# or shell, or if you need Windows support today. Before committing, verify three things on your own machine: that seagoat-server start indexes a representative directory without errors, that a query such as gt "Where are the numbers rounded" returns useful hits while indexing is still running, and that the port in .seagoat.yml does not collide with anything else you run locally.

## FAQ

### What is SeaGOAT?

SeaGOAT is a local code search engine that uses vector embeddings to let you search a codebase semantically, and it also provides regular expression and keyword matches through ripgrep. It is distributed as the seagoat Python package and installs with pipx install seagoat.

### Does SeaGOAT send my code to a remote service?

No. The README states that SeaGOAT does not rely on third party or remote APIs and executes all functionality locally using a server you run on your own machine, with telemetry disabled by default. The FAQ adds that the current version does not send data to remote servers, while noting that optional remote features could be added later if they proved useful.

### Why does SeaGOAT need a server running?

The FAQ says the server exists for response speed, because SeaGOAT relies heavily on vector embeddings and vector databases that cannot currently be replaced by an architecture processing files on the fly. You start it with seagoat-server start /path/to/your/repo and stop it with the matching stop command.

### Which programming languages does SeaGOAT support?

The FAQ says SeaGOAT is hard coded to process only .txt, .md, .py, .c, .h, .cpp, .cc, .cxx, .hpp, .ts, .tsx, .js, .jsx, .html, .go, .java, .php and .rb files. Binary files are ignored.

### Why is SeaGOAT indexing so slowly while barely using my CPU?

The FAQ describes this as an intentional design choice so that processing large repositories does not block or slow down your computer. It does not affect query performance, and you can query the repository while indexing is still in progress.

### Can I search my repository while SeaGOAT is still processing files?

Yes. The FAQ says regular expression and full text results are displayed from the very beginning, and when files are not processed yet you receive a warning with an estimation of how accurate the semantic results are.

## Sources

- [kantord/SeaGOAT on GitHub](https://github.com/kantord/SeaGOAT)
- [License: MIT](https://github.com/kantord/SeaGOAT/blob/main/LICENSE)
- [Project website](https://kantord.github.io/SeaGOAT/)
- [README](https://github.com/kantord/SeaGOAT/blob/main/README.md)
- [Releases](https://github.com/kantord/SeaGOAT/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kantord-seagoat
