# Marmot catalogs data assets and hands them to agents over MCP

> Marmot is an MIT-licensed Go data catalog that ships as a single binary with its own interface, stores rich metadata for tables, topics, queues, APIs and dashboards through a plugin system, and exports certified context to agents over the Model Context Protocol, an API and the UI, with SDKs for Go, Python and TypeScript generated from one OpenAPI document.

**marmotdata/marmot** — The open-source context layer for your AI. Catalog your tables, topics, queues and APIs then expose real metadata to your AI agents.

- Repository: https://github.com/marmotdata/marmot
- Website: https://marmotdata.io
- Stars: 618 · Forks: 32
- Language: Go
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/marmotdata-marmot

## One binary against the enterprise catalog

The pitch is deliberately against a category of software rather than for a feature. Marmot is described as an open-source data catalog for teams that want data discovery without enterprise complexity, and the argument is about what you have to run: traditional catalogs require extensive infrastructure and configuration, while Marmot ships as a single binary with an interface and can be deployed and used for cataloging within minutes.

Two deployment routes are documented. One is the deployment guide, and the other is downloading the single binary for your platform and running it. Both are first-class, which says something about the audience: this is aimed at teams that do not want to stand up a stack to find out what tables they have.

The licence supports that audience too. It is MIT, so it can be used, modified and distributed freely, and self-hosting carries no licensing cost.

The framing of the whole project is that an agent should be able to do what a person does in the catalog. The tagline makes the ambition explicit: discover any data asset in seconds, then let your AI do the same.

## Full-text search plus structured filters

The first of the five headline features is search, and its composition matters more than its existence. Marmot offers full-text search together with structured queries, boolean logic and metadata filters.

Those four things solve different problems. Full text finds the asset whose name you half remember. Structured queries find assets by property. Boolean logic lets you combine or exclude terms rather than accepting a fuzzy match. Metadata filters narrow by the fields you have actually filled in, which is the part that makes a catalog usable once it holds more than a few hundred tables.

The second feature is interactive lineage: tracing data flows from source to destination, with the stated purpose of analysing impact before making a change. In a catalog that also holds topics, queues and APIs, that is the feature that turns a list of names into a map.

The third is metadata-first storage for any asset type, from tables and topics to APIs and dashboards, which is what makes the asset list below possible.

## Asset types arrive through a plugin system

The catalog is not limited to tables in a warehouse, and the plugin system is how the other types get in.

Tables cover databases and data warehouses. Topics cover message queues and event streams. Queues cover job queues and messaging systems. APIs cover REST endpoints, GraphQL endpoints and internal services. Dashboards cover visualization tools and BI platforms.

That list is the practical shape of a data landscape, and it also explains where the project's extension story sits: the contributing guide explicitly asks for new plugins for data sources, alongside documentation improvements and bug reports. If your asset type is missing, writing a plugin is the sanctioned route rather than waiting for a release.

The fourth feature, team collaboration, sits alongside it. Ownership is assigned rather than implied, business context is documented, and shared glossaries are maintained, which are the three things that stop a catalog from becoming a stale list of table names.

## Certified context leaves through MCP

The fifth feature is the one that distinguishes this catalog from a wiki, and the question the project answers about it is worth reading in full.

Marmot exposes certified context through MCP, the Model Context Protocol, the API and the interface. That lets an AI agent query real metadata about your data assets, so the tool understands the data landscape and can trace lineage, and so it makes decisions based on actual metadata rather than guessing.

The word certified is doing work in that sentence. It implies a distinction between context a human has reviewed and context that has not, which is the difference between an agent you can let near production and one you cannot. The documentation does not define the certification mechanism, so treat the word as a stated intention rather than a described workflow.

The container image agrees with the framing. The Dockerfile sets an MCP server name label pointing at the project, and the Go module depends on the Model Context Protocol Go SDK, so the server identity travels with the image rather than being configured per deployment.

## Three SDKs from one generated document

The API surface is not hand-maintained in three languages, and the build shows how it is generated.

A swagger target in the Makefile runs the swag generator against the internal API directory, pointing it at the general info file for the v1 server, parsing dependencies, and writing the result into docs. The SDK targets in the same Makefile then cover Go, Python and TypeScript, and each language gets its own generate, lint, test and build steps, plus dependency and clean steps. There is a shared sdk target above them.

That shape means a new endpoint flows through as one change: the Go handler, the generated document, and three client libraries. It also means the generated clients are the only clients, so anyone integrating against the catalog is using the same versioned surface rather than a hand-written approximation of it.

The rest of the build is conventional Go with a version string injected at link time from the git description, and the release target runs the clean, swagger and frontend steps before building with a production build tag.

## Development runs unencrypted with telemetry off

The development recipe is written into the Makefile rather than left to a guide, and it sets three environment variables before starting the server.

Logging is set to debug. Transport is allowed to be unencrypted, through a server flag that exists so a local run does not need certificates. And telemetry is switched off with an explicit enabled flag set to false, which is a choice worth noticing: the ability to turn it off is a default in the development path rather than something you have to discover.

The rest of the repository supports that workflow. A Tiltfile sits at the top level for local Kubernetes development, a charts directory holds Helm charts for deployment, a goreleaser configuration and two release Dockerfiles, including a distroless variant, cover publishing, and an install script covers installation.

The release line is currently 0.11.0, published on September 23, 2026 after two preview builds earlier that month, and the last push to main is dated September 30, 2026. The container image is Alpine-based, runs as a non-root user, exposes port 8080, and starts with the run command.

## Conclusion

Marmot fits a team that has data spread across warehouses, queues, APIs and dashboards and needs one place to answer what exists, who owns it and what a change would affect, including from an agent rather than only from a browser. It does not fit an organisation that already runs a full enterprise catalog and needs its governance features. Before you deploy, decide how the catalog will be populated, since the asset types arrive through plugins rather than being auto-discovered, and note that the development recipe starts the server with unencrypted transport allowed and telemetry disabled, which is a development setting rather than a production one.

## FAQ

### What is Marmot?

Marmot is an open-source data catalog written in Go that ships as a single binary with its own interface, under the MIT licence. It is built for teams that want data discovery without enterprise infrastructure, and it catalogs tables, topics, queues, APIs and dashboards.

### How does Marmot connect to AI agents?

It exposes certified context through MCP, the Model Context Protocol, as well as through the API and the interface, so an agent can query real metadata about data assets and trace lineage instead of guessing. The container image carries an MCP server name label.

### What can Marmot catalog?

Through its plugin system: tables from databases and data warehouses, topics from message queues and event streams, queues for job and messaging systems, APIs including REST, GraphQL and internal services, and dashboards from visualization tools and BI platforms.

### How do I deploy Marmot?

Follow the deployment documentation for a guided setup, or download the single binary for your platform and run it. The container image is Alpine-based, exposes port 8080 and starts as a non-root user with the run command.

### Is Marmot free to use?

Yes. It is MIT licensed, so you can use, modify and distribute it freely, and self-hosting it costs nothing in licensing terms.

### Can Marmot show what a change to a data asset would affect?

Yes. Interactive lineage traces data flows from source to destination, which the project describes as a way to analyse impact before making changes.

## Sources

- [License: MIT](https://github.com/marmotdata/marmot/blob/main/LICENSE)
- [marmotdata/marmot on GitHub](https://github.com/marmotdata/marmot)
- [Project website](https://marmotdata.io)
- [README](https://github.com/marmotdata/marmot/blob/main/README.md)
- [Releases](https://github.com/marmotdata/marmot/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/marmotdata-marmot
