Marmot: a single-binary data catalog that exposes metadata to AI agents over MCP
The open-source context layer for your AI. Catalog your tables, topics, queues and APIs then expose real metadata to your AI agents.
At a glance
- What is it?
- Marmot is an MIT-licensed Go data catalog that indexes tables, topics, queues and APIs and serves that metadata to AI agents through MCP. It ships as one binary, but the current releases are previews and the README documents no installation commands.
- Who is it for?
- Adopt Marmot if you want a self-hosted catalog whose metadata is reachable by AI agents over MCP, and you are willing to run preview builds while v0.11 stabilizes. Do not adopt it if you need a documented, copy-paste install path today: the README hands deployment to the docs site and shows no commands, so verify the Deploy guide and the v0.11.0-preview3 release notes before committing.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Marmot targets: metadata that AI agents cannot see
Ask an AI agent to write a query against your warehouse and it will guess column names, invent join keys and miss the table that was deprecated last quarter. The model has no access to the catalog, so it works from the prompt alone. Marmot's stated purpose is to close that gap: catalog every data asset, attach the context that matters, and expose the result to both people and agents. The README frames the audience as teams that want data discovery without enterprise complexity, and it positions the project against catalogs that need extensive infrastructure and configuration. The asset types it names are tables from databases and warehouses, topics from message queues and event streams, queues from job and messaging systems, APIs including REST and GraphQL services, and dashboards from BI platforms. That list matters more than it looks. A catalog that only understands warehouse tables cannot answer an agent asking which Kafka topic feeds a given dashboard, and lineage across those boundaries is exactly where impact analysis gets hard.
The mechanism: Go server, Postgres, OpenSearch, plugins over gRPC and MCP on top
The repository layout tells you most of the architecture. cmd/ holds the entry point, internal/ holds the server, pkg/ holds shared code, plugins/ holds data source integrations and sdk/ holds the interfaces for writing new ones. The dependency list in go.mod is the clearest signal of how it runs. jackc/pgx and jackc/tern appear for PostgreSQL access and migrations, so a relational database is the system of record. opensearch-project/opensearch-go appears alongside the README's promise of full-text search with boolean logic and metadata filters, so search is delegated to OpenSearch rather than implemented in the binary. hashicorp/go-plugin and hashicorp/yamux point to out-of-process plugins talking over gRPC, which is how a Go catalog can ingest from systems written in anything. modelcontextprotocol/go-sdk is the MCP server implementation, and the Dockerfile labels the image with io.modelcontextprotocol.server.name set to io.github.marmotdata/marmot, so the container advertises itself as an MCP server. The data flow is therefore: a plugin connects to a source, reads its assets and metadata, writes them through the server into Postgres, the search index in OpenSearch makes them queryable, and MCP, the HTTP API and the web UI are three front doors onto the same store. That last point is the design bet. The UI is not the product; it is one client of an API that agents also use.
Installing Marmot: what the repository gives you and what it does not
The README does not contain install commands. It says new users should follow the Deploy documentation at marmotdata.io/docs/Deploy, and its FAQ offers two routes: that guide, or downloading and running the single binary for your platform. The repository also contains install.sh at the top level and a charts/ directory, which implies a shell installer and a Helm chart exist, but neither is documented in the README, so read them before running anything. What the repository does show precisely is the container build. The Dockerfile produces an image that runs as a non-root user and listens on port 8080, with the entrypoint fixed to the marmot binary and the default command set to run. If you already have a container runtime and a Postgres instance, that is the most concrete path the repository offers.
FROM alpine:3.23.3
LABEL io.modelcontextprotocol.server.name="io.github.marmotdata/marmot"
RUN apk upgrade --no-cache && apk add --no-cache ca-certificates tzdata
RUN adduser -D -u 10001 marmot
WORKDIR /app
COPY --from=builder /app/marmot /usr/local/bin/
RUN chmod +x /usr/local/bin/marmot && \
chown marmot:marmot /usr/local/bin/marmot
USER marmot
EXPOSE 8080
ENTRYPOINT ["/usr/local/bin/marmot"]
CMD ["run"]Building from source is documented in the Makefile. The build target compiles the binary into bin/marmot from cmd/main.go, and the dev target regenerates the Swagger spec first, builds with the swagger tag, then starts the server with debug logging, unencrypted connections allowed and telemetry disabled. Those three environment variables are the ones the project itself uses for local work, and they are worth copying exactly.
go build -ldflags '-s -w' -o bin/marmot cmd/main.go
MARMOT_LOGGING_LEVEL=debug MARMOT_SERVER_ALLOW_UNENCRYPTED=true MARMOT_TELEMETRY_ENABLED=false ./bin/marmot runThe first real use, based on the feature list, is to register a source through the plugin system, let it crawl, and then search. The README does not show the command or the config format for that, so the honest instruction is: open the Deploy guide, add a plugin for one source, and confirm the assets appear in the UI before wiring any agent to the MCP endpoint. The Dockerfile exposes port 8080, so that is where the server answers, but the README does not document the MCP endpoint path, and you should not guess it.
Where Marmot will disappoint you
The release history is the first warning. The three most recent releases are v0.11.0-preview1, v0.11.0-preview2 and v0.11.0-preview3, published on 2026-09-03, 2026-09-09 and 2026-09-10. Preview builds across a single week suggest an unstable interface, and the README does not describe a migration path between versions even though tern migrations are in the dependency tree. Back up the Postgres database before upgrading, because the project does not document rollback. The second issue is operational weight. The README's pitch is a single binary, and that is true of the server process, but the search experience depends on OpenSearch and the catalog depends on Postgres. A single binary that needs two backing services is a different proposition from a single binary you can run on a laptop, and the README does not discuss whether either is optional or how they are configured. Third, the plugin model means coverage is whatever plugins/ actually contains. The README lists asset categories, not connectors, so check the directory before assuming your source is supported. Finally, the documentation surface is split: the README defers deployment and local development to the docs site, and the FAQ's own links to GitHub Issues point at github.com/marmotdata/marmot while the Contributing section still says github.com/your-org/marmot, a placeholder that was never replaced.
Marmot versus a search-first catalog such as DataHub
The obvious comparison is with DataHub, the widely deployed open source catalog from LinkedIn's lineage. The difference is architectural rather than feature-by-feature. DataHub is a platform: a metadata service backed by Kafka for event streaming, a separate storage layer, a frontend, and an ingestion framework with its own CLI, deployed through Helm charts that assume a cluster. Marmot's README takes the opposite position, promising a single binary and an intuitive UI with no complex infrastructure. That trade shows up in the dependency list: Marmot uses OpenSearch for search and Postgres for storage, which is a smaller footprint than an event-streaming metadata backbone, and its ingestion is a plugin system over go-plugin rather than a separate framework. The cost is ecosystem depth. DataHub has had years of connectors and a public metadata model; Marmot's plugins/ directory is what it is today. If you need broad source coverage and a mature model, DataHub is the safer choice. If you want something a small team can run and point an MCP-capable agent at this week, Marmot's shape is the more direct fit.
Maintenance, licence and what an upgrade actually costs
The repository is not archived and the last push was on 2026-09-10, so the project is being worked on right now, with preview releases landing in the same week. That is a double-edged fact: fixes arrive quickly, and so do breaking changes. The version string is injected at build time through the internal/cmd.Version variable via ldflags, so a binary built without the Makefile will report v0.0.0 unless you set VERSION yourself. Upgrades run through tern migrations against Postgres, which means a schema change is a database operation, not just a binary swap. The licence is MIT, stated in both the README and the LICENSE file. MIT is permissive: you can use, modify and distribute the software, and self-hosting carries no licensing cost. What MIT does not give you is any warranty or support obligation from the maintainers, and it does not cover the plugins you might load, which live in their own repositories and may carry their own terms. Check each plugin's licence separately. This is not legal advice; read the LICENSE file and your own counsel's guidance.
Editorial conclusion
Adopt Marmot if you want a self-hosted catalog whose metadata is reachable by AI agents over MCP, and you are willing to run preview builds while v0.11 stabilizes. Do not adopt it if you need a documented, copy-paste install path today: the README hands deployment to the docs site and shows no commands, so verify the Deploy guide and the v0.11.0-preview3 release notes before committing. Check first whether the plugin you need exists in plugins/, and whether the chart in charts/ fits your cluster.
Frequently asked questions
What is Marmot and who is it for?
Marmot is an open-source data catalog that indexes tables, topics, queues, APIs and dashboards, then exposes that metadata to people and AI agents. The README targets teams that want data discovery without the infrastructure and configuration that enterprise catalogs require.
How do I install Marmot?
The README does not list install commands. It directs new users to the Deploy guide at marmotdata.io/docs/Deploy and mentions downloading and running the single binary for your platform; the repository also contains install.sh and a charts/ directory, neither of which the README documents.
Is Marmot free to self-host?
Yes. The README states Marmot is open-source software licensed under the MIT License, that you can use, modify and distribute it freely, and that self-hosting is completely free with no licensing costs.
How does Marmot connect to AI agents?
It exposes certified context through MCP, the Model Context Protocol, alongside an API and a UI. The dependency list includes the official Go MCP SDK, and the Docker image carries the label io.modelcontextprotocol.server.name set to io.github.marmotdata/marmot.
Is Marmot stable enough for production?
The three most recent releases are all v0.11.0 previews published between 2026-09-03 and 2026-09-10, and the README does not document a rollback or migration path between versions. The repository is not archived and the last push was on 2026-09-10, so work is ongoing, but preview versioning is the signal to test before depending on it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/marmotdata-marmot)