Self-hosted service
GreptimeTeam/greptimedb avatar
GreptimeTeam/greptimedb

GreptimeDB: one columnar engine for metrics, logs and traces on object storage

The open-source observability database. One columnar engine for metrics, logs, and traces, on object storage.

6,724 stars553 forksRustApache-2.0

At a glance

What is it?
GreptimeDB puts metrics, logs and traces in one columnar engine over object storage, with SQL and PromQL on top. It is a real consolidation play for small observability teams, and a poor fit if you expect PromQL to behave exactly like Prometheus.
Who is it for?
Adopt GreptimeDB if you already run Prometheus plus Loki or Elasticsearch and want one Apache-2.0 backend that ingests OTLP, Prometheus remote write and Loki push while answering SQL and PromQL against the same tables. Do not adopt it if exact PromQL parity with Prometheus is a hard requirement, or if you are not prepared to operate object storage and its caches.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The three-backend problem GreptimeDB targets

The README's "Why You Might Use It" list is unusually direct about the target user: someone running Prometheus plus Loki or Elasticsearch who wants one backend instead of three, or someone who has outgrown Prometheus on cardinality or retention and does not want the Thanos or Mimir operational surface. That is a narrower audience than the phrase "observability database" suggests. GreptimeDB is not aimed at teams that have not yet built a telemetry stack; it is aimed at teams that already have one and are paying for it in operational attention.

The consolidation argument rests on a shared table model: tags, a timestamp, and fields. When metrics, logs and traces carry common identifiers such as service, host, or trace ID, the README states you can correlate them in SQL without moving data between databases. That is the concrete claim worth testing. A trace ID in a log line and the same ID in a span table only help if both land in the same engine, and that is the thing a Prometheus plus Loki pair cannot do without a join in a third system.

The README also lists two less obvious audiences: teams storing GenAI or agent telemetry under the OpenTelemetry GenAI conventions alongside infrastructure signals, and teams that need the same engine and semantics on resource-constrained devices. The second one is a real constraint, not marketing: it means the same code path is expected to run where a full distributed deployment would not fit.

How the engine is put together: datanode, frontend, meta, object store

The Cargo workspace is the clearest architecture document in the repository. It contains separate crates for src/frontend, src/datanode, src/meta-srv, src/meta-client, src/object-store, src/mito2, src/promql, src/flow, src/log-store, src/metric-engine and src/standalone. Read together, they describe a disaggregated system: a frontend that accepts queries and ingestion, datanodes that hold compute and caches, a meta service for cluster metadata, and an object store crate that abstracts S3, GCS, Azure Blob and S3-compatible endpoints as primary storage.

The README states compute and storage are disaggregated: object storage holds the data, while memory and local-disk caches keep recent and frequently queried data close to compute. That sentence carries the main performance trade-off. Object storage is cheap and effectively unbounded, but every query that misses the cache pays a fetch. The caches are therefore not an optimization detail, they are part of the capacity plan.

The protocol surface is split across crates too. There is a dedicated src/promql crate and a src/log-query crate, which matches the README's claim that SQL and PromQL are both first-class query paths rather than a translation layer bolted onto one of them. On the ingest side the README lists OpenTelemetry (OTLP), Prometheus Remote Write, Loki Push, Elasticsearch Bulk, InfluxDB line protocol and gRPC. Built-in features include retention policies, downsampling, continuous aggregation, explicit table partitioning, and inverted, skipping and fulltext indexes. The src/flow crate is presumably where continuous aggregation lives, though the README does not spell out that mapping.

Installing GreptimeDB with Docker and pointing a client at it

The README points to the user guide at docs.greptime.com for the overview, and the repository ships a docker/ directory plus a Docker Hub image under the greptime/greptimedb namespace. The README does not reproduce a full quickstart command sequence, so treat the user guide as the authoritative install path rather than a blog post. What the repository does show is the image name and the MySQL wire protocol address used in test configuration.

The .env.example file sets GT_MYSQL_ADDR to localhost:4002, which tells you the MySQL-compatible endpoint is expected on port 4002. That is the address a client uses to connect, and it is the fastest way to confirm the server is up after starting it. The repository also ships a config/ directory and a rust-toolchain.toml, so a source build is reproducible against a pinned toolchain.

Once the container is running, connect with any MySQL client to port 4002. The README describes the shared table model as tags, a timestamp, and fields, and lists SQL as a supported query path, but it does not include a CREATE TABLE example, so the exact schema syntax belongs in the user guide rather than here. If a query returns nothing, the first thing to check is whether you connected to port 4002 rather than a gRPC port; the README lists gRPC as a separate ingest protocol, and mixing them up is the usual cause of an empty result.

Where the compatibility table stops being a comfort

The README is candid in a way that most database READMEs are not. It states plainly that compatibility is per protocol, and that query-side coverage is narrower than ingestion. The table has a Compatible column and a Not compatible column, and the excerpt cuts off mid-header, so the list of gaps is not visible here. That truncation is itself the warning: the gaps exist and are enumerated, and anyone planning a migration needs to read that table in full before committing.

The practical consequence is that "drop-in replacement" is the wrong mental model. Ingestion is the easy half. If your collectors speak OTLP, Prometheus remote write or Loki push, they will likely keep working. The query half is where teams get surprised, because dashboards and alerts encode assumptions about a specific query dialect and function set. The README's own migration advice reflects this: migrate ingestion one signal at a time without rebuilding your collectors. It does not promise that your queries survive unchanged.

A second limitation is structural rather than documented. Because object storage is primary storage and caches are memory and local disk, a deployment is only as fast as its cache hit rate for the queries you actually run. The README does not publish cache sizing guidance in the excerpt available, and it does not claim a specific retention or cardinality ceiling. Anyone who needs a number before committing should look for it in the deployment and administration documentation rather than in the README.

GreptimeDB vs InfluxDB and vs Prometheus: different bets

The most common comparison question around this project is GreptimeDB vs InfluxDB, and the difference is architectural rather than cosmetic. InfluxDB is a time series database built around its own line protocol and query languages. GreptimeDB accepts the InfluxDB line protocol as one of six listed ingest paths, which means it can sit behind existing InfluxDB-style writers, but its own query surface is SQL and PromQL. The bet GreptimeDB makes is that SQL is the durable interface for telemetry, and that a columnar engine over object storage is the right storage substrate. InfluxDB's bet is closer to a purpose-built time series store with its own query model. If your team already writes Flux or InfluxQL heavily, that code does not transfer.

The GreptimeDB vs Prometheus question is a different kind of comparison, because Prometheus is not really a long-term store. The README frames the migration case as outgrowing Prometheus on cardinality or retention without taking on the Thanos or Mimir operational surface. GreptimeDB accepts Prometheus remote write and answers PromQL, so the collector side stays put. The trade is that you are now operating an object storage backed system with caches, which is a different failure profile from a single Prometheus process with a local disk. Whether that trade is good depends on whether your pain is cardinality and retention or operational simplicity.

Against ClickHouse, the split is about what comes in the box. ClickHouse is a general columnar engine that teams assemble into an observability stack with surrounding services for ingestion, query translation and retention. GreptimeDB ships PromQL, Loki push, OTLP and Jaeger-compatible trace queries as part of the same binary, and the README lists retention policies, downsampling and continuous aggregation as built-in. That is convenience bought at the cost of the flexibility a general engine gives you.

Licence, releases and what upgrades cost

The core is Apache-2.0, stated in both the README and the workspace manifest, and the repository carries a LICENSE file plus an AUTHOR.md. There is also a LICENSE-ENTERPRISE file and a licenserc-enterprise.toml alongside licenserc.toml, which indicates an enterprise edition boundary exists. The README has a section titled "Limitations and Edition Boundary", so the split between the Apache-2.0 core and the commercial edition is documented rather than implied. Anyone evaluating this for production should read that section before assuming every listed feature is in the open source build. This is a description of what the repository contains, not legal advice; licence questions belong with your own counsel.

Release cadence is visible in three channels. The README describes stable as the production channel, canary as including pre-releases, and nightly as a weekly snapshot of main. The recent releases bear this out: v1.2.0 is a stable release, while v1.3.0-alpha.1 and a nightly build tagged v1.3.0-alpha.1-nightly-20260907 sit alongside it. The workspace manifest pins the development version at 1.3.0-alpha.1, so main is ahead of the stable line.

Upgrade cost depends on which channel you run. Staying on stable means waiting for the next stable tag. Running canary or nightly means tracking a moving target, and the nightly naming convention makes it obvious that those builds are snapshots rather than supported releases. The repository ships rust-toolchain.toml and a Cargo.lock, so source builds are reproducible against a pinned toolchain, but the README does not document rollback or downgrade procedures in the excerpt available.

Editorial conclusion

Adopt GreptimeDB if you already run Prometheus plus Loki or Elasticsearch and want one Apache-2.0 backend that ingests OTLP, Prometheus remote write and Loki push while answering SQL and PromQL against the same tables. Do not adopt it if exact PromQL parity with Prometheus is a hard requirement, or if you are not prepared to operate object storage and its caches. Verify first which protocols your collectors actually use and how much of that surface the compatibility table marks as not compatible, then check the storage options in the deployment configuration before sizing anything.

Frequently asked questions

How does GreptimeDB compare with InfluxDB?

GreptimeDB accepts the InfluxDB line protocol as one of its ingest paths, but its own query surface is SQL and PromQL rather than InfluxQL or Flux. The deeper difference is storage: GreptimeDB uses object storage as primary storage with memory and local-disk caches, and it presents metrics, logs and traces through one shared table model of tags, timestamp and fields.

How does GreptimeDB compare with Prometheus?

GreptimeDB accepts Prometheus remote write and answers PromQL, so existing collectors can keep working. The README frames the case for switching as having outgrown Prometheus on cardinality or retention without wanting to operate Thanos or Mimir, which means the trade is added storage and cache infrastructure in exchange for longer retention and higher cardinality.

What are the alternatives to GreptimeDB?

The README positions it against running Prometheus plus Loki or Elasticsearch, and against adding the Thanos or Mimir operational surface to Prometheus. It also notes that teams wanting a general columnar engine assemble their own stack around one, whereas GreptimeDB ships PromQL, Loki push, OTLP and Jaeger-compatible trace queries in the same system.

Official sources

  1. GreptimeTeam/greptimedb on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/greptimeteam-greptimedb.svg)](https://hysenlabs.com/projects/greptimeteam-greptimedb)