# Chronon makes point-in-time correctness the feature you can measure

> Airbnb's Chronon is a feature platform where you define features as transformations of raw data and the system handles batch, streaming, backfills and serving. The part that matters is not the aggregations, it is the claim that backfilled training data and online serving agree, and that Chronon will show you when they do not.

**airbnb/chronon** — Chronon is a data platform for serving for AI/ML applications.

- Repository: https://github.com/airbnb/chronon
- Stars: 1,059 · Forks: 108
- Language: Scala
- License: Apache-2.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/airbnb-chronon

## Features are declared as transformations, and that is the whole interface

The platform description is a single idea stated plainly: users define features as transformations of raw data, and Chronon performs batch and streaming computation, backfills, low-latency serving, correctness and consistency, plus monitoring. The readme then says the point is that you can use data from batch tables, event streams or services without handling the orchestration that usually entails. The quickstart makes the interface concrete. A source is declared with an event source holding a table name, a topic, a query with a select list and a time column. A feature set is then a group-by declaration: a list of sources, a key, and a list of aggregations, each with an input column, an operation and a set of windows. Window sizes are themselves declared as values, in the example three windows of three, fourteen and thirty days. That is a declarative surface, not a procedural one, and the consequence is that the same declaration can be executed twice by different engines. That is the mechanism behind the rest of the platform: if batch and streaming read the same declaration, they cannot disagree about what a feature means, only about how recently it is current. Whether they agree on the values is what the consistency tooling then measures, and that distinction between shared definition and verified values is the thing to hold on to.

## Point-in-time accuracy is the claim worth interrogating

The backfill section is short and the adjectives are strong: scalable for large time windows, resilient to highly skewed data, and point-in-time accurate such that consistency with online serving is guaranteed. The third claim is the one that matters and the one that deserves a concrete definition, because point-in-time accuracy has a precise meaning in feature stores and a vague one everywhere else. It means that when you compute a feature value for a training example, you use only the source rows that existed as of the timestamp of that example, not the whole history. Get that wrong and your model is trained on information it will not have at inference, which is the training-serving skew problem, and the symptom is a model that performs worse in production than its offline evaluation predicted, for reasons that are hard to trace. A system that computes backfills point-in-time correctly is making a specific promise: the value a model sees during training is the value the serving path would have produced at that moment. The readme also claims the backfills are resilient to highly skewed data, which is the other hard part, since a key with one enormous history will not parallelise the way a uniform key will. Both claims are architectural rather than incidental, and both are the reason to consider this platform at all.

## Consistency monitoring turns a correctness claim into an alert

The observability section lists two things and both are the interesting kind. The first is data freshness, described as ensuring that online values are being updated in realtime, which is a liveness check on the pipeline rather than a correctness check. The second is online versus offline consistency, described as ensuring that backfill data for model training and evaluation is consistent with what is being observed in online serving. That second item is the platform's most valuable output and the reason to adopt it even if you end up writing your own transformation layer. A correctness guarantee that only holds on average is not a guarantee anyone can act on, so a system that continuously compares what serving returns against what a backfill would produce for the same key and timestamp is doing something your CI cannot. The output is a divergence rate, and the engineering discipline is deciding what to do when it moves. In practice that means a threshold, an owner, and a decision about whether to block a model promotion or merely file a ticket. The readme does not describe the alerting surface or the query interface for the comparison, so the mechanics are documentation work rather than something the readme settles. What is clear from the structure is that this is a first-class feature of the platform rather than a diagnostic script.

## The engine story is plural, and that is both strength and cost

The top-level layout answers a question the readme does not: which compute engines does this actually run on. There are separate modules for Spark, for Flink, for Airflow, plus an aggregator module, an online serving module, a service module, an API directory, a JVM directory, a project directory and a tools directory. So a feature declared once can be computed by batch, by streaming, and orchestrated by a scheduler, with the online side reading from a key-value store. That plurality is a genuine advantage over building the pipeline yourself, and it is also the cost: a platform with a Spark path, a Flink path and an online path has three places for a version mismatch to appear, and you will pick one rather than use all of them. Look at the quickstart compose file to see which one the project itself tests. It runs a main service whose command starts a Spark shell against a data loader script, sets a Spark version of 3.1.1, sets a local job mode with a parallelism of two, and configures executor memory and cores. The compose file also brings up ZooKeeper and Kafka, with a topic created for return events, and MongoDB as the key-value store. So the reference path is Spark plus Kafka plus a document store, and the streaming path is present in the layout even though the quickstart explicitly excludes running streaming jobs.

## The quickstart is a shell in a container, and the compose file is the install

There is no package install step, and the readme is explicit about that. The setup instructions are to download a compose file and run it, and the command given fetches the compose file from a hosted URL before bringing the stack up:

```bash
curl -o docker-compose.yml https://chronon.ai/docker-compose.yml
docker-compose up
```

The shell used for the rest of the tutorial runs inside the main container:

```bash
docker-compose exec main bash
```
 Once you see printed data with a notice that only the top twenty rows are shown, the tutorial considers you ready. That framing is honest about what the quickstart is: a self-contained demonstration with fabricated data, not a deployment. Docker is the only stated requirement. To work on features you then open a shell inside the main container, and the readme notes that the Python feature definitions are already in the image, with nothing to run until the backfill step, which tells you the authoring step is editing files rather than executing commands. The sample data is four tables chosen to model the problem well: a users table that is a daily batch source, purchases and returns as log tables each with a streaming topic counterpart, and a checkouts log that drives the model prediction. That last table is the modelling insight in the tutorial, since the prediction fires at checkout and the features are whatever was known about the user at that moment. A sample directory in the repository holds this data, and a sample group-by directory holds the feature definitions the quickstart refers to.

## Building the jar yourself, and what that says about the build

The container file at the top level is where a prospective adopter should stop and read carefully, because it is not a build of the repository. The base image is a slim Java 8 runtime, and the very next line is an environment variable holding a path to a jar the comment says you must set manually and which requires a local build of that jar. So this image is a development and test environment built on top of an artefact you produce yourself, not something that fetches a released binary. From there the file installs Python and build tooling, downloads and compiles Thrift from an Apache archive at a pinned version with two language bindings disabled, downloads and installs a Scala package from a Lightbend download, and then fetches Spark and Hadoop archives at pinned versions and unpacks them. Every external artefact is pinned and every download comes from an archive, which is the pattern of a project that wants reproducible environments at the cost of carrying several years of pinned infrastructure. The build itself is Bazel, with configuration files for Bazel and a version file, alongside an sbt build, a Maven template, formatter configuration for Scala, a tox configuration for the Python side, and both a requirements file and a lock file. A shell script in the compose file runs a Spark data loader at startup.

## Version numbering is the honest signal here

The release history deserves plain statement rather than reassurance. The published versions are v0.0.101, v0.0.93 and v0.0.89, with the newest released in July 2025 after earlier ones in April and March of the same year. So the project is below one hundredth of a version, three digits into the patch position, and the numbering suggests a project that increments on every merge rather than cutting releases at compatibility boundaries. The last push to the master-equivalent branch was on 2026-09-18, so the code is moving. For an adopter, the honest reading is that you should pin an exact version and expect the API to move between pins, and that you should read the release notes for the specific version rather than assuming a minor bump is safe. There is a release document and a development notes file in the repository, plus a governance file, a project management committee reference and an authors file, which describes a project with process rather than a solo experiment. The licence is Apache-2.0, which is the permissive grant most companies require for platform code, and the project also publishes to a package repository with shell scripts for uploading artefacts, so it is consumable as a dependency rather than only as a fork.

## Conclusion

Adopt Chronon if you have a machine learning feature pipeline where training data and serving features are computed by different code and have silently drifted, since point-in-time backfill correctness with a built-in consistency check is the one problem this platform is shaped around and the one that quietly poisons models. Do not adopt it for a small team with a handful of features, because the operational surface includes Spark, Kafka, a JVM toolchain and Thrift, and the version numbering below one with a hundred releases is a signal about API stability. Four things to verify. That your compute engines are ones Chronon already speaks, since the repository has separate modules for Spark, Flink, Airflow and a streaming path. Which online store you will use, because the quickstart wires MongoDB through a class you supply and production deployments will differ. That the release you pin has documentation matching your engine versions. And what your consistency monitoring will alert on, since the platform reports online versus offline divergence and that signal is only useful if something acts on it. The licence is Apache-2.0, the last push was on 2026-09-18, and the newest published release is v0.0.101 from 2025-07-10.

## FAQ

### What is Chronon and who is it for?

Chronon is described as a data platform for AI and ML applications that abstracts away the complexity of data computation and serving. Users define features as transformations of raw data, and the platform performs batch and streaming computation, backfills, low-latency serving and consistency checks.

### How do I start the Chronon quickstart?

Docker is the only stated requirement. Download the compose file with curl and bring the stack up with docker-compose, then open a shell in the main container. You are ready once data prints with a notice that only the top twenty rows are shown.

### What is point-in-time accurate backfilling in Chronon?

The readme states that backfills are point-in-time accurate such that consistency with online serving is guaranteed. It also claims they are resilient to highly skewed data, and the observability section covers data freshness plus online versus offline consistency checks.

### Which compute engines does Chronon support?

The repository has separate modules for Spark, Flink, Airflow, an aggregator, an online serving module and a service module. The quickstart compose file runs Spark and Kafka with MongoDB as the key-value store, and the quickstart itself does not cover running streaming jobs.

### What licence is Chronon released under?

Apache-2.0. The published versions are all below one, with the newest at v0.0.101 from 2025-07-10, while the last push to the main branch was on 2026-09-18, so you should pin an exact version rather than track the branch.

## Sources

- [airbnb/chronon on GitHub](https://github.com/airbnb/chronon)
- [Issues](https://github.com/airbnb/chronon/issues)
- [License: Apache-2.0](https://github.com/airbnb/chronon/blob/main/LICENSE)
- [README](https://github.com/airbnb/chronon/blob/main/README.md)
- [Releases](https://github.com/airbnb/chronon/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/airbnb-chronon
