# Timely Dataflow in Rust: A Low-Latency Cyclic Dataflow Engine

> Timely dataflow is a Rust workspace for building data-parallel programs that scale from one thread to a cluster. It is expressive and fast, but it is a framework for writing your own operators, not a ready-made query engine.

**TimelyDataflow/timely-dataflow** — A modular implementation of timely dataflow in Rust

- Repository: https://github.com/TimelyDataflow/timely-dataflow
- Stars: 3,649 · Forks: 295
- Language: Rust
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/timelydataflow-timely-dataflow

## What timely dataflow solves, and who it is for

The README describes timely dataflow as a low-latency cyclic dataflow computational model, introduced in the Naiad paper, and this repository as an extended and more modular implementation in Rust. The problem it addresses is coordination. In a data-parallel program you need to know when a round of input has been fully processed before you emit a result, and that question gets harder as you add threads and machines. Timely answers it with progress tracking on timestamps rather than with a central coordinator.

The audience is narrow and specific. If you are building a streaming or iterative compute engine in Rust, and you want to define your own operators rather than pick from a fixed catalogue, this is the layer you build on. The README states the main goals are expressive power and high performance. It also says the project is intended to support multiple levels of abstraction, from manual dataflow assembly up to higher-level declarative layers, and that the set of such layers is expected to expand as people write them. That sentence is the honest description of the project's position: it is infrastructure, and the ecosystem above it is thin.

If you want a system that answers SQL or runs a job you describe in a config file, this is not it. Nothing in the README promises a query planner, a scheduler, or a storage layer.

## How progress tracking and worker exchange actually work

The hello example in the README shows the whole mechanism in about twenty lines. A worker creates an InputHandle and a ProbeHandle. It builds a dataflow with input_from, then exchange, then inspect, then probe_with. The exchange operator routes records between workers using the data itself: the closure returns the value that decides the destination. In the example the record is its own routing key, so worker 0 sends 0, 2, 4 and so on and worker 1 receives 1, 3, 5.

The interesting part is the loop below the dataflow. Each round, worker zero sends a value, then every worker calls advance_to on its input handle with round + 1. That advances the worker's notion of the current timestamp. Then each worker spins on worker.step() while probe.less_than(input.time()) is true. The probe reports whether any data at a timestamp earlier than the current one is still in flight anywhere in the cluster. When the probe says no, the round is complete and the next one can start.

That is the whole coordination story: no barrier service, no global lock, just each worker stepping its own dataflow and asking a local probe about global progress. It is why the README can claim low latency. It is also why the model demands discipline from you. If you send data without advancing the timestamp, or advance it too far, the probe's answer stops meaning what you think it means. The README does not document what happens when you get that wrong.

## Installing timely and running the hello example with two workers

Timely is published on crates.io. The README gives the dependency line as timely="*", which pulls the current release. For a real project you would pin a version, but the README does not recommend one, so the wildcard is what is documented.

```toml
[dependencies]
timely="*"
```

With that in Cargo.toml, the smallest program is the simple example. It uses timely::example, which builds a scope for you, turns a range into a stream, and inspects each record.

```rust
use timely::dataflow::operators::*;

fn main() {
    timely::example(|scope| {
        (0..10).to_stream(scope)
               .container::<Vec<_>>()
               .inspect(|x| println!("seen: {:?}", x));
    });
}
```

From the root of the repository you run cargo run --example simple and the README shows ten lines of output, seen: 0 through seen: 9. Note that simple always uses one worker thread, because timely::example ignores user-supplied arguments.

For the multi-worker case, use hello. The worker count is set with -w or --workers. The README's two-worker run is cargo run --example hello -- -w2, and the output alternates worker 0: hello 0, worker 1: hello 1, and so on through nine. That alternating output is the exchange operator doing its job, not a printing artifact.

For multiple processes, the README specifies a hostfile of hostname:port lines, an -n or --processes count, and a -p or --process index per process. Each process also takes -w. The README states the number of workers should be the same for each process, which is a constraint worth respecting on the first attempt.

```text
host0% cargo run -- -w 2 -h hostfile.txt -n 4 -p 0
```

## The limitation: it is a library, and the documentation says so

The README is unusually candid. It calls the documentation for timely dataflow a work in progress, though mostly improving. It points to long-form text in mdbook format with examples tested against current builds, and separately to a three-part blog series, with the warning that the examples there may need tweaks to build against the current code. If you are the kind of reader who wants a stable, complete reference before writing code, take that warning seriously.

The deeper limitation is scope. Timely gives you streams, timestamps, progress tracking and exchange. It does not give you windowing, joins, aggregation or persistence. Those exist in higher layers, and the README says the set of layers is expected to expand as interested people write their own. That is a statement about the present, not a roadmap. If your problem is a join over two streams, you will be writing the join.

There is also a versioning cost visible in the repository itself. The workspace has six members: bytes, communication, container, logging, mdbook and timely. The recent releases show timely_container, timely_communication and timely_logging all at v0.31.0 on the same date. The crates move as a set. A wildcard dependency in your Cargo.toml will follow that set forward, and the README does not describe an upgrade path or a compatibility policy. Pin your versions deliberately.

## Timely dataflow compared with differential dataflow and Flink

The most common comparison is with differential dataflow, and it is not really a rivalry. Differential dataflow is a layer built on top of timely dataflow, aimed at incremental computation where inputs change over time and you want to maintain a result rather than recompute it. If your problem is incremental, you want the layer, not the substrate. Choosing timely means you are building the layer.

The other comparison people reach for is Flink, and the difference in approach is structural. Flink is a complete system: you submit a job, it manages execution, state and recovery. Timely is a Rust crate you link into your own binary. There is no cluster daemon to run and no job submission protocol. You compile a program, and that program is the cluster. That gives you control over the dataflow graph and the data representation, and it takes away every operational convenience a job manager provides. Neither is better in the abstract; they answer different questions.

One more boundary worth naming: timely dataflow is not a Python tool. The README documents Rust only, and the repository's workspace members are Rust crates. Searches for a Python interface have no answer in this material.

## Licence, workspace layout and upgrade cost

The repository is MIT licensed, with a LICENSE file and a COPYRIGHT file at the top level. MIT is permissive: it allows commercial use and modification, and it requires that the copyright notice and permission notice be preserved in copies. That is the standard reading of the text, not legal advice; if you are redistributing a modified timely inside a product, have your own counsel read the file rather than this paragraph.

The workspace pins edition 2021 and rust-version 1.86, so a toolchain older than that will not build the workspace. The workspace also carries a long list of clippy lints set to warn, including clone_on_ref_ptr and needless_pass_by_ref_mut. If you vendor these crates into your own workspace, expect those warnings to surface in your builds unless you override them.

Upgrade cost is the piece the repository does not document. There is a CHANGELOG.md and a release-plz.toml, which suggests releases are automated, but the README says nothing about a stability guarantee or a migration process between the v0.31.0 line and whatever follows. Treat the version as something you pin explicitly and bump on your own schedule, testing your dataflow against the new version rather than assuming the operators behave identically.

## Conclusion

Adopt timely dataflow if you are writing a Rust data-parallel engine and need control over timestamps, progress tracking and worker-to-worker exchange. Do not adopt it if you want a query language, a scheduler, or a Python API; there is none here. Before committing, build the hello example with two workers and confirm the probe behaves as the README describes, and check that the crate versions you pin are compatible across the workspace.

## FAQ

### What is timely dataflow used for?

It is used to build data-parallel programs that scale from a single thread to a cluster, with progress tracking that tells each worker when a round of data is fully processed. The README describes it as a low-latency cyclic dataflow computational model and says the main goals are expressive power and high performance.

### How do I install timely dataflow in a Rust project?

Add timely to the dependencies section of your Cargo.toml. The README gives the line as timely="*", which brings in the timely crate from crates.io.

### Can I run timely dataflow across multiple machines?

Yes. The README describes a hostfile of hostname:port lines, with -n or --processes giving the number of processes and -p or --process giving each process its index. Each process also takes -w for its worker count, and the README states the number of workers should be the same for each process.

### Is there a Python interface for timely dataflow?

The README documents Rust only, and the repository's workspace members are Rust crates. No Python binding is described in this material.

## Sources

- [Issues](https://github.com/TimelyDataflow/timely-dataflow/issues)
- [License: MIT](https://github.com/TimelyDataflow/timely-dataflow/blob/master/LICENSE)
- [README](https://github.com/TimelyDataflow/timely-dataflow/blob/master/README.md)
- [Releases](https://github.com/TimelyDataflow/timely-dataflow/releases)
- [TimelyDataflow/timely-dataflow on GitHub](https://github.com/TimelyDataflow/timely-dataflow)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/timelydataflow-timely-dataflow
