# Ten traces, twelve directories, and a survey between you and the data

> Alibaba's cluster trace programme is a research dataset rather than software, and its value is the specificity of what each release contains: 1300 machines for twelve hours, 4000 for eight days, 6500 GPUs for two months, 155,410 GPUs across six months. The friction is that access to the earlier releases goes through a web form, and the licence is an instruction in the readme rather than a file.

**alibaba/clusterdata** — cluster data collected from production clusters in Alibaba for cluster management research

- Repository: https://github.com/alibaba/clusterdata
- Stars: 2,211 · Forks: 485
- Language: Jupyter Notebook
- License: not declared
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/alibaba-clusterdata

## Each release is a paper with a dataset attached, and the scale is stated

The readme lists every trace with its scale, its duration and the paper that motivated it, and that pairing is the format to expect from a research programme rather than a product. The first release covers about 1300 machines over twelve hours and introduces the co-location of long-running online services with batch workloads, which is the structural feature that makes it interesting. The second is larger at about 4000 machines over eight days and adds the directed acyclic graph structure of production batch jobs, so you can see not just that a job ran but what it depended on. Each subsequent entry extends a different axis, and reading them in order shows the questions the group was working on. A microarchitecture release adds fine-grained processor metrics, with the stated purpose of studying processor performance, architectural contention and memory bandwidth contention in a co-located datacentre. A microservices release adds runtime metrics for more than 20000 services over twelve hours, including call dependencies, response times and call rates. The GPU releases then take over the list, moving from a machine-learning-as-a-service platform with over 6500 GPUs on about 1800 machines across two months, to a heterogeneous cluster with over 6200 GPUs on about 1200 machines, to a disaggregated serving study of over 150 inference services and more than 20k inference instances. The scale numbers are the reason to use this data, and they are in the readme rather than hidden in a schema.

## A survey form stands between you and most of the downloads

The access mechanism is the practical detail and it is described in passing, with a form to complete before the link is released:

```bash
https://goo.gl/forms/eOoe6DwZQpd2H5n53
```
 For the earliest release the readme says a download link is available after a short survey and gives a form link. For the second it says a download link is available after a survey, and adds that it takes less than a minute. Later entries do not repeat the requirement, which may mean the arrangement changed or that the readme was not updated, and either way you should not assume a direct link. What the survey is for is not stated, and that is worth thinking about before you fill one in: the motivation section links to a page about the challenges Alibaba faces in datacentres, and the stated purpose of releasing the data is to help people in the field understand modern datacentres and to provide production data for researchers to verify their ideas. So the programme's motivation is external validation of their own operational understanding, and a form before the download is consistent with tracking who is using the data. The readme also asks to be told when a publication using the trace is available, because the maintainers keep a list of related publications so researchers can find each other. That is a request rather than a licence, and it tells you the relationship the programme wants with users: a visible research community rather than anonymous consumption.

## Twelve directories, ten descriptions, and a gap worth noticing

The top-level listing is a directory per trace and it is worth counting against the readme, because the two do not match. The readme describes ten releases, and the listing contains twelve directories. The two that are not described are a second microservices release with a 2022 date, and a 2026 spot GPU release. So either the readme is behind the data, or those two are staged and not yet announced. Either way, if you are looking for the most recent data the readme's own list is not authoritative and the directory listing is. The listing also confirms that each trace gets its own subdirectory with its own readme, which the overview text supports by pointing at subdirectories for the data, schema and processing scripts of the GPU releases. That per-trace structure matters in practice. A schema from 2017 and a schema from 2026 describing GPU utilisation and network topology are not the same shape, and a research pipeline that reads both needs to handle that, so the per-trace documentation is not boilerplate. The two newest entries also promise something the older ones do not: figure reproduction scripts, named in the descriptions of both the newest GPU release and the model-as-a-service release, alongside the data and schema. That is a real improvement, because it means a paper's figures can be regenerated rather than reimplemented.

## The newest traces describe two different kinds of system

The last two entries are the ones that tell you where this programme is now, and they are not the same kind of measurement. One is a six-month GPU cluster trace from a serverless infrastructure, covering up to 155,410 GPUs across 37,707 servers, and it comes with fact tables for the workload mix, job and model types, priority classes, resource requests and utilisation, the GPU and server inventory, network topology, and execution-time analysis. That is a scheduling study, and the accompanying paper is about heterogeneity at scale and how such a cluster should be scheduled. The other is described as a top-down view of a stable diffusion serving system, with performance data captured at three layers: the application layer with user requests and end-to-end latency, the middleware layer with gateway queues, schedulers and pipeline management, and the infrastructure layer with container resources, GPU utilisation and memory usage. A three-layer view is a different kind of dataset from a fact table. It lets you correlate an end-to-end latency with a queue depth and a resource measurement, which is how you find out whether slowness is the model, the queue or the hardware. The third new one is production large language model inference, online and offline batch, again over six months, with the stated purpose of characterising inference clusters at scale. Together the three say the programme has moved from describing datacentre workloads in general to instrumenting specific inference pipelines layer by layer.

## A repository with no code, and a licence that is a sentence

The file listing is worth pausing on. There is a readme, a desktop metadata file that should not have been committed, and twelve trace directories. There is no library, no package manifest for the data, no build configuration, no test suite and no continuous integration, and the primary language is recorded as a notebook format, which reflects the analysis material inside the trace directories rather than anything you would depend on. This is a data repository, and the way to use it is to read the per-trace schema, then load the files into whatever you already use. The absence of a licence file is the one thing that would stop me short of using it. The readme's terms are stated in a sentence in the motivation section: you may use the trace however you want as long as it is for research or study purposes, and it is encouraged that you use it for research or study. That is a permission with a condition, not a licence, and there is no file to point a lawyer at. For academic work, where the condition matches your intent, that is sufficient in practice. For anything else, from a benchmark in a commercial product to a derived dataset you intend to publish, you would be relying on a sentence in a readme, and the right move is to email the address the readme gives before you build on it. The readme in fact invites exactly that, and asks for issues as the preferred channel because the discussion helps the whole community.

## Conclusion

Use the Alibaba traces if you are doing systems research where synthetic workloads will not do, because each release is tied to a published paper with a stated scale and duration, which means you can check whether the data matches your question before you build a pipeline on it. Do not use it as a benchmark of Alibaba's own hardware, since it is anonymised, aggregated and shaped by what the researchers chose to record, so it tells you about workload structure rather than about performance you could reproduce. Three things to know before you start. That access is not a download link, since the readme points at a survey form and asks you to complete it, so budget for that step rather than assuming a clone gives you the data. That the newest releases describe themselves differently from the older ones, with several now promising figure reproduction scripts alongside schemas, which is a meaningful improvement for anyone trying to check a paper's numbers. And that no licence file is declared, so the terms are the readme's statement that the data is for research or study, which is a permission rather than a licence. The last push was on 2026-09-24 and no releases are published.

## FAQ

### How do I download the Alibaba cluster traces?

Access goes through a survey form rather than a direct link for at least the earlier releases, which the readme describes as taking under a minute. Later entries are linked from their own subdirectories. The readme asks you to file an issue with questions rather than emailing, so discussion benefits everyone.

### What is in each cluster trace release?

Scale and duration are stated per release: about 1300 machines over twelve hours in the first, about 4000 machines over eight days in the second with batch job dependency graphs, over 6500 GPUs on about 1800 machines across two months, over 6200 GPUs on about 1200 machines, over 150 inference services, and a six-month trace covering up to 155,410 GPUs across 37,707 servers. Each has its own subdirectory with schema and processing details.

### What licence applies to the Alibaba cluster trace data?

No licence file is declared. The readme states you may use the trace however you want as long as it is for research or study purposes, which is a permission with a condition rather than a formal licence. The readme asks to be notified when a publication using the trace is available.

### Which cluster traces come with figure reproduction scripts?

The two newest releases described in the readme, the six-month GPU trace from the serverless infrastructure and the model-as-a-service trace of production large language model inference, both mention scripts for reproducing the paper's figures alongside the data download and schema.

### What is the newest trace in the Alibaba cluster data programme?

A six-month GPU cluster trace from the serverless infrastructure covering up to 155,410 GPUs across 37,707 servers, released with the OSDI paper on heterogeneity at hyperscale. The repository's last push was on 2026-09-24, and there are no published releases.

## Sources

- [alibaba/clusterdata on GitHub](https://github.com/alibaba/clusterdata)
- [Issues](https://github.com/alibaba/clusterdata/issues)
- [README](https://github.com/alibaba/clusterdata/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alibaba-clusterdata
