Model or dataset
microsoft/aurora avatar
microsoft/aurora

microsoft/aurora: a foundation model for Earth system forecasting

Implementation of the Aurora model for Earth system forecasting

1,020 stars175 forksPythonNOASSERTION

At a glance

What is it?
Aurora is a Python package that wraps pretrained weather, air pollution and ocean wave models behind a Batch and Metadata API. It is built for research reproducibility, not for operational forecasting.
Who is it for?
Adopt Aurora if you are reproducing the Nature paper, fine-tuning a foundation model on a specialised atmospheric task, or benchmarking against ERA5 in a research setting. Do not adopt it for operational forecasting, safety-critical planning or automated decision pipelines: the README states the code has not been developed or tested for non-academic purposes and that outputs are not meant to be used directly to plan operations.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem Aurora solves, and who it is actually for

Numerical weather prediction has traditionally meant solving the primitive equations on a supercomputer. Aurora takes a different route: it is a machine learning model that predicts atmospheric variables such as temperature, and it is a foundation model, meaning it was first trained broadly and can then be adapted to specialised forecasting tasks with relatively little task-specific data. The README lists four specialised versions: medium-resolution weather prediction, high-resolution weather prediction, air pollution prediction, and ocean wave prediction.

The intended audience is stated plainly in the Responsible AI documentation. The goal in publishing the code is to facilitate reproducibility of the paper and to support further research into foundation models for atmospheric forecasting. The same document says the code has not been developed nor tested for non-academic purposes. So the audience is researchers and engineers working on forecasting methods, not meteorologists running a shift. If you need a forecast you can hand to a dispatcher, this is the wrong layer of the stack.

How the model is structured: Batch, Metadata and forward

The public API is small. Inputs are wrapped in a Batch object that separates surface variables, static variables and atmospheric variables. Metadata carries the grid: latitude, longitude, a timestamp tuple, and the pressure levels for atmospheric variables. Calling model.forward(batch) returns a prediction object whose surf_vars dictionary holds the output tensors.

That separation matters because the four model variants share the same input contract while differing in what they predict and at what resolution. The repository layout reflects the same split: aurora/ holds the package, finetuning/ holds adaptation code, tests/ holds the test suite, and docs/ is built with Jupyter Book. The Makefile exposes install, test and docs targets, so the maintainers' own workflow is visible from the repository root rather than hidden in CI configuration. The dependency list in pyproject.toml is a reasonable signal of scope: numpy, scipy, torch, einops, timm, huggingface-hub, pydantic, xarray, netcdf4 and azure-storage-blob. Checkpoints are pulled through huggingface-hub, and azure-storage-blob appears because some data paths go through Azure.

Installing microsoft-aurora and running the small pretrained model

The README gives two install routes. The pip route is the one most readers will take:

bash
pip install microsoft-aurora

A conda or mamba route exists as well, using the conda-forge channel:

bash
mamba install microsoft-aurora -c conda-forge

The package requires Python 3.10 or newer according to pyproject.toml. For development work, the Makefile runs an editable install with the dev extra and sets up pre-commit hooks:

bash
pip install -e ".[dev]"
pre-commit install

The README's first real example loads the small pretrained model and runs it on random tensors, so you can confirm the install works before downloading real data:

python
from datetime import datetime

import torch

from aurora import AuroraSmallPretrained, Batch, Metadata

model = AuroraSmallPretrained()
model.load_checkpoint()

The README notes that this incurs a 500 MB download. After that, you build a Batch with surface variables 2t, 10u, 10v and msl, static variables lsm, z and slt, and atmospheric variables z, u, v, t and q, then call model.forward(batch) and read prediction.surf_vars["2t"]. The documentation points to a fuller example that runs the model on ERA5 at microsoft.github.io/aurora/example_era5.html. That page, not the README, is where the real end-to-end workflow lives.

The limitation the documentation itself admits

The Responsible AI section is unusually direct, and it is the most important part of the repository for anyone evaluating adoption. Aurora is based on neural networks, so there are no strict guarantees that predictions will always be accurate. The documentation goes further: altering the inputs, providing a sample that was not in the training set, or even providing a sample that was in the training set but is simply unlucky may result in arbitrarily poor predictions. It also notes the model may inherit biases present in any of the training datasets.

The stated out-of-scope uses are direct operational decision-making without expert review, applications requiring guaranteed forecast accuracy, and non-environmental prediction tasks. Safety-critical planning or automated decision pipelines should be accompanied by appropriate domain validation. That is a wide exclusion. If your product depends on a forecast being right, Aurora is a component to be validated, not a service to be trusted. The documentation also says the models published here are streamlined versions, which is a hint that what ships may not be identical to what the paper evaluated.

Aurora against numerical weather prediction and task-specific models

The obvious alternative is a traditional numerical weather prediction pipeline, where the forecast comes from solving physical equations on a large compute cluster. The difference in approach is fundamental: NWP encodes physics explicitly and produces a deterministic, inspectable computation, while Aurora learns a mapping from input fields to output fields and offers no such guarantees. NWP is expensive to run and hard to modify; Aurora is cheap to run once a checkpoint is loaded and easy to fine-tune, which is exactly why the finetuning/ directory exists.

The second alternative is a single-task machine learning model trained only for the variable you care about. Aurora's bet is the opposite: general pretraining first, adaptation second, so a specialised task needs relatively little data. That bet pays off when your task is close to one of the four supported ones and you lack the data to train from scratch. It pays off less when your target variable is far from anything in the pretraining mixture, because you are then fine-tuning a large model on a task it was never shaped for. The README does not document rollback or version pinning for checkpoints, so reproducibility across model versions is something you have to manage yourself.

Maintenance, licensing and upgrade cost

The last push to the repository was on 2026-08-19, and the most recent release is v2.0.1 from 2026-08-04, following v2.0.0 on 2026-07-09 and v1.8.0 on 2025-10-17. The jump from v1.8.0 to v2.0.0 in roughly nine months is a major-version boundary, which normally signals breaking API changes; anyone pinning to v1.8.0 should expect migration work rather than a drop-in upgrade. The version is derived from git tags through hatch-vcs, so the installed version reflects the tag it was built from.

The licence field in the repository metadata is NOASSERTION, and the README only says to see LICENSE.txt. That means the actual terms are not summarised anywhere in the repository's own metadata, and the package metadata does not resolve them to a recognised identifier. The README also asks that anyone interested in commercial applications email [email protected], which suggests the licence is not a straightforward permissive one. Read LICENSE.txt before you build anything on top of this, and treat the commercial-use question as open until you have an answer in writing.

Editorial conclusion

Adopt Aurora if you are reproducing the Nature paper, fine-tuning a foundation model on a specialised atmospheric task, or benchmarking against ERA5 in a research setting. Do not adopt it for operational forecasting, safety-critical planning or automated decision pipelines: the README states the code has not been developed or tested for non-academic purposes and that outputs are not meant to be used directly to plan operations. Before committing, verify three things: the exact terms in LICENSE.txt, which pretrained checkpoints are available and how large each download is, and whether your task matches one of the four specialised versions (medium-resolution weather, high-resolution weather, air pollution, ocean waves) rather than something adjacent.

Frequently asked questions

How do I install microsoft-aurora?

The README gives two routes: pip install microsoft-aurora, or mamba install microsoft-aurora -c conda-forge. The package requires Python 3.10 or newer according to pyproject.toml.

How much does the first run of the small pretrained model download?

The README states that running the AuroraSmallPretrained example on random data will incur a 500 MB download. The documentation is the place to check for the sizes of the other model variants.

Can I use microsoft-aurora for commercial applications?

The README asks that you email [email protected] if you are interested in using Aurora for commercial applications, and it points to LICENSE.txt for the licence itself. The repository metadata does not resolve the licence to a recognised identifier.

What variables does a Batch need to contain?

The README's example builds a Batch with surface variables 2t, 10u, 10v and msl, static variables lsm, z and slt, and atmospheric variables z, u, v, t and q. Metadata supplies latitude, longitude, the timestamp and the atmospheric levels.

Official sources

  1. Issues
  2. microsoft/aurora on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-aurora.svg)](https://hysenlabs.com/projects/microsoft-aurora)