The Well: a 15TB physics simulation dataset collection for ML surrogate models
A 15TB Collection of Physics Simulation Datasets
At a glance
- What is it?
- The Well packages 16 numerical simulation datasets into a single PyTorch-facing loader. It is built for researchers training PDE surrogates, not for small projects or interactive analysis.
- Who is it for?
- Adopt The Well if you are training or evaluating neural surrogates for spatiotemporal PDEs and can give it the disk space and GPU time; the benchmark configs and the published baseline checkpoints make it a fair comparison target. Do not adopt it if you need a small sample dataset, if your storage budget is measured in tens of gigabytes, or if you expect the bundled models to be state of the art, since the README states they are a simple baseline.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 69 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem The Well solves: surrogate models need a shared physics benchmark
Machine learning for spatiotemporal physical systems has a comparison problem. A paper that trains a Fourier Neural Operator on one fluid dataset and a paper that trains a transformer on another cannot be compared, because the data, the splits and the preprocessing all differ. The Well addresses that by collecting numerical simulations from domain scientists and numerical software developers into one package with a consistent interface. The README describes 15TB of data across 16 datasets, spanning biological systems, fluid dynamics, acoustic scattering, magneto-hydrodynamic simulations of extra-galactic fluids and supernova explosions. The intended user is a researcher who wants to train or evaluate a PDE surrogate model and needs either a single dataset or the full benchmark suite. The repository is not a general data-processing library. If your data is experimental measurements rather than simulation output, or if your sequence lengths and grid geometries fall outside what the packaged datasets provide, the collection gives you nothing to load.
How the WellDataset loader and registry work
The package is a Python library, not a service. Its central object is WellDataset, which takes three arguments in the README example: well_base_path, well_dataset_name and well_split_name. The base path points either at a local directory or at a Hugging Face URI, and the dataset name selects one of the 16 datasets. The repository layout shows the data classes under the_well/data and a registry file at the_well/utils/registry.yaml, which pyproject.toml declares as package data, so the mapping from dataset names to their metadata ships inside the installed package rather than being fetched at runtime. The loader returns batches that plug into a standard torch.utils.data.DataLoader. The benchmark side is separate: the_well/benchmark holds a train.py script that uses hydra to instantiate the dataset, model, optimizer and other components from YAML configs under the_well/benchmark/configs, with some state-of-the-art model implementations under the_well/benchmark/models. That split matters. Installing the base package gives you data access; the benchmark script and its dependencies are an extra install.
Installing the_well and running a first training job
The README recommends a machine with enough computing resources and a fresh Python environment of version 3.10 or newer, which matches the requires-python field in pyproject.toml. Creating that environment with venv looks like this:
python -m venv path/to/env
source path/to/env/activate/binThe README gives that activation path literally, so if your virtual environment is laid out differently the command will fail and you should use the activate script your platform generated. The package itself installs from PyPI:
pip install the_wellIf you want the benchmark tooling, the README specifies an extra:
pip install the_well[benchmark]Data does not come with the package. A console script named the-well-download, declared in pyproject.toml as the entry point for the_well.utils.download, fetches it:
the-well-download --base-path path/to/base --dataset active_matter --split trainOmit --dataset and --split and the README warns that everything will be downloaded, which for the full collection is 15TB. Once data is in place, the README's usage example loads a split and wraps it in a DataLoader:
from the_well.data import WellDataset
from torch.utils.data import DataLoader
trainset = WellDataset(
well_base_path="path/to/base",
well_dataset_name="name_of_the_dataset",
well_split_name="train"
)
train_loader = DataLoader(trainset)
for batch in train_loader:
...To run the bundled benchmark instead, the README gives this command from inside the benchmark directory, using hydra config groups for the experiment, server and data:
cd the_well/benchmark
python train.py experiment=fno server=local data=active_matterHere server=local selects configs/server/local.yaml, which the README says just declares the relative path to the data. The README notes the same command can be placed in an sbatch script for Slurm.
Streaming from Hugging Face versus downloading 15TB locally
Most datasets are also hosted on Hugging Face, and the loader accepts a hub URI as the base path. The README shows well_base_path set to hf://datasets/polymathic-ai/ and warns that the line may take a couple of minutes to instantiate the datamodule. That delay is the visible cost of resolving a remote dataset, and it is not the only one. The README advises downloading locally for better performance in large training. Streaming makes sense for inspection, for a small fine-tuning run, or for checking that a dataset matches your model's expected channels and resolution before you allocate disk. It is the wrong choice for a multi-epoch training job, where every batch would otherwise travel over the network. Note also that the README says most datasets are on the hub, not all of them, so a dataset you want may only be available through the-well-download.
Where The Well is the wrong tool
The size range is the first constraint. Individual datasets run from 6.9GB to 5.1TB, and the collection totals 15TB. A laptop or a shared department fileserver is not a realistic home for this. The README tells you to ensure enough free disk space, which is the only sizing guidance it gives; it does not document a partial-download or streaming-subset mechanism beyond per-dataset and per-split selection. The second constraint is what the benchmark models are. The README is explicit that the models benchmarked in the original paper were designed as a simple baseline and should not be considered state of the art. If you need a strong pretrained surrogate out of the box, these checkpoints are a starting point for comparison, not a solution. The third is scope: the datasets are numerical simulations. Anyone whose problem is experimental data, irregular sensor placement, or a physical regime not represented among the 16 datasets will spend their time adapting the loader rather than using it. The README does not document rollback or versioning of the downloaded data itself, so reproducibility of a specific data revision rests on the package version you pin, not on the download command.
The Well compared with a general dataset hub
The obvious alternative is pulling simulation data from a general-purpose hub or from the original simulation codes' own output. The difference in approach is curation versus generality. A general hub gives you a place to publish and fetch arrays with no opinion about their structure; you write the normalization, the train/validation/test split and the channel metadata yourself, and you own the consequences when two papers make different choices. The Well instead ships a registry and a dataset class that impose a consistent interface across 16 heterogeneous physical systems, plus a hydra-driven benchmark harness so that an FNO run and a transformer run differ by a config group rather than by a fork. The cost of that consistency is that you inherit the collection's choices about splits, resolutions and variable naming. If your research question depends on a split or a preprocessing step the package does not expose, a plain hub download plus your own loader is less friction than fighting the registry.
Editorial conclusion
Adopt The Well if you are training or evaluating neural surrogates for spatiotemporal PDEs and can give it the disk space and GPU time; the benchmark configs and the published baseline checkpoints make it a fair comparison target. Do not adopt it if you need a small sample dataset, if your storage budget is measured in tens of gigabytes, or if you expect the bundled models to be state of the art, since the README states they are a simple baseline. Before committing, verify the actual on-disk size of the datasets you plan to use (they range from 6.9GB to 5.1TB each), confirm your Python is 3.10 or newer, and check whether the dataset you want is among those hosted on Hugging Face if you intend to stream rather than download.
Frequently asked questions
How do I install The Well?
The README recommends a Python 3.10 or newer environment, then pip install the_well from PyPI, or pip install the_well[benchmark] if you also want the benchmark dependencies. Installing from source means cloning the repository and running pip install . inside it.
How large is The Well and what disk space do I need?
The README states the collection totals 15TB across 16 datasets, with individual datasets ranging from 6.9GB to 5.1TB. It advises checking that your system has enough free disk space for the datasets you intend to download.
Can I stream The Well data instead of downloading it?
Most datasets are hosted on Hugging Face and the loader accepts hf://datasets/polymathic-ai/ as the base path. The README warns that instantiating the datamodule this way may take a couple of minutes and advises downloading locally for better performance in large training.
Are the benchmarked model checkpoints in The Well state of the art?
No. The README states the models benchmarked in the original paper were designed as a simple baseline and should not be considered state of the art, and expresses the hope that the community builds better architectures for PDE surrogate modeling.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/polymathicai-the-well)