Open-source project
aimhubio/aim avatar
aimhubio/aim

Aim: A Self-Hosted ML Experiment Tracker Built for Scale

Aim 💫 — An easy-to-use & supercharged open-source experiment tracker.

6,272 stars413 forksPythonApache-2.0

At a glance

What is it?
Aim is an open-source, self-hosted experiment tracker for Python ML teams, designed to handle tens of thousands of training runs. It stores metadata in a RocksDB-backed store, exposes a web UI for comparing runs, and provides a Python SDK for querying results programmatically.
Who is it for?
Aim is the right choice for ML teams who need self-hosted experiment tracking at scale and want to query run metadata programmatically after training. It is the wrong choice for teams who need a fully managed cloud tracking service with zero infrastructure overhead, or for projects where training runs are so few that a simple CSV log is enough.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Aim Solves for ML Teams

Aim is a self-hosted, open-source experiment tracking system for machine learning teams. The project description states it is designed to handle tens of thousands of training runs. That scale target separates it from simple logging approaches: the architecture is built around fast writes and fast queries across large experiment collections rather than around ease of embedding in a single notebook.

The primary audience is Python ML engineers who run multiple experiments per day and need to compare hyperparameters, loss curves, and system resource usage across a history of runs. The README describes the tool as covering logging metadata across an ML pipeline, visualising and comparing metadata through a UI, running trainings effectively with real-time alerting, and organising experiments with tags and grouping.

The self-hosted nature is a defining constraint. Aim stores data on the machine or server where you run it. There is no managed cloud service embedded in the open-source version, though AimStack, the company behind the project, offers enterprise support via [email protected]. Teams with strict data-residency requirements benefit from the self-hosted model. Teams that want zero infrastructure management do not.

RocksDB Storage and the Programmatic SDK

Aim stores experiment data using aimrocks, its Python binding for RocksDB. The pyproject.toml build configuration lists aimrocks 0.5.* as a required build dependency alongside Cython 3.0.10. RocksDB is a log-structured merge-tree key-value store developed for write-heavy workloads. Using it as the backend means Aim can ingest metrics from a running training loop without the overhead of a relational database write.

The v3.29.1 release notes describe an improvement in query performance achieved by reading from a single unified database and applying constant data indexing. This suggests earlier versions split data across multiple store segments, and this release consolidates those into one indexed structure. The implication for users is that queries across large run histories should be faster on v3.29.1 than on earlier releases.

Beyond the UI, the SDK enables programmatic access to all tracked metadata. The README describes querying using Python expressions, which means you can write code that filters runs by hyperparameter values or metric thresholds and processes the results without opening the browser. This is the path for automation pipelines that need to select the best checkpoint after a sweep or export results to a downstream reporting tool.

Installing Aim and Starting the Tracking Server

The package name is `aim` and it is published on PyPI. The setup.py in the repository confirms this name and requires Python 3.7 or newer.

bash
pip install aim

The setup.py lists the runtime dependencies that are installed alongside the package: aim-ui at the same version as the core, aimrecords 0.0.7, aimrocks 0.5.*, cachetools, click, cryptography, filelock, numpy, psutil, RestrictedPython, tqdm, and aiofiles. These dependencies cover the storage layer, the web UI assets, the CLI, and the runtime security sandbox for query expressions.

The repository includes a docker/ directory, which indicates a containerised deployment path exists, though the docker/ directory contents are not detailed in the README portion available. For development and research environments without Docker, the pip install path with a local Python environment is the documented approach.

The v3.27.0 release notes mention enhancements for S3ArtifactsStorage, which points to an integration for storing large artefacts such as model checkpoints in S3-compatible object storage rather than on the local filesystem where the tracking server runs. The README does not document the configuration syntax for this in the available text, but the release notes confirm the capability exists.

Framework Integrations and Distributed Training Support

The examples/ directory in the repository lists integration files for a wide range of ML frameworks. The available list includes pytorch_track.py, keras_track.py, hugging_face_track.py, catboost_track.py, fastai_track.py, lightgbm_track.py, mxnet_track.py, optuna_track.py, paddle_track.py, prophet_track.py, pytorch_ignite_track.py, pytorch_lightning_track.py, sb3_track.py, tensorboard_aim_sync.py, tensorflow_keras_track.py, and xgboost_track.py. There is also a pytorch_track_images.py, which suggests image logging is supported alongside scalar metrics.

Each file is named for its framework, which means the integration pattern for each follows the same naming convention. The tensorboard_aim_sync.py file is notable: it suggests you can sync existing TensorBoard logs into Aim without rewriting training code, which matters for teams migrating from TensorBoard-based setups.

The v3.28.0 release notes describe a new callback for HuggingFace distributed runs, addressing a gap in multi-GPU or multi-node HuggingFace Trainer setups. The v3.27.0 notes describe enhancements for the PyTorch Lightning logger and fixes for metric aggregations. These two entries show the integration layer receives active maintenance alongside the core storage and query components.

Real-time alerting is listed as a feature in the README under the heading of running trainings effectively. Logging and configurable notifications are also listed, though the README does not document which notification channels are supported in the available text.

Where Aim Is the Wrong Choice

Aim is not a managed tracking service. There is no hosted version embedded in the open-source distribution. Running it means running a process and keeping storage available. Teams that cannot operate a persistent server, or that need tracking available without any infrastructure work, will find this a friction point.

The Python minimum version is 3.7, but the build requirements include Cython 3.0.10 and aimrocks, which requires a compiled C extension. On platforms where pre-built wheels are not available, installation requires a working C compiler and RocksDB build dependencies. The setup.py note states that users are expected to install aim from wheels rather than building from source; if your platform has no published wheel, this creates an obstacle.

Aim is also not oriented toward data engineering pipelines. The README does not describe DAG-based workflow orchestration or data versioning features. It tracks training run metadata and metrics. Teams looking for a tool that also versions datasets, orchestrates data preprocessing steps, or manages model deployment should look beyond Aim for those capabilities.

The README does not document a built-in access control model for multi-user deployments. For teams where multiple researchers share a single Aim server and need to isolate their experiment data or limit write access, the README does not describe how to configure that.

Aim vs. MLflow: Different Storage and Query Approaches

MLflow is a widely used open-source experiment tracking platform that stores run data in a SQL database or on a local filesystem, depending on configuration. The primary architectural difference from Aim is the storage backend. MLflow's default store is SQLite or PostgreSQL; Aim uses RocksDB via aimrocks. For write-heavy workloads with many metrics per step and thousands of steps per run, the log-structured RocksDB approach handles concurrent writes differently from a relational store.

MLflow's query interface is primarily through its Python API and its web UI, with filtering by run parameter or tag. Aim's README describes querying using Python expressions, which the documentation positions as a more flexible query language for complex filtering. The tradeoff is that Aim's query model requires familiarity with Python expression syntax rather than a simple key-value filter.

MLflow has a plugin ecosystem for logging models in a standard format and for deploying them through MLflow-compatible serving infrastructure. The Aim README does not describe a model deployment capability in the available text. If model packaging and deployment are as important as experiment tracking in your workflow, that gap is worth investigating before choosing Aim over MLflow.

Licence, Maintenance, and Enterprise Support

Aim is licensed under Apache-2.0. This licence is permissive for commercial use and does not require publishing modifications, though it does include a patent grant clause that MLflow and some other Apache-licensed tools also carry. For teams with legal review processes, the Apache-2.0 patent terms are worth confirming with counsel before use in commercial products.

The last push to the repository was on 2026-09-26, which is two days before this writing. The project has a recent release history: v3.29.1 in May 2025, v3.28.0 in March 2025, and v3.27.0 in December 2024. This pace shows consistent maintenance at the level of several releases per year.

AimStack offers enterprise support beyond the open-source core via [email protected]. The README does not describe what enterprise support covers in the available text, so teams with specific SLA or on-site support requirements should contact AimStack directly to clarify scope.

The repository includes a CHANGELOG.md, a CITATION.cff for academic citation, a CONTRIBUTING.md, and a CODE_OF_CONDUCT.md, which together indicate a project operating with standard open-source community practices.

Editorial conclusion

Aim is the right choice for ML teams who need self-hosted experiment tracking at scale and want to query run metadata programmatically after training. It is the wrong choice for teams who need a fully managed cloud tracking service with zero infrastructure overhead, or for projects where training runs are so few that a simple CSV log is enough. Before adopting it, verify that your Python environment meets the 3.7 minimum and that your deployment host can run the web server component, since Aim requires a running process to serve the UI and handle incoming metric writes.

Frequently asked questions

How does Aim store experiment metadata compared to file-based logging approaches?

Aim stores metadata in a RocksDB-backed store via the aimrocks package, rather than writing CSV or JSON files to a directory. The v3.29.1 release notes describe a query performance improvement achieved by reading from a single unified database with constant data indexing.

Does Aim integrate with HuggingFace Trainer and PyTorch Lightning?

Yes. The examples/ directory includes hugging_face_track.py and pytorch_lightning_track.py. The v3.28.0 release notes describe a new callback specifically for HuggingFace distributed runs, and v3.27.0 describes enhancements for the PyTorch Lightning logger.

Can Aim sync existing TensorBoard logs without rewriting training code?

The examples/ directory includes a tensorboard_aim_sync.py file, which the naming suggests provides a path for syncing TensorBoard-format logs into Aim. The README does not detail the sync process in the available text, so the file itself is the starting point for understanding how the sync works.

Official sources

  1. aimhubio/aim on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/aimhubio-aim.svg)](https://hysenlabs.com/projects/aimhubio-aim)