Open-source project
root-project/root avatar
root-project/root

ROOT: C++ and Python Toolkit for Scientific Data Analysis

The official repository for ROOT: analyzing, storing and visualizing big data, scientifically

3,305 stars1,573 forksC++NOASSERTION

At a glance

What is it?
ROOT is CERN's C++ framework for storing, processing and plotting scientific data at petabyte scale, with Python bindings through cppyy. It is a strong fit for particle physics and large columnar datasets, and a poor fit for small tabular workloads.
Who is it for?
Adopt ROOT if you work with columnar physics data, need histogramming and fitting in C++ or Python, or want RDataFrame to scale an analysis across cores and distributed backends. Do not adopt it for small CSV or relational workloads where pandas, DuckDB or plain NumPy already fit; the build and dependency surface is not worth it.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ROOT solves, and for whom

ROOT exists for one class of problem: scientific datasets too large to hold in memory and too structured to treat as flat text. The README states that exabytes of scientific data are written in columnar ROOT format at the Large Hadron Collider experiments, and that the package covers the path from data acquisition to publication-ready plots. That is the scope. It is not a general data science library.

The audience follows from that scope. Experimental physicists and analysts who need histograms in an arbitrary number of dimensions, curve fitting, statistical modelling and minimization in one coherent environment. The README also positions the package as performance critical software written in C++, with rapid prototyping through Cling, a C++ compliant interpreter. That combination matters: you can iterate interactively and still ship compiled C++ when the analysis stabilizes.

If your data is a few hundred megabytes of CSV and your output is a bar chart, ROOT is the wrong instrument. The columnar format, the interpreter and the C++/Python bridge all carry weight that only pays off at scale.

Cling, cppyy and the C++/Python execution model

The mechanism worth understanding is the interpreter. Cling does not merely evaluate C++ snippets; the README describes it as enabling performant C++ type introspection, which is the building block of automatic interoperability with Python. That introspection is what makes the Python bindings work without hand-written wrappers for every class.

The Python side is described in the README as dynamic bindings leveraging cppyy technology, giving efficient on-demand C++/Python interoperability in a uniform cross-language execution environment. In practice this means a Python session can instantiate C++ objects that the interpreter has just parsed, rather than a pre-generated binding layer. The trade-off is that the boundary is dynamic: type resolution happens at runtime, so errors that a compiled binding would catch at build time can surface later.

RDataFrame sits on top of this as the parallel processing framework. The README says it can considerably speed up an analysis by taking advantage of multi-core and distributed systems. The repository's requirements.txt lists pyspark as the Spark backend and dask plus distributed as the Dask backend, which tells you the distributed story is Python-driven even though the core engine is C++.

Installing ROOT and running a first histogram

The README does not embed install commands. It points to https://root.cern/install for installation instructions and to https://root.cern/install/build_from_source for building from source. Those two pages are the authoritative source, and platform specifics live there rather than in the repository root.

What the repository does show is a Python packaging path. The pyproject.toml declares the project name as root, requires Python 3.10 or newer, and builds through scikit-build-core with Ninja as the CMake generator. The wheel version is derived from the version macros in core/foundation/inc/ROOT/RVersion.hxx, with a pre-release suffix appended, so a wheel for ROOT 6.42.02 is published as version 6.42.2a1.

The pyproject.toml lists numpy as the only hard runtime dependency, so the core install is comparatively light. Everything else in requirements.txt is optional and grouped by feature: pandas for PyROOT array interoperability, onnx and onnxscript for TMVA SOFIE, scikit-learn, tensorflow, torch and xgboost for PyMVA, jax for RDataLoader, numba and cffi for the ROOT.Numba.Declare decorator, and IPython, jupyter and metakernel for the C++ notebook kernel.

For a first real use, the README points to the Getting started page at https://root.cern/learn and to the tutorial index at https://root.cern/doc/master/group__Tutorials.html. Those tutorials are the intended entry point, and the repository keeps a tutorials/ directory alongside them. The install page at https://root.cern/install is where the executable ends up on your PATH once the platform-specific steps are done.

Where ROOT gets in the way

The first limitation is the one the packaging already admits. The wheel requires Python 3.10 or newer. If your analysis environment is pinned to an older interpreter, the PyPI route is closed and you are back to the install page and a source build.

The second is dependency sprawl on the feature side. requirements.txt is a long list, and several entries are heavy: tensorflow, torch, jax, pyspark and dask are all there. None are required by the core wheel, but the moment you want a TMVA interface or a distributed RDataFrame backend, you inherit that weight. The file also carries a version marker, tensorflow is restricted to python_version < "3.14", which is a signal that the optional stack tracks Python releases unevenly.

The third is conceptual. ROOT's own README frames it around scientific data at exabyte scale. That framing is honest, and it also means the abstractions are tuned for that case. If your workload is a join across two modest tables, the columnar format and the interpreter add ceremony without adding throughput. A dataframe library that fits in memory will be simpler to reason about.

Finally, the licence metadata is worth a direct look. The repository's top-level entries include both LGPL2_1.txt and LICENSE, and the README badge reads LGPL v2.1+. GitHub reports the licence as NOASSERTION, so the automated classifier did not resolve it. Read the LICENSE file yourself rather than trusting the badge.

ROOT versus a general-purpose dataframe stack

The obvious alternative for Python users is pandas, or a dataframe engine like DuckDB, and the difference is not speed, it is the storage and execution model. pandas holds a table in memory and operates on it in the host process. ROOT writes columnar files designed to be read in chunks, and RDataFrame executes over those chunks with an explicit parallel backend.

That distinction drives everything else. With pandas you load, then transform. With RDataFrame the README describes a framework that takes advantage of multi-core and distributed systems, and the repository's requirements.txt shows the backends it targets: pyspark for Spark, dask and distributed for Dask. You choose where the work runs.

The second difference is the fitting and statistics layer. ROOT ships histogramming in an arbitrary number of dimensions, curve fitting, statistical modelling and minimization as part of the same package, and roofit/ and tmva/ are top-level directories in the repository. A pandas plus SciPy stack can do pieces of this, but you assemble it. ROOT ships it assembled, at the cost of learning ROOT's own APIs.

The third is language. ROOT is C++ first, with Python as a dynamically bound layer. If your team is Python-only and unwilling to touch C++ semantics, the cppyy boundary will occasionally surprise you. If your team already writes C++ for detector or reconstruction code, the same environment serves both.

Maintenance, releases and what upgrading costs

The repository is not archived, and the last push was on 2026-09-10. Recent releases show a maintenance pattern rather than a single line: v6-40-04 on 2026-08-27, v6-36-14 on 2026-09-01, and v6-32-24 on 2026-06-18. Three release lines receiving updates means older series get patches, which is good news if you cannot move to the newest minor version immediately.

That also defines the upgrade cost. You are not forced onto the latest release to stay supported, but you do have to pick a line and track it. The wheel versioning adds a wrinkle: pyproject.toml derives the wheel version from RVersion.hxx and appends a pre-release suffix, with a comment stating the suffix is bumped manually only to re-publish a wheel for a ROOT version already on PyPI, and reset to a1 for the next release. A version like 6.42.2a1 is therefore an alpha marker in the packaging sense, not a statement about the C++ release's stability.

On licensing, the README badge points to LGPL v2.1+, and LGPL2_1.txt is present at the repository root. LGPL is a copyleft licence with a linking exception in its usual form, which matters if you redistribute a product that links against ROOT. This is not legal advice; if you ship ROOT inside a commercial product, have counsel read LGPL2_1.txt and the LICENSE file together, since the two files coexist at the root and GitHub's classifier did not resolve them.

Editorial conclusion

Adopt ROOT if you work with columnar physics data, need histogramming and fitting in C++ or Python, or want RDataFrame to scale an analysis across cores and distributed backends. Do not adopt it for small CSV or relational workloads where pandas, DuckDB or plain NumPy already fit; the build and dependency surface is not worth it. Before committing, verify the install path for your platform at root.cern/install, check whether your Python version is 3.10 or newer for the wheel, and confirm which optional dependencies from requirements.txt your analysis actually needs, since tensorflow, torch, jax, pyspark and dask are all listed there.

Frequently asked questions

What is ROOT software?

ROOT is a unified software package for the storage, processing and analysis of scientific data, from acquisition to final visualization in the form of customizable, publication-ready plots. It is written in C++ and provides Python interoperability through dynamic bindings.

What does ROOT mean?

The README does not expand the name into a phrase. It describes ROOT only as a unified software package for scientific data storage, processing and analysis, and the citation credits it as ROOT - An Object Oriented Data Analysis Framework.

How do I download and install ROOT?

The README directs readers to https://root.cern/install for installation instructions, and to https://root.cern/install/build_from_source for building from source. A Python wheel is also published under the name root, requiring Python 3.10 or newer.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. root-project/root on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/root-project-root.svg)](https://hysenlabs.com/projects/root-project-root)