Open-source project
mlr-org/mlr3 avatar
mlr-org/mlr3

mlr3: an object-oriented machine learning toolkit for R

mlr3: Machine Learning in R - next generation

1,080 stars97 forksRLGPL-3.0

At a glance

What is it?
mlr3 rebuilds the mlr package around R6 classes, with tasks, learners, resamplings and measures as separate objects. It targets R users who want explicit control over the experiment rather than a single modelling call, and it installs from CRAN in one line.
Who is it for?
mlr3 suits R users who need to assemble training, resampling and evaluation as separate, inspectable objects, and who accept that learners live in extension packages such as mlr3learners and mlr3extralearners.
Can I use it commercially?
Yes, with conditions. LGPL-3.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly R, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem mlr3 solves for R users

The README is blunt about why mlr3 exists. The predecessor, mlr, reached CRAN in 2013, and the project states that its core design dates back even further. Years of added features produced what the README calls feature creep, which it says made mlr hard to maintain and hard to extend. The rewrite is not a cosmetic refactor. It is an attempt to fix the extension points that the old package got wrong.

The README names the specific failure: mlr was extensible in some parts (learners, measures) but other parts were less easy to extend from the outside. That is the problem mlr3 addresses. If you have ever wanted to add a new resampling scheme or a new evaluation measure to an R modelling package and found the internal contract undocumented or closed, this is the design decision mlr3 was written to correct.

The audience follows from that. mlr3 is for R users who treat a machine learning experiment as a set of named parts rather than a single function call. The README lists the intended users as both users and developers, and the repository ships a CONTRIBUTING.md and an AGENTS.md alongside the R source, which signals that extension authors are a first-class audience, not an afterthought.

Tasks, learners and resamplings as separate objects

The mechanism is visible in the README example. You do not hand a data frame to a model. You build a task first, then a learner, then optionally a resampling, and you connect them explicitly.

The README constructs a classification task from the palmerpenguins data with as_task_classif(species ~ ., data = ...). The printed object reports the target, the properties, the feature types and the class distribution. That printout is the point: the task is an object you can inspect before any model exists.

Learners are named by string, for example lrn("classif.rpart", cp = .01), and hyperparameters are set at construction. Training takes a task and a row index vector, not a formula. Prediction returns a prediction object with a confusion matrix and a score method. Resampling is its own object: rsmp("cv", folds = 3L), passed to resample() together with the task and learner. The result carries per-iteration scores and an aggregate.

This is the design principle the README states directly: only the basic building blocks are implemented in this package. Learners themselves are elsewhere. The README points to mlr3learners for recommended core regression, classification and survival learners, and mlr3extralearners for everything else, with a learner search page on the mlr-org site. That split is what keeps the core small, and it is also the first thing a new user trips over, because lrn("classif.randomForest") will not work until the package that provides it is installed.

Installing mlr3 and running a first resampling

The README gives two installation routes. The release version comes from CRAN. The development version comes from GitHub through pak. For a first real use, the README recommends the mlr3verse meta-package, which installs mlr3 together with some of the most important extension packages.

r
install.packages("mlr3")

If you want the development version instead, the README shows the pak route, which requires pak to be installed first.

r
# install.packages("pak")
pak::pak("mlr-org/mlr3")

For a working setup that includes learners, the README recommends the meta-package. This is the command a beginner should run rather than the single-package install, because the core package on its own ships only the building blocks.

r
install.packages("mlr3verse")

A first run follows the README example: build a task, choose a learner, define a resampling, and aggregate a measure. The palmerpenguins data is used in the README, so it must be available.

r
library(mlr3)

task = as_task_classif(species ~ ., data = palmerpenguins::penguins)
learner = lrn("classif.rpart", cp = .01)
resampling = rsmp("cv", folds = 3L)

rr = resample(task, learner, resampling)
measure = msr("classif.acc")

rr$score(measure)[, .(task_id, learner_id, iteration, classif.acc)]
rr$aggregate(measure)

What you should see is a table with one row per fold and an aggregate accuracy. The README shows exactly this shape of output: three rows for three folds and a single aggregated value. If your learner is not installed, lrn() fails before resample() is reached, which is the extension-package split showing up at the worst moment.

Where mlr3 gets in the way

The extension-package split is the main friction. The core package deliberately implements only building blocks, so a fresh install of mlr3 alone gives you a framework with few usable learners. The README's own advice, to install mlr3verse for a better user experience, is an admission that the bare package is not a complete working environment. Anyone who skips that advice will spend their first hour resolving missing learner packages.

The object model is also heavier than a formula interface. You must decide on a task class, name learners as strings, and pass row indices for train and test rather than letting a function split internally. For a one-off model on a small data frame, that ceremony buys nothing. The README's own example shows the cost: constructing the task, splitting it, training, predicting and scoring takes several distinct calls where a single modelling function would take one.

There is a documentation boundary worth naming. The README is not the tutorial. It states that the book at mlr3book.mlr-org.com should be the central entry point, and it defers to the mlr-org FAQ, a gallery, cheatsheets and the reference manual. The README does not document migration from mlr, so if you have an existing mlr codebase, the README will not tell you how to move it. The repository also does not document rollback or version pinning in the README, so reproducibility across mlr3 releases is something you have to establish from NEWS.md and your own lockfile practice.

mlr3 compared with tidymodels

The two are the obvious comparison in R, and the difference is in the object model rather than the feature list. tidymodels is a collection of packages built around the tidymodels and parsnip modelling interfaces, with recipes for preprocessing and workflows for bundling steps. mlr3 puts tasks, learners, resamplings and measures in one package built on R6 classes, and pushes learners out to mlr3learners and mlr3extralearners.

The practical consequence is where the abstraction sits. In the tidymodels approach, preprocessing is a first-class stage through recipes, and a workflow ties a model to that recipe. In mlr3, the core README describes tasks, learners, resamplings and measures; preprocessing and pipeline composition live in extension packages such as mlr3pipelines, which the README lists among the cheatsheets. So mlr3 asks you to compose more, and install more, to reach the same place.

The extension model is the second difference. mlr3's learner catalogue is deliberately external and searchable through the mlr-org learner search page, with recommended learners separated from the long tail. That is a clear contract for contributors, and it means the core package changes less often. It also means the learner you want may live in a package with a different release cadence.

Maintenance, licence and upgrade cost

The repository is not archived and the last push was on 2026-08-31, so development is current. The release history shows v1.8.0 on 2026-08-21, v1.7.1 on 2026-06-11 and v1.7.0 on 2026-06-10. That is a steady cadence of minor and patch releases rather than long quiet gaps.

The upgrade cost is spread across the ecosystem, not concentrated in mlr3. Because learners live in mlr3learners and mlr3extralearners, a core release and an extension release can move independently, and a version bump in one does not guarantee the other has caught up. The repository ships a NEWS.md, which is the file to read before upgrading, and a DESCRIPTION that pins the dependency contract. Neither the README nor the release list documents a rollback path, so pinning versions in your own project is the only mechanism available to you.

The licence is LGPL-3.0, and the repository carries a LICENSE.md. For most users this is a non-issue: linking to the package from your own R code does not trigger the copyleft obligations that a GPL-style licence would impose on your own source. The obligation attaches to modified versions of the library itself. This is a general description of how LGPL-3.0 is usually read, not legal advice; if you are redistributing a modified mlr3 or embedding it in a product, read the licence text and get your own counsel.

Editorial conclusion

mlr3 suits R users who need to assemble training, resampling and evaluation as separate, inspectable objects, and who accept that learners live in extension packages such as mlr3learners and mlr3extralearners. It is the wrong choice if you want one function call that fits and predicts without naming a task or a resampling, or if you work outside R. Before adopting it, check that the learner you need is available in mlr3learners or mlr3extralearners, and read the mlr3book rather than the README, which points to the book as the central entry point.

Frequently asked questions

What is the purpose of mlr3?

The README describes mlr3 as efficient, object-oriented programming on the building blocks of machine learning, and as the successor to mlr. It provides tasks, learners, resamplings and measures as separate objects, with learners supplied by extension packages.

What does the acronym MLR stand for in mlr3?

The README does not expand the acronym. It only presents mlr3 as machine learning in R and as the successor of mlr, so the letters are never defined.

What is MLR healthcare, and is mlr3 related to it?

Nothing in the README connects mlr3 to healthcare. The package is a machine learning framework for R, and its listed resources are a book, a gallery, cheatsheets and courses on machine learning.

Official sources

  1. License: LGPL-3.0
  2. mlr-org/mlr3 on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mlr-org-mlr3.svg)](https://hysenlabs.com/projects/mlr-org-mlr3)