libact: pool-based active learning in Python, and when a bandit picks your strategy
Pool-based active learning in Python
At a glance
- What is it?
- libact implements thirteen query strategies behind a shared interface and adds ActiveLearningByLearning, a meta-algorithm that chooses among them while labelling runs. It suits small labelled budgets and scikit-learn pipelines, not deep learning or Windows.
- Who is it for?
- Adopt libact when your labelling budget is the bottleneck and your model is a scikit-learn classifier; the shared interface and the ActiveLearningByLearning meta-algorithm are the parts worth having. Do not adopt it if you need deep learning query strategies, PyTorch tensors, or a native Windows wheel, because the cibuildwheel configuration skips Windows entirely and the README points Windows users at WSL.
- Can I use it commercially?
- Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 40 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The labelling budget problem libact was written for
Active learning starts from the assumption that unlabelled data is cheap and labels are expensive. You have a pool of candidate examples, an oracle that can label one at a time, and a fixed budget of queries. The question is which example to ask about next. libact exists to answer that question with code rather than by reimplementing the same loop for every paper you want to try.
The README frames the package as making active learning easier for real-world users, and the repository layout backs that up: libact/ holds the library, examples/ holds runnable scripts such as label_digits.py and plot.py, and docs/ feeds the Read theDocs site. The audience is someone who already has a scikit-learn classification pipeline and wants to swap the training-set selection step, not someone building a labelling product from scratch. There is no annotation UI, no database, and no queueing here. It is a library that returns an index into your pool.
How the pool, the strategy and the model fit together
The design is pool-based, which the README states in the title. You hold a dataset where every example already has features, some subset carries labels, and the rest sit in an unlabelled pool. A query strategy scores the pool and returns the index it wants labelled. You label that example, add it to the labelled set, retrain, and repeat until the budget runs out.
What makes the package more than a bag of scoring functions is the unified interface. The README says libact provides one interface for implementing strategies, models and application-specific labelers, so a strategy written against that interface can be swapped for another without rewriting the loop. The strategy table groups the thirteen implementations by intent: UncertaintySampling and EpsilonUncertaintySampling exploit the current model, CoreSet and InformationDensity go after diversity and representativeness, QueryByCommittee and BALD use disagreement across an ensemble, QUIRE and DWUS mix uncertainty with density, and RandomSampling is the baseline you should always compare against.
ActiveLearningByLearning is the unusual one. The README describes it as a meta-algorithm that assists users in selecting the best strategy on the fly, and the strategy table calls it a multi-armed bandit over the other strategies. Instead of committing to uncertainty sampling for the whole run, it treats each candidate strategy as an arm and shifts weight toward whichever is performing. That is a real design decision with a real cost: you are now running several strategies per round instead of one, and the bandit itself has behaviour you have to trust. The linked technical report on arXiv is where the selection rule is argued; the README does not reproduce it.
Installing libact and labelling your first pool
The README gives the official release as a plain pip install. Python 3.9 through 3.12 are listed as supported, and numpy>=2, scipy>=1.13, scikit-learn>=1.6, matplotlib>=3.8 and joblib come along automatically.
pip install libactOn Debian or Ubuntu the README asks for BLAS and LAPACKE headers before building, because the variance reduction and hintsvm modules compile C extensions. Arch users install lapacke; macOS users install openblas with Homebrew.
sudo apt-get install build-essential gfortran libatlas-base-dev liblapacke-dev python3-devIf you would rather not compile those extensions, the README documents build options that disable them. This is the configuration to reach for when a build fails on a machine you do not control.
pip install libact --config-settings=setup-args="-Dvariance_reduction=false" \
--config-settings=setup-args="-Dhintsvm=false"For development the README recommends the conda environment file at the repository root, then an editable install with build isolation turned off so that meson, ninja and cython stay available for rebuilds.
conda env create -f environment.yml
conda activate libact
pip install --no-build-isolation -e .The README warns that importing libact after an editable install with build isolation can fail on missing ninja, and the fix it gives is either reinstalling without isolation or falling back to a regular install. Once the import works, examples/label_digits.py is the shortest path to a working loop, and examples/plot.py is the one to read if you want to see learning curves rather than a single run.
Where libact stops being the right tool
The strategy list is classical. Uncertainty sampling, committee disagreement, k-center greedy, density weighting and the two SVM-flavoured modules all assume a scikit-learn estimator with predict_proba or a decision function. There is no query strategy for a neural network's penultimate layer, no embedding-based diversity, and no batch acquisition beyond what a strategy returns one index at a time. If your model is a transformer, you will be writing the strategy yourself against libact's interface, and at that point you have to ask whether the interface is worth more than the framework you already use.
Windows is the other boundary. The README recommends WSL as the primary environment for installing and running libact on Windows, and the pyproject.toml cibuildwheel section skips "*-win*" along with musllinux and PyPy. So there is no official Windows wheel and no promise of one. A native Windows user is expected to build from source or move to WSL, and the README does not document a supported native path.
There is also a version mismatch worth noticing. The pyproject.toml declares version 0.2.0, while the most recent release listed is v0.1.5 from 2019. The classifiers block still lists Python 3.8 even though requires-python is >=3.9. Those are small inconsistencies, but they mean you should read pyproject.toml rather than the release page when you need to know what you are actually getting.
modAL and the difference in approach
The closest comparison in this space is modAL, another Python active learning library built around scikit-learn. The difference is where the abstraction sits. modAL wraps a model in an ActiveLearner object that owns the estimator, the query strategy and the labelled data, so the loop is a method call on the learner. libact keeps the dataset and the strategy separate, and the strategy's job is to return an index into the pool. That separation is what lets ActiveLearningByLearning treat strategies as interchangeable arms; a design where the learner owns everything makes a meta-algorithm over strategies harder to express.
The trade-off runs the other way too. A learner object that owns the labelled set is easier to serialize and easier to hand to a colleague than a loop where you manage the labelled indices yourself. libact also carries compiled extensions for VarianceReduction and HintSVM, which means BLAS and LAPACKE headers on Linux and a build step that modAL does not require. If you want the smallest possible dependency surface, that is a point against libact. If you specifically want the bandit meta-algorithm or the QUIRE and DWUS density mixtures, libact is where those live.
Maintenance, licensing and the cost of upgrading
The repository is not archived and the last push was on 2026-08-21, so the codebase is being touched. That is not the same as a release cadence. The three releases listed are v0.1.3 in 2017, v0.1.4 in 2019 and v0.1.5 in 2019, while pyproject.toml already declares 0.2.0. The practical reading is that the project moves on master and that PyPI lags behind it. If you need a feature that only exists on master, the README documents the development install directly from the git URL, and you should expect to pin a commit rather than a version.
The upgrade cost is concentrated in the build. The dependency floors are aggressive: numpy>=2, scipy>=1.13, scikit-learn>=1.6, matplotlib>=3.8, and joblib pinned to exactly 1.5.1. That exact joblib pin is the kind of thing that collides with another package's requirement in a shared environment, and the README does not discuss it. The build backend moved to meson and meson-python, so anyone with an old setup.py-based workflow has to relearn the editable install, which is why the README spends a troubleshooting paragraph on ninja errors.
Licensing is BSD-2-Clause, stated in the repository and referenced from pyproject.toml as a file. That is a permissive licence, and the practical implication for a commercial user is that redistribution and modification are allowed provided the copyright notice and licence text travel with the source. The compiled modules for VarianceReduction and HintSVM are part of the same distribution, so disabling them with the build options is a size and dependency decision, not a licensing one. This is a description of what the repository states, not legal advice; check the LICENSE file against your own obligations.
Editorial conclusion
Adopt libact when your labelling budget is the bottleneck and your model is a scikit-learn classifier; the shared interface and the ActiveLearningByLearning meta-algorithm are the parts worth having. Do not adopt it if you need deep learning query strategies, PyTorch tensors, or a native Windows wheel, because the cibuildwheel configuration skips Windows entirely and the README points Windows users at WSL. Before committing, verify that your scikit-learn version satisfies the >=1.6 floor and that the model you plan to wrap is one of the available models or can be expressed through SklearnAdapter, since those two checks decide whether the library fits your stack at all.
Frequently asked questions
Does libact run on Windows?
The README recommends Windows Subsystem for Linux as the primary environment for installing and running libact on Windows, and the cibuildwheel configuration in pyproject.toml skips Windows wheels. There is no documented native Windows path.
What Python versions does libact support?
The README lists Python 3.9, 3.10, 3.11 and 3.12, and pyproject.toml sets requires-python to >=3.9. The classifiers block still mentions 3.8, which conflicts with that floor.
What does ActiveLearningByLearning do in libact?
The README describes it as a meta-algorithm that selects the best strategy on the fly, and the strategy table lists it as a multi-armed bandit over the other strategies. The selection rule itself is in the linked arXiv technical report, not the README.
How do I install libact without the compiled extensions?
The README documents build options that turn off the optional modules. Passing -Dvariance_reduction=false and -Dhintsvm=false through --config-settings=setup-args skips both C extensions.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ntucllab-libact)