facebookresearch/balance: survey weights for biased samples in Python
The balance python package offers a simple workflow and methods for dealing with biased data samples when looking to infer from them to some target population of interest.
At a glance
- What is it?
- balance fits a weight to each respondent so a biased sample can stand in for a target population. It is a small, opinionated Python API on top of pandas, aimed at survey and observational data, and it is still labelled beta.
- Who is it for?
- Adopt balance if you have a sample and a target frame that share covariates and you need weights you can inspect, plot and export, and if you are comfortable that the missing at random assumption holds for your data. Do not adopt it for causal identification: the package adjusts for bias under MAR, and if the bias mechanism is not captured by the covariates you supply, weighting will not recover the population.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What balance solves, and who ends up using it
A survey sample is rarely a miniature of the population it is meant to describe. Some groups answer more often, some are harder to reach, and the resulting estimates are skewed in ways that no amount of modelling downstream will fix. balance addresses that specific gap: given a sample with covariates and a separate sample from the target population with the same covariates, it fits a weight per unit in the sample. The README describes the interpretation directly: each weight can be loosely read as the number of people from the target population that this respondent represents.
The package targets researchers rather than production engineers. The README names survey methodologists, demographers, UX researchers and market researchers, plus data scientists, statisticians and machine learners. That audience matters for judging the API. It is built around pandas DataFrames and notebook-style exploration, with plotting helpers wired in, rather than around a serving path. If you are looking for a component to drop into an inference pipeline, the workflow will feel heavier than you want.
The Sample object and the adjust() call
The mechanism is a small object model. You wrap your sample and your target in Sample objects, attach one to the other, and then call adjust(). The README's quickstart shows exactly this shape, and the resulting adjusted object carries the weights.
What the package does underneath is inverse propensity weighting, referred to in the README as ipw. The idea is that the sample and the target are two groups, and a model is fitted to distinguish them from the covariates. Units that look more like the target get more weight; units that look over-represented relative to the target get less. The README frames the whole exercise under the missing at random assumption, which is the load-bearing condition: the bias has to be explainable by covariates you have measured in both the sample and the population. That is a strong claim about your data, and the package cannot verify it for you.
The workflow the README lays out has seven steps: load the respondents, load the target population, run diagnostics on the covariates, adjust, evaluate, use the weights for population-level estimates, and save the weights. The order is deliberate. Diagnostics come before adjustment, and evaluation comes after, which is the right sequence for anyone who has been burned by fitting weights and reporting the result without checking them.
Installing balance and running the quickstart
balance requires Python 3.9 through 3.14, and the README states it builds and runs on Linux, OSX and Windows. The recommended install is from PyPI, which pulls prebuilt wheels for those three platforms.
python -m pip install balanceIf you want the current main branch instead of the released version, the README gives a git install. Note that this bypasses the wheel and builds from source, so you will need a working build environment.
python -m pip install git+https://github.com/facebookresearch/balance.gitFrom a local clone, the README offers the plain install and a dev variant that adds tooling such as pytest, sphinx and pre-commit.
cd balance
python -m pip install .[dev]Once installed, the quickstart loads simulated data and builds the two objects. The sample carries an outcome column, here named happiness, while the target does not. That asymmetry is the whole point: outcomes exist only in the sample, covariates exist in both.
from balance import load_data, Sample
target_df, sample_df = load_data()
sample = Sample.from_frame(sample_df, outcome_columns=["happiness"])
target = Sample.from_frame(target_df)
sample_with_target = sample.set_target(target)After that, the README suggests checking diagnostics before adjusting, with the plot call commented out in the snippet. Then comes the adjustment itself, a single call that defaults to ipw.
adjusted = sample_with_target.adjust()What you get back is an adjusted object with weights attached. The README points to separate documentation pages for pre-adjustment diagnostics and for the adjustment process, and to a quickstart tutorial for the full walkthrough. Those pages, not the README, are where the options for adjust() live.
The missing at random assumption is the real limit
Every weighting method lives or dies by its assumptions, and balance states its own plainly. The README says bias can sometimes be mitigated under MAR by relying on covariates present for all items in the sample and in a sample from the population. The word sometimes is doing a lot of work there.
If non-response depends on something you did not measure in both frames, the weights will not correct it, and they will not warn you either. You will get a fitted weight per respondent and a set of diagnostics that look reasonable, because the diagnostics only see the covariates you gave the model. This is the failure mode to plan for: a clean-looking adjustment on a covariate set that misses the mechanism driving the bias.
There is a second constraint in the setup itself. You need a target sample with the same covariates as your survey sample. In many real projects that frame does not exist, or exists at a different granularity, or has different variable definitions. balance gives you no way around that; without the target, there is nothing to adjust toward. If your only information about the population is a set of known marginal totals, this is not the tool for the job.
How balance differs from generic propensity score tooling
The closest alternative in practice is not another balancing package but the general propensity score machinery in scikit-learn, which is already a dependency here. With that route you fit a classifier to distinguish sample from target, take the predicted probabilities, convert them into weights yourself, and then handle trimming, diagnostics and evaluation on your own. You get full control and you own every decision.
balance sits one level up. It fixes the workflow, provides the Sample abstraction so the sample and target travel together, ships diagnostics and plots, and exposes the adjustment as a single call. The trade-off is the usual one for opinionated wrappers: less flexibility in exchange for not rebuilding the same pipeline in every project. If your weighting scheme is unusual, or you need a custom estimator, the wrapper becomes an obstacle and the scikit-learn path is the shorter route. If your need is the standard survey reweighting case the README describes, the wrapper saves real work.
The package is also explicitly in beta, per the README note, and is described there as actively supported. That is a statement about support, not about API stability. Expect the surface to move.
Versioning, dependencies and the MIT licence
The project is MIT licensed, with the repository carrying a separate LICENSE-DOCUMENTATION file alongside LICENSE, which suggests the docs are covered under different terms than the code. If you are redistributing the documentation or the site content, read that second file rather than assuming the MIT terms extend to it. This is a note about what the repository contains, not legal advice.
Dependency pinning is the practical cost to budget for. The pyproject.toml splits requirements by Python version: numpy is capped below 2.0 on Python earlier than 3.12, and scipy is capped below 1.14.0 on the same range, while 3.12 and later get looser floors. If you are on an older Python and your environment already runs a newer numpy or scipy, installing balance can force a downgrade. That is the kind of conflict that surfaces at install time, not at import time, so it is worth resolving in a clean environment first.
Release cadence is visible in the changelog: 0.21.0 on 2026-06-02, 0.22.0 on 2026-07-15, and 0.23.0 on 2026-08-04, which is also the date of the last push to main. The version numbers are still in the 0.x range, consistent with the beta label. Budget for reading the changelog before each upgrade rather than assuming patch-level compatibility.
Editorial conclusion
Adopt balance if you have a sample and a target frame that share covariates and you need weights you can inspect, plot and export, and if you are comfortable that the missing at random assumption holds for your data. Do not adopt it for causal identification: the package adjusts for bias under MAR, and if the bias mechanism is not captured by the covariates you supply, weighting will not recover the population. Before you commit, install it in the Python version you actually run, reproduce the quickstart on your own two frames, and look at the pre-adjustment diagnostics before trusting any adjusted estimate.
Frequently asked questions
How do I install balance?
The README recommends installing the latest stable version from PyPI with python -m pip install balance, which uses prebuilt wheels for OSX, Linux and Windows. You can also install the bleeding edge version from Git with python -m pip install git+https://github.com/facebookresearch/balance.git.
Which Python versions does balance support?
The README states you need Python 3.9, 3.10, 3.11, 3.12, 3.13 or 3.14, and that balance can be built and run on Linux, OSX and Windows.
What does the weight that balance fits actually mean?
The README says balance fits a weight for each unit in the sample that can be loosely interpreted as the number of people from the target population that this respondent represents.
What assumption does balance rely on to reduce bias?
The README places the method under the missing at random assumption, where bias can sometimes be partially mitigated using covariates that are present both for the sample and for a sample of items from the population.
Is balance still in beta?
Yes. A note in the README states that balance is currently in beta and is actively supported, and the pyproject.toml classifier is Development Status :: 4 - Beta.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-balance)