# Splink: probabilistic record linkage in Python with DuckDB and Spark backends

> Splink is a Python package for deduplicating and linking records that lack unique identifiers. It fits Fellegi-Sunter models without training data and runs the same code against DuckDB, Spark, PostgreSQL or SQLite.

**moj-analytical-services/splink** — Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends

- Repository: https://github.com/moj-analytical-services/splink
- Website: https://moj-analytical-services.github.io/splink/
- Stars: 2,461 · Forks: 269
- Language: Python
- License: MIT
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/moj-analytical-services-splink

## The record linkage problem Splink targets

Two tables describe the same people and neither carries a shared key. Names are spelled differently, dates of birth are mistyped, cities are abbreviated. Deterministic joins on exact equality throw away most true matches; joins on fuzzy string similarity alone produce too many false ones. Splink's answer is a probabilistic model over record pairs: each pair gets a match probability, and pairs above a threshold are treated as links and then grouped into clusters that stand in for a unique entity ID.

The audience is analysts who work in Python and can write SQL-shaped logic but do not want to hand-tune a matching pipeline. The README points to government, academia and the private sector, and the Office for National Statistics has published a case study about using Splink to link 2021 Census data to itself. The package is MIT licensed, maintained by the UK Ministry of Justice's analytical services team, and the repository's last push was on 2026-09-28, with v5.0.0 released on 2026-09-18.

## Fellegi-Sunter, blocking rules and the SQL backend split

The statistical core is the Fellegi-Sunter model, the same family of model described in the academic paper that accompanies the R fastLink package. Every comparison column contributes a match weight, and the model combines those weights into a probability. Splink's customisations are term frequency adjustments, which down-weight common values like the surname Smith, and user-defined fuzzy matching logic such as Jaro-Winkler thresholds.

Training is unsupervised. The quickstart estimates three things in sequence: the probability that two random records match, given a recall value for a chosen blocking rule; the u parameters, the probability of observing a value among non-matches, via random sampling; and the m parameters, the same probability among matches, via expectation maximisation. No labelled pairs are required, which is the practical reason this fits administrative data where ground truth does not exist.

Blocking rules decide which pairs are ever compared. Comparing all pairs is quadratic, so the settings list rules like block_on("first_name") and block_on("surname") to restrict candidate pairs. The backend is separate from the model: DuckDBAPI() runs in-process on a laptop, and the README states Spark is used for 100+ million records. PostgreSQL and SQLite are supported through optional extras. sqlglot sits underneath to translate the generated logic per dialect, and pyarrow handles data interchange.

## Installing Splink and running the quickstart

Splink installs from PyPI. The README also gives a conda-forge alternative. The pyproject.toml requires Python >=3.10.0,<4.0.0, though the README text still says 3.9+; trust the packaging metadata when pinning an environment.

```bash
pip install splink
```

Backend support is opt-in through extras. Install the Spark or PostgreSQL extra only if you need that engine.

```bash
pip install 'splink[spark]'
pip install 'splink[postgres]'
```

The quickstart below builds a deduplication model on a bundled fake dataset, estimates parameters, predicts pairwise links and clusters them. The README gives this example, truncated here at the Linker construction.

```python
import splink.comparison_library as cl
from splink import DuckDBAPI, Linker, SettingsCreator, block_on, splink_datasets

db_api = DuckDBAPI()
df = splink_datasets.fake_1000
df_sdf = db_api.register(df, dataset_display_name="fake_1000")

settings = SettingsCreator(
    link_type="dedupe_only",
    comparisons=[
        cl.JaroWinklerAtThresholds("first_name", [0.9, 0.7]),
        cl.JaroAtThresholds("surname", [0.9, 0.7]),
        cl.ExactMatch("city").configure(term_frequency_adjustments=True),
    ],
    blocking_rules_to_generate_predictions=[
        block_on("first_name"),
        block_on("surname"),
    ]
)

linker = Linker(df_sdf, settings)
```

Training then runs through the three estimation steps described in the README, followed by predict() and clustering. The documentation states that a million records can be linked on a laptop in around a minute; that figure comes from the project's own feature list, not from an independent run. The interactive visualisations are the part worth opening first, because they show which comparison levels are pulling weights in unexpected directions.

## Where Splink breaks down: correlated columns and single-column data

The README is unusually direct about the input shape Splink wants: multiple columns that are not highly correlated. City is a poor companion to postcode because one predicts the other, and a model built entirely from correlated columns has little independent evidence to weigh. Correlation is not a bug in the implementation; it is a property of the Fellegi-Sunter assumption that comparison columns contribute conditionally independent evidence.

The harder boundary is the bag-of-words case. A table with one company name column and nothing else is explicitly outside the design. If that is your data, a string-similarity index or a dedicated entity resolution service is a better starting point than a probabilistic model with one comparison.

Two operational limits are worth naming. Blocking rules that are too narrow silently drop true matches before the model ever sees them, and the recall value passed to estimate_probability_two_random_records_match is a user assumption, not something Splink derives. The README does not document rollback of a trained model or a supported downgrade path between major versions, so pin your version in production.

## Splink versus fastLink and recordlinkage

The README itself names fastLink, an R package, as sharing the same underlying model, and links the academic paper describing it. The difference is execution: fastLink is R-native, while Splink pushes the pairwise comparison work into a SQL engine and keeps only orchestration in Python. That is what lets the same settings object run on DuckDB locally and on Spark at scale.

Against Python's recordlinkage package, the divergence is in the training step. recordlinkage expects you to supply or construct labelled match and non-match pairs for supervised classification. Splink's expectation maximisation estimates parameters from unlabelled data, which matters when no ground truth exists. The trade-off is interpretability of the fitting process: EM converges on a local optimum, and the visualisations exist precisely because the parameters are not directly inspectable as a confusion matrix would be.

Senzing is a commercial entity resolution product rather than a Python library, so the comparison is mostly about control and cost. Splink gives you the model and the SQL; it does not give you a managed service, a support contract or a curated knowledge base of name variants.

## Version 5, maintenance and licence terms

Splink 5 was released on 2026-09-18, ten days before the most recent push to master. The README flags the release with a link to new syntax examples and a release announcement, and the CHANGELOG.md at the repository root is the place to check what moved. The 4.x line continued in parallel, with v4.0.17 published on 2026-09-03, so a 4.x to 5.x upgrade is a real migration rather than a patch.

Upgrade cost is concentrated in the settings and API surface, since the backend abstraction and the Fellegi-Sunter core are stable across the change. Budget time for re-running training and re-checking thresholds after any major bump; match weights are not guaranteed to be numerically identical across versions.

The licence is MIT, which permits commercial use and modification. The practical implication is that you can embed Splink in a proprietary pipeline, but you inherit the obligation to keep the copyright notice. This is a description of the licence text, not legal advice; check with your own counsel for anything consequential.

## Conclusion

Adopt Splink if you have several partly independent columns per entity and no labelled training pairs, and you are comfortable reading interactive diagnostics before trusting a match threshold. Skip it if your only signal is a single free-text column, since the README states that a bag-of-words table is outside its design. Before committing, install the backend extra you actually need, run the quickstart against splink_datasets.fake_1000, and check the pairwise and cluster outputs on your own data.

## FAQ

### Is Splink open source?

Yes. Splink is released under the MIT licence, and the repository is public under moj-analytical-services/splink. The package is published on PyPI and can also be installed from conda-forge.

### What are the limitations of using Splink?

The README states that Splink performs best with multiple columns that are not highly correlated, and that it is not designed for linking a single column containing a bag of words. Correlation between columns is particularly problematic when all input columns are highly correlated.

### How do I install Splink?

Install the released version from PyPI with pip install splink, or use conda install -c conda-forge splink. Backend-specific support is added through extras such as splink[spark] or splink[postgres].

### How do I use Splink for deduplication?

The README quickstart registers a dataframe with a backend such as DuckDBAPI, builds a SettingsCreator with comparisons and blocking rules, trains with estimate_u_using_random_sampling and estimate_parameters_using_expectation_maximisation, then calls predict and clustering on the Linker.

## Sources

- [License: MIT](https://github.com/moj-analytical-services/splink/blob/master/LICENSE)
- [moj-analytical-services/splink on GitHub](https://github.com/moj-analytical-services/splink)
- [Project website](https://moj-analytical-services.github.io/splink/)
- [README](https://github.com/moj-analytical-services/splink/blob/master/README.md)
- [Releases](https://github.com/moj-analytical-services/splink/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/moj-analytical-services-splink
