Open-source project
OpenRefine/OpenRefine avatar
OpenRefine/OpenRefine

OpenRefine: a local data cleaning workbench for messy spreadsheets

OpenRefine is a free, open source power tool for working with messy data and improving it

12,016 stars2,173 forksJavaBSD-3-Clause

At a glance

What is it?
OpenRefine is a Java-based tool that runs in your browser but keeps your data on your own machine. It is best suited to exploratory cleaning, faceting and reconciliation of tabular data, not to scheduled pipelines.
Who is it for?
Adopt OpenRefine if you need to inspect and repair a messy table interactively, especially when the data cannot leave your machine or when you want to reconcile names against an external source before exporting.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OpenRefine actually solves, and for whom

The README describes OpenRefine as a Java-based power tool that lets you load data, understand it, clean it up, reconcile it, and augment it with data from the web, all from a web browser and on your own computer. The phrase that matters most is "your own computer". The application starts a local server and you interact with it through a browser, so the dataset never has to be uploaded to a third-party service. That constraint alone decides the audience: data journalists working with leaked or embargoed records, researchers handling personal data under institutional rules, and librarians or open-data maintainers aligning names against a shared authority.

The topics list on the repository points in the same direction: data-wrangling, datacleansing, reconciliation, wikidata, datajournalism. These are tasks where a human needs to look at values, group them, spot the near-duplicates, and decide what the canonical form should be. OpenRefine is built around that loop. It is not a spreadsheet replacement for arithmetic and charting, and it is not a database. It is a workbench for the stage between "we received a file" and "we can load this into something else".

How the local server, facets and reconciliation fit together

From the repository layout, OpenRefine is a Maven multi-module Java project. The top level holds main/, modules/, extensions/, server/, conf/, packaging/ and a pom.xml, with two launcher scripts, refine and refine.bat. The README states that running ./refine on macOS and Linux, or refine.bat on Windows, is what starts the application from a cloned repository. The browser is the interface; the Java process is the engine.

Two mechanisms carry most of the cleaning work. Facets let you summarise a column and then filter or edit the rows behind a chosen value, which is how you turn "I think there are spelling variants" into "here are the 14 variants and their counts". Reconciliation is the second: the topics list includes reconciliation and wikidata, and the README lists augmenting data with data coming from the web as a core capability. In practice that means matching your values against an external authority and pulling back identifiers or extra fields. The data flow is therefore local-first: load a file into the project, transform it in memory through facets and edit operations, reconcile against a remote service when you explicitly ask for it, then export.

The extensions/ directory is where optional capabilities live, and the README's licensing section notes that each extension carries its own licenses folder. That matters when you evaluate what is actually bundled versus what you add.

Installing OpenRefine and running a first facet

The README points to the GitHub releases page as the download location. There is no package manager command documented in the README, so the honest instruction is to fetch a release archive from that page and unpack it. If you prefer to run from a clone, the README gives the launcher scripts and the toolchain requirements.

To run from source on macOS or Linux, the README specifies JDK 11 or newer, Apache Maven, and Node.js 18 or newer. The command is a single script:

bash
./refine

On Windows the equivalent is refine.bat. After the build finishes, OpenRefine opens a browser tab pointing at the local instance. What you should see is the create-project screen, where you paste or upload data and get a preview before committing to a project.

Once a project exists, the first useful action is a facet on a column you suspect is dirty. The point is not to change anything yet, just to see the distribution of values. If a column that should hold a dozen country names shows 60 distinct values, you have found your first cleanup target. From there you can cluster near-identical values and apply a merge, or edit individual cells. The README does not document a specific menu path for clustering, so follow the user manual at openrefine.org/docs for the exact clicks.

One practical note on versions: the most recent release listed is 3.10.1, published on 2026-03-04. If you are following a tutorial written against an older release, expect the interface to differ.

Where OpenRefine is the wrong tool

The clearest limitation follows from the architecture. Because the application is driven through a browser and started by a script, it is an interactive tool, not a batch processor. If your requirement is a nightly job that ingests a CSV, applies the same transformations, and writes to a warehouse, OpenRefine is a poor fit. There is no documented headless mode in the README, and the operations you perform are recorded as steps in a project rather than as a declarative spec you can check into version control and run unattended.

Scale is the second boundary. OpenRefine is designed for datasets a person can inspect. Work that involves hundreds of millions of rows belongs in a query engine, not in a browser tab. The project's own framing supports this reading: it is a "power tool" for understanding and improving data, language that implies a human in the loop.

The third limitation is reconciliation itself. Matching against an external authority depends on a remote service being reachable and on your values being close enough to something in that authority. The README does not describe offline reconciliation or a bundled authority file, so a network-restricted environment will limit that capability. If your environment blocks outbound requests, verify what reconciliation services are reachable before you build a workflow around them.

OpenRefine against a scripted pandas workflow

The most common alternative for the same job is a Python notebook with pandas. The difference is not capability, it is where the state lives. In pandas, every transformation is code, so the cleaning is reproducible, diffable and runnable in CI. The cost is that you cannot see a column until you write the expression to display it, and a non-programmer on the team cannot participate.

OpenRefine inverts that trade-off. The transformations are operations you apply through the interface, the effect is visible immediately, and a domain expert can do the work without writing Python. The price is reproducibility: the README does not describe exporting a project's operation history as a script that another tool can execute, so the record of what you did lives with the project rather than in a repository.

A spreadsheet is the other comparison people reach for. A spreadsheet is fine for a few hundred rows with consistent types. It becomes painful when you need to see the distinct values of a column with counts, apply the same normalisation across thousands of rows, or match values against an external list. Facets and reconciliation are exactly the operations a spreadsheet makes awkward, which is the gap OpenRefine occupies.

Maintenance, licensing and the cost of staying current

The repository is not archived, and the last push was on 2026-09-18. The release cadence visible here is modest: 3.10.0 on 2026-02-26, followed by 3.10.1 on 2026-03-04. The README states plainly that OpenRefine is maintained by a small core team and relies on grants plus community support, with fiscal sponsorship by Code for Science and Society since 2020. That is the real maintenance picture: a small team, funded project by project, supported by a community forum rather than a commercial vendor with an SLA.

For an operator, that translates into a specific cost. Upgrades come as release archives you download and unpack, because no package manager command appears in the README. There is no documented automatic update path, so someone has to watch the releases page and repeat the manual step. Running from source adds a toolchain to keep current: JDK 11 or newer, Apache Maven, and Node.js 18 or newer.

On licensing, OpenRefine is BSD-3-Clause, with the licence text in LICENSE.txt. The README points to the licenses folders under main/webapp/ and inside each extension for the libraries OpenRefine depends on. If your organisation has a policy about bundled dependencies, those folders are where you check, and the extensions directory is where the set of dependencies can grow. This is a description of what the repository says, not legal advice.

Editorial conclusion

Adopt OpenRefine if you need to inspect and repair a messy table interactively, especially when the data cannot leave your machine or when you want to reconcile names against an external source before exporting. Do not adopt it as a scheduled ETL component: the README describes a browser-driven tool, not a headless pipeline, and the project is maintained by a small core team that relies on grants and community support, so plan for manual upgrades between releases rather than automatic ones. Before committing, verify that your JDK is 11 or newer, that Node.js is at least 18 if you intend to run from source, and that the BSD-3-Clause licence and the bundled library licences under main/webapp/licenses and each extension's licenses folder are acceptable to your organisation.

Frequently asked questions

What is OpenRefine used for?

The README describes it as a tool to load data, understand it, clean it up, reconcile it, and augment it with data from the web. The repository topics point at data cleaning, data wrangling and reconciliation for journalism, open data and Wikidata work.

How do I install OpenRefine?

The README points to the GitHub releases page for downloads. To run from a cloned repository instead, the README gives ./refine on macOS and Linux or refine.bat on Windows, which requires JDK 11 or newer, Apache Maven and Node.js 18 or newer.

Is OpenRefine free and open source?

Yes. The README states OpenRefine is open source software licensed under the BSD license located in LICENSE.txt, and the repository licence is BSD-3-Clause.

Is OpenRefine safe to use with sensitive data?

The README says the tool works from a web browser and on your own computer, so the data stays local. Reconciliation and augmentation features do contact the web, so the README does not support a blanket claim that no data ever leaves the machine.

What is a facet in OpenRefine?

The README does not define facets. It describes the tool as loading data, understanding it, cleaning it up, reconciling it and augmenting it, and the user manual at openrefine.org/docs is where the README sends readers for documentation.

Official sources

  1. License: BSD-3-Clause
  2. OpenRefine/OpenRefine on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openrefine-openrefine.svg)](https://hysenlabs.com/projects/openrefine-openrefine)