Library / SDK
yuchenlin/rebiber avatar
yuchenlin/rebiber

rebiber has three version numbers and a PyPI name frozen in 2021

A simple tool to update bib entries with their official information (e.g., DBLP or the ACL anthology).

3,037 stars165 forksPythonMIT

At a glance

What is it?
A command line tool that replaces unofficial preprint bibliography entries with the published record, keeping your cite keys, using a bundled set of conference dumps it does not query live by default. The mechanics are unusually carefully specified, and two operational facts sit beside them: the output flag defaults to overwriting your input file, and the version you get from the package index is four years old.
Who is it for?
Use this if your bibliography is full of preprint entries with journal names that were never published, because that is the exact problem it solves and the matching rules are conservative enough that a wrong replacement is unlikely. Three things to know before you run it.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Three version numbers and a package index entry from 2021

There are three version numbers in play and they do not agree. The manifest says one point four. The newest release tag is one point three, published in August 2026, and the two before it are a one point one point three from September 2021 and a one point one point one from February 2021. And the documentation's install section says in bold that installation is from the git host only, because the name on the public package index is a release from 2021 at one point one point three. That last fact is the one that bites. The project owns that name and has not published to it since 2021, so anyone who types the obvious install command gets a four year old version of a tool whose current release requires a newer Python floor and a newer parser library, with no warning from the package manager. The install lines the file does give, both of which name the git address explicitly, are the correct ones, and the reason is spelled out one sentence above them.

The two supported install routes are:

bash
uv tool install git+https://github.com/yuchenlin/rebiber
# or
pip install "rebiber @ git+https://github.com/yuchenlin/rebiber"

Omitting the output flag rewrites your bibliography in place

The flag table gives the output flag a default of in place, and the usage section says the same thing in prose: omitting it overwrites each input file. The usage block then says, in two words, start with the dry run flag. That combination is the most important operational fact in the file and it is stated in the right order, which is to say the destructive default comes first and the advice comes after it. The dry run flag prints the change report and writes nothing, and a companion flag can write that same report to a file, which is what you would actually want before accepting a batch of replacements. The output flag also accepts a directory when several input files are given, and one example passes a glob as the input and a directory as the output, so the batch form is supported. The dry run is not optional politeness here; it is the difference between a diff you read and a bibliography you rewrite.

A short flag whose false value means do the thing

Six of the flags are single letters with non obvious defaults, and one is genuinely counter intuitive. The duplicates flag defaults to keeping duplicate cite keys, and setting it to false means do not keep them, which is what the change notes for one point four then describe: field removal, venue abbreviation, field selection, title protection and sorting are all applied to entries with repeated cite keys when that flag is false. So a false value on that flag switches on five transformations for the entries most likely to be affected by them. The venue abbreviation flag takes an explicit true and defaults to false, so abbreviations are opt in and driven by a table shipped inside the package, and a separate flag lets you point at your own table instead. A flag for custom venue lists behaves the same way, defaulting to the packaged list. A short flag refreshes the packaged dumps from the main branch, and another prints the version.

Digit aware titles, because the old dumps merged 16x16 and 32x32

The matching rules are the substance of this tool and they are written down precisely. A title hit converts only if the first author's last names match, or if at least two last names overlap; empty or missing authors never match at all, and a flag that skips the author check still will not convert a record with no authors. Trailing forms of et al and the phrase and others are stripped before comparing. Titles are keyed digit preserving first, so a sixteen by sixteen matrix and a thirty two by thirty two matrix are different titles, and only if that fails does it fall back to a letters only key, which the file attributes to older dumps. That fallback is an admission of a real historical bug: the earlier data could not tell those two apart. For the live search path a published record is preferred over a preprint one, and if several published candidates remain the file says it skips, which is a refusal rather than a guess.

Fully offline needs a flag, because the preprint server is queried by default

Two flags look like they turn networking off and neither does on its own. The live lookup flag, off by default, is the one that sends a title to the bibliographic database when a local dump misses, capped at five hits and with a note about respecting that service's rate limits. The flag that claims to be formatting only, also off by default, skips the database and preprint rewrites entirely and just pretty prints, which is the one to use for a genuinely offline run. But even with both off, the file notes that leftover unofficial preprint entries may still query the preprint server for a year and a subject class. So the default path is not offline, the flag that sounds like it makes it offline does, and the reason is that the year and subject class are not in the local dumps. That is a reasonable design, stated in one sentence, and it is the sentence to read before running this on a machine with no route out.

ACL gets workshops and main tracks, everything else gets main only

The venue table covers two dozen conference families with the year each dump covers, and the file is careful about what those years mean. It says outright that the years are what is in this repository rather than whatever the upstream source has tomorrow, which is the honest framing for data refreshed by automation. The coverage is uneven in two ways a reader will notice. The anthology, which is the computational linguistics family, includes main conferences and workshops; every other family is main tracks only, so computer vision and robotics workshops are simply absent rather than wrong. And within the table the ranges are irregular where the series are irregular: one conference is listed with a starting year and then a jump, skipping two consecutive years with no note, while a nearby vision family lists only two biennial years because the most recent one is not upstream yet. A final line names what is missing and why, including one venue whose upstream table of contents is empty.

The data arrives by monthly pull request, so the ranges have a date

The data is not fetched on demand. It is a set of per conference, per year files inside the package, refreshed by a scheduled job that opens a pull request with fresh data, and the year ranges in the table are whatever that job last produced. Adding a venue by hand is one command, which downloads a range for a named conference, writes the file, and appends the conference name to the list of what is packaged when it is not already there; the invitation is for pull requests rather than for people to run it locally. Three details of the layout matter if you write a script against it. The computational linguistics dump is split across three files rather than one per year, so a naive glob over the data directory will not find it. The venue list and the abbreviation table are plain text files inside the package and can be overridden by pointing the flags at your own copies. And a separate flag refreshes the packaged dumps from the main branch at run time, which is the one feature in this list that reaches the network without asking.

Editorial conclusion

Use this if your bibliography is full of preprint entries with journal names that were never published, because that is the exact problem it solves and the matching rules are conservative enough that a wrong replacement is unlikely. Three things to know before you run it. Start with the dry run flag, because the output flag defaults to writing over the file you pointed at, and the documentation's own advice is to start there. Install from the git address rather than the package index, because the name on that index is a release from 2021 and installing it gets you the tool from before the current parser migration and the current Python floor. And if you need a fully offline run, pass the format-only flag, because leftover preprint entries otherwise query the preprint server for a year and a subject class. The matching rules are the best part of this project: it refuses on ambiguity rather than guessing, and it never drops an entry it could not parse.

Frequently asked questions

What does Rebiber do?

It replaces unofficial or preprint bibliography entries with the official published record from a bibliographic database or the computational linguistics anthology, keeping your existing cite keys so nothing in your text has to change. It also offers formatting only mode, field removal, venue abbreviation, field allowlisting and sorting.

How do I install Rebiber?

From the git host, not the package index. The file says installation from GitHub only and warns that the package index name resolves to a release from 2021. It gives two commands, one for the tool installer and one for pip with an explicit git address, plus a development path that clones and syncs an extra before running the tests.

Does Rebiber modify my .bib file in place?

Yes, if you omit the output flag, because its default is in place and the usage notes say that omitting it overwrites each input file. The file's own advice is to start with the dry run flag, which prints the change report and writes nothing, optionally to a report file.

Which venues does Rebiber cover?

About two dozen conference families, with a per year range for each, and the file states that the ranges are what is in the repository rather than what the upstream database has. The computational linguistics anthology covers main conferences and workshops; the other families are main tracks only, and one historical set is frozen at around 2020.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. yuchenlin/rebiber on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yuchenlin-rebiber.svg)](https://hysenlabs.com/projects/yuchenlin-rebiber)