Congress-legislators ships YAML at the root and generates the JSON and CSV people download
Members of the United States Congress, 1789-Present, in YAML/JSON/CSV, as well as committees, presidents, and vice presidents.
At a glance
- What is it?
- A public-domain legislative dataset whose current and historical files split on a date comparison, whose records carry twelve cross-walk identifiers, one of them documented against a retired service, and whose only automated import listed alongside a publication that no longer publishes.
- Who is it for?
- This is the reference dataset it claims to be, with two caveats a user should know about. The file split between serving and former members is not reliable on its own, and the documentation says so, so load both.
- Can I use it commercially?
- Yes. CC0-1.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The root holds eight YAML files and none of the JSON or CSV
The file table lists nine rows, and the repository root contains eight YAML files and nothing else in that family.
Every row in the table offers download links. Seven of the eight offer three formats. One, the current committees file, offers two. It is the only file in the table without a CSV link, and it is easy to miss because every neighbouring row has three.
More interesting is where the other two formats live. The documentation says the files are maintained in YAML in the main branch, and that the CSV and JSON files are also provided, in a separate branch, with a link to the project's own hosted site. So the repository root holds the sources and the downloads are generated artefacts.
That is a sound arrangement and it has one practical consequence. If a consumer downloads a malformed CSV, the cause is in the generator rather than in the data, and fixing it means a commit to a different branch from the one a data contributor would be editing. It also means a stale download and a stale source fail differently: a source edit lands in the main branch immediately, while a generated file reaches the download site only after the build runs.
For anyone reading this as a plain data source, the practical rule is that the file you are editing is not the file you are downloading.
The current and historical split is a date comparison, and the docs say not to trust it
The documentation is unusually candid about the boundary between the two legislator files, and what it says is worth reading twice.
It says the split between the current file and the historical file is somewhat arbitrary, because those files may not be updated immediately when a legislator leaves office. And it says that if it matters to you, just load both files.
The mechanism underneath is a date comparison rather than a flag. Each legislator record has a list of terms served, in chronological order, and the current term is always last. To check whether somebody is currently serving, you check whether the end date on that last term is in the future.
So there is no boolean anywhere in the schema. A member who resigned in March and whose record was updated in May sits in the current file until the update lands, and a member whose term end date has passed but whose record is still marked as a future end date stays in the current file for as long as nobody corrects it.
That is a defensible design for a file maintained by volunteers, and it is also why the documentation tells you to load both. A consumer that treats the current file as authoritative will eventually treat a departed member as serving. Merging the two files is the documented answer, and the deduplication falls out of the identifier structure.
Each record carries twelve cross-walk identifiers, and the documentation names a primary key
Look at the identifier block in the example record and count it.
There is the biographical directory identifier, then a numeric identifier from the project's own former web service, then four identifiers from campaign finance and voting-record projects, a list of one or more finance identifiers, a cable network identifier, two encyclopedia and wiki identifiers including a name rather than a number, a mapping identifier, a house history identifier, and a scholarly identifier.
Twelve keys, spanning official government directories, finance disclosure, voting records, encyclopedias and an academic dataset. That block is the interoperability contract: it is what lets a researcher join this file to a finance dataset or a voting dataset on a key that both sides agree on.
The documentation also says which one to use, and the answer is the biographical directory identifier. It calls that the best field to use as a primary key, and it explains why only one is included per person: at one point that directory had two entries for some members, specifically women who changed their name on marriage, and the project chose one rather than carrying both.
That is a documented data-cleaning decision, and documenting it is the difference between a dataset you can trust with a join key and one you cannot.
One caution. The numeric identifier is documented against the project's former web service and its successor beta domain, and the same service is listed further down as an automated import source. If you are building a pipeline on that identifier, check its current status before you depend on it.
The data is a merge of volunteer edits and seven automated imports, two of them academic
The provenance section is the most valuable paragraph in the documentation and it is short.
The dataset is maintained by a combination of manual edits by volunteers and automated imports. The volunteer side names four organisations, three of them journalism or transparency projects, plus others. The automated side lists seven sources.
Two of those seven are academic rather than governmental. One is a historical standing committees data set attributed to two named researchers at a Massachusetts university department, and one is a roll-call voting records dataset attributed to two other named authors, covering the period from 1789 to 1990. Both are the kind of dataset that took years to assemble and are maintained by people who will not update them again.
The other five are the biographical directory, a project that tracks every vote by every member, a congress application programming interface from a transparency lab, a legislative tracking website, and a cable network's chronicle of proceedings.
Read together, this tells you something useful about the file: it is a consensus artefact, assembled from about a dozen sources of uneven reliability, and its accuracy in any given field depends on which source won that merge.
It also tells you where to report a problem. If a term's dates are wrong, the likely culprit is an import rather than a volunteer.
The licence is a public-domain dedication and the records contain phone numbers and street addresses
The repository is released under a public-domain dedication, the most permissive licence available, and the example record in the documentation contains a full contact block: a street address in a congressional office building with a room number, a telephone number, a facsimile number, a contact form URL and a member's personal website.
That combination deserves a sentence from the project that it does not currently have.
A congressional office address and telephone number are official work contact details, published by the House and the Senate, and their inclusion in a public dataset is unremarkable. What is worth thinking about is the interaction with a public-domain dedication: once a dataset is dedicated to the public domain, a downstream redistributor has no licence obligation to preserve anything, and there is no mechanism for a person to ask for their number to be removed from a copy already in circulation.
There is one place where the project draws a line, and it draws it well. The social media file is restricted to official accounts and explicitly excludes campaign and personal accounts, which is exactly the right call for a dataset that distributes handles.
So the reasoning is understood. It is applied to handles and not to phone numbers. Whether that distinction is worth making is a question for the maintainers, and it is a question nobody has asked on the page.
Coverage is asymmetric across the files, and the table does not say so
The opening line states the coverage of everything in one sentence, and the table underneath states it again, slightly differently.
Members of Congress are covered from 1789 onward, which is the whole of the republic. Committees are covered from 1973 onward, attributed to the ninety-third Congress. Committee membership is current only. District offices are for current members. Presidents and vice presidents are in their own file with no stated date range.
So the oldest data in this repository is two and a half centuries of legislators, and the newest structured data is about seventy committees, and between those two extremes sits one file that is a snapshot and will be a snapshot forever. A user building a time series over committee assignments has almost nothing to work with, and a user building one over legislator terms has everything.
The table's descriptions do say the date ranges, but they say them per row rather than as a shape. Read as eight independent files, the coverage looks uniform. Read as a dataset, it is one deep archive, one medium archive and one snapshot.
That asymmetry is not a defect. It reflects what the upstream sources actually contain, and the two academic datasets that cover committees and votes both stop well before the present. But it is the thing to know before choosing this repository for a question about how committee work has changed over time.
There is a test directory, and the readme never mentions it
The repository is mostly data, but it is not only data.
At the root there is one Python script, a lockfile for a Python environment manager with its lock, and three directories that are not data: a miscellaneous directory, a scripts directory, and a test directory. There is also a continuous integration badge at the top of the page, pointing at a hosted service rather than at the platform's own workflow system.
None of that appears in the documentation. The documentation describes the schema, the formats, the provenance and how to join the files, and it says nothing about the fact that the YAML is validated somewhere.
For a dataset maintained by manual edits across a dozen upstream sources, that test directory is the most important thing in the repository. It is what catches a duplicated identifier, a term that ends before it starts, a record whose last term has an end date in the wrong format, a state that is spelled two ways, and a field that drifted from the data dictionary. A volunteer editing one record by hand is exactly the case where those checks matter and exactly the case where nobody reviews the diff carefully.
So the credit is due and the documentation is missing. A contributor fixing a typo has no way to know from the page that there is a test suite to run, and the pull request that adds a malformed record looks exactly like the one that adds a good one.
Editorial conclusion
This is the reference dataset it claims to be, with two caveats a user should know about. The file split between serving and former members is not reliable on its own, and the documentation says so, so load both. And the formats most people consume are generated from YAML on a different branch, so a bad download points at the generator rather than at the data. For the identifier question, use the one the documentation tells you to and you will not go wrong.
Frequently asked questions
What does unitedstates/congress-legislators contain?
Eight data files covering members of Congress from 1789 onward, current committee assignments, committee history from 1973, district offices for current members, and presidents and vice presidents. Each is maintained as YAML in the main branch, with JSON and CSV generated for download from a separate branch.
How current is the legislators-current file?
The documentation says the split from the historical file is somewhat arbitrary because the files may not be updated immediately when a legislator leaves office, and advises loading both. Whether somebody is serving is determined by checking that the end date on their last term is in the future, since there is no flag for it.
Which identifier should I use as a primary key for legislators?
The biographical directory identifier, which the documentation calls the best field for the job. Each record also carries cross-walk identifiers for around a dozen other databases covering finance disclosure, voting records, encyclopedias and an academic dataset.
Does the dataset include social media accounts?
Yes, in a separate file for current members, and it is restricted to official accounts. The description explicitly excludes campaign and personal accounts.
Where does the data in congress-legislators come from?
A combination of manual edits by volunteers affiliated with several named organisations and automated imports from seven sources, two of which are long-running academic datasets on historical committees and on roll-call voting records, plus a project that tracks every vote, a congress API, and a cable network's chronicle of proceedings.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/unitedstates-congress-legislators)