# A corpus of conference metadata with no licence file, no licence field, and a default input path that misses its own file

> The data repository behind a commercial paper-search site: JSON records per conference per year, plus a Streamlit browser and a small extraction script. The table of years, the licensing gap and one default argument are the parts worth reading.

**papercopilot/paperlists** — Processed / Cleaned Data for Paper Copilot

- Repository: https://github.com/papercopilot/paperlists
- Website: https://papercopilot.com/
- Stars: 958 · Forks: 48
- Language: Python
- License: not declared
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/papercopilot-paperlists

## No licence field and no licence file

This is the first thing to establish and the repository does not establish it.

The repository's own metadata carries no licence value. The root listing contains no licence file. There is no licence text in the readme either, and no note pointing anywhere else.

That is a different situation from the common one where a repository carries a permissive licence in its metadata but is easy to miss. Here there is nothing to find. The data itself is derived: the readme says records are combined from multiple sources, naming a submission review platform, official conference sources and open access sites, so the underlying facts belong to whoever published them, and the page offers no terms for the assembled files.

For anyone planning to redistribute, merge into a larger corpus, or train on this, that is a blocker rather than a formality. The honest reading is that the omission may be an oversight in a repository that exists to serve one website, but the burden is on the user to establish permission, and nothing in the repository helps with that.

The one piece of formal metadata that is present is a Croissant descriptor at the root, which is a machine-readable dataset format. It describes the shape of the data. It is not a licence.

## The default input path points at a file that is one directory down

The command line tool takes a keyword and four optional arguments, and one of those defaults is wrong for this repository.

```bash
python extract.py [keyword] [-i INPUT_PATH] [-o OUTPUT_FILE] [-f FIELDS...]
```

The default for the input path is given as a bare filename, `iclr2025.json`, but the actual file in the repository is inside a conference directory, and the readme's own worked example has to say so explicitly:

```bash
python extract.py retrieval -i iclr/iclr2025.json -o results.json -f title keywords
```

So the tool as documented fails on the documented corpus unless you remember to add the directory prefix. A user who reads the flag description, omits the input argument and gets nothing, has no reason from the output to guess that the fix is a path segment.

The other two defaults are sound. Output is optional, and the field list defaults to four names, which is the part of this tool that tells you what the data actually looks like.

## Four default field names, and they are the review platform's vocabulary

The search defaults to four fields, and the list is more informative than the flag description:

The flag description gives the default set as keywords, title, primary area and topic.

A primary area, a topic and a keyword list are the field names a submission review platform uses on its own records, not a schema someone designed for this repository. That tells you the records were taken from such a platform rather than assembled from several independent sources at the field level, and it tells you the primary area is a small closed set of broad subjects while keywords are author-supplied and therefore inconsistent.

The worked example narrows the search to two of the four, title and keywords, which drops the venue-level subject classification. That is a sensible thing to do when looking for a specific technique, and it is worth knowing that the classification field is the noisiest of the four.

One thing the tool is not: the page calls the tool a search over conference papers and gives no indication of any semantic or embedding component. It is field-limited substring or token matching, which is fast and predictable and will miss a paper that uses different words for the same idea.

## Two ICLR years have no row, and one NeurIPS year has data but no statistics

The listing is the reference for what exists, so its gaps are the interesting part.

ICLR is presented in two tables. The first covers 2020 through 2025 with a JSON link and a statistics link for every year. The second covers 2013 through 2019, and two of those columns are empty in both rows. 2016 and 2015 have neither a data file nor a statistics page.

NeurIPS has a different kind of gap. Its data rows run from 2000 to 2024 with a link in every cell, so the files exist across twenty-five years. The statistics rows are thinner: the main statistics row is empty for 2020, and the separate datasets-and-benchmarks statistics row only has links for 2024, 2023 and 2022.

So the pattern is that the data is broader than the analysis. That is a reasonable division of labour for a repository that feeds a website, but it means you cannot tell from the listing whether a missing statistics page means the year was not processed or the page was simply never published.

Timing is worth noting too. The newest ICLR file listed is 2025 and the newest NeurIPS file is 2024, while the last commit to the default branch is dated 2026-07-01. Either the most recent cycle has not been added or it exists without being linked from the table.

## The directory is called nips and holds twenty-five years of files

NeurIPS is presented under a heading that carries both names, and its data lives in a directory using the older abbreviation.

Inside that directory every file carries the same prefix followed by a year, so the newest is named nips2024.json and the oldest in the second table is nips2010.json.

Every file from 2000 to 2024 is named with that prefix, and the directory is named the same way, while the statistics links on the same page use the current name. The conference was renamed years before most of these files were collected, so this is a legacy name preserved for continuity rather than a mistake.

It matters more for tooling than for reading. A script that discovers directories by name will find a conference called nips and not one called neurips, and a join against any external dataset that keys on the modern abbreviation will not match without a mapping. Given that the whole point of the repository is to make records from different sources agree, the one place where the project's own naming has not been normalised is a small irony worth knowing about.

The other thirty directories follow the same convention, one per conference, with the year in the filename. That uniformity is the repository's main strength and the reason a single extraction script covers all of them.

## A Streamlit browser, a local URL, and a requirements file the root listing does not show

The local tool has two front ends. The web one is a Streamlit application:

```bash
cd paperlists/tools
streamlit run app.py
# a corresponding local url will popsup, e.g. `Local URL: http://localhost:8501`
```

The default port is the Streamlit default and the example URL is loopback, which is the right shape for a tool that indexes your whole corpus. No host binding is specified in the command, so what the server actually listens on is the framework's default rather than a decision recorded here.

The install instruction is thinner than the tool deserves:

```bash
git clone https://github.com/papercopilot/paperlists.git
pip install -r requirements.txt
```

Two things about it. The requirements file does not appear in the repository's root listing, so a reader cannot confirm from the tree that the pinned dependency set is committed or where it lives. And the only version guidance anywhere is inside a comment, a commented-out conda command creating an environment with Python 3.10, so the interpreter floor is a suggestion rather than a requirement.

The tool itself is credited to a named outside contributor, which is worth noting because it is the one part of this repository that is not simply the data.

## Thirty-two directories, one metadata descriptor, and the analysis lives on a website

The repository's shape is a list of conference directories and very little else.

There is no source directory, no build configuration, no test suite and no continuous integration at the root. Alongside the thirty-one conference directories there is a tools directory, and alongside that a single Croissant descriptor file.

The Croissant file is the one piece of forward thinking here. It is a machine-readable format for describing a dataset, so an external tool can discover what these JSON files are, what shape they take and how to obtain them without parsing this page.

What the repository does not contain is the analysis. Every statistics link in every table points at the product's own site, and the site also hosts the per-year pages and the conference landing pages. So the division of labour is: this repository holds the records, and the website holds everything computed from them.

That is a sensible split, and it is also the reason a user should decide early whether they want the data or the numbers. Downloading this repository gets you JSON and a local browser. It does not get you acceptance rates, topic distributions, or any of the other statistics the tables link to, and there is no script in the repository that produces them.

## Conclusion

paperlists is worth using if you want a local, offline copy of a conference's paper metadata in one consistent schema, because the whole point is that records normally scattered across a submission platform, a conference site and open access indexes arrive in one shape you can filter. The extraction script is small enough to read, and the Streamlit browser is enough to find out whether the coverage is what you need. Three things to check first. Settle the licensing question, because the repository records no licence and its root holds no licence file, so there is no stated basis on which to redistribute, merge or build on these records even though the underlying fields are third-party metadata. Confirm the years you need are actually present, because the listing has gaps in both the machine-readable rows and the statistics rows, and the newest data listed is 2025 for one conference and 2024 for another while the last commit is dated this past July. And pass the input path explicitly, because the script's default filename does not match where the file actually sits in the repository. Who should not use it: anyone who needs the statistics rather than the records, since every analysis link points at the product's own site and none of that computation is in the repository. What this is not is a dataset with a documented provenance chain. The page says records are combined from multiple sources to make them coherent, and the fields are the submission platform's own vocabulary, but per-record source attribution is not described anywhere.

## FAQ

### What license does the paperlists data have?

None is stated. No licence value appears in the repository metadata, the root contains no licence file, and the readme points nowhere for terms. Since the records are described as combined from third-party sources, there is no stated basis on which to redistribute or build on the assembled files.

### How do I search the paperlists data from the command line?

Run the extraction script with a keyword and, if you want anything other than the defaults, pass an input path, an output file and a field list. Pass the input path explicitly, because the documented default is a bare filename while the file actually lives inside a conference directory.

### Which fields does the paperlists search cover by default?

Four: keywords, title, primary_area and topic. The names are the submission review platform's own field names, and the worked example narrows the search to two of them, title and keywords, dropping the subject classification.

### Which conferences and years are in the paperlists repository?

Thirty-one conference directories, with ICLR listed from 2020 to 2025 and again from 2013 to 2019 with two years empty, and the older abbreviation directory carrying files from 2000 to 2024. The newest files listed are ICLR 2025 and the 2024 set, while the last commit is dated 2026-07-01.

### Can I get the acceptance statistics from the paperlists repository?

No. Every statistics link in the repository's tables points at the product's own website, and no script in the repository produces those numbers. The repository holds the records; the site holds everything computed from them.

### How do I run the paperlists local search tool?

Clone the repository, install the requirements file, change into the tools directory and run the Streamlit application, which serves on the framework's default port. The only Python version guidance given is a commented-out conda command creating an environment with Python 3.10.

## Sources

- [Issues](https://github.com/papercopilot/paperlists/issues)
- [papercopilot/paperlists on GitHub](https://github.com/papercopilot/paperlists)
- [Project website](https://papercopilot.com/)
- [README](https://github.com/papercopilot/paperlists/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/papercopilot-paperlists
