# pandas-datareader: the README's pandas floor is two generations stale

> The package gives pandas remote access to macroeconomic and factor data, and its documentation has drifted from its own packaging in specific, checkable ways: a pandas floor stated twice with different numbers, a distutils requirement that requirements.txt answers with setuptools, a license the index declines to assert, and three different descriptions of what the library is for.

**pydata/pandas-datareader** — Extract data from a wide range of Internet sources into a pandas DataFrame.

- Repository: https://github.com/pydata/pandas-datareader
- Website: https://pydata.github.io/pandas-datareader/stable/index.html
- Stars: 3,271 · Forks: 692
- Language: Python
- License: NOASSERTION
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/pydata-pandas-datareader

## The README states a pandas floor two generations behind requirements.txt

Two files give a pandas floor and they do not agree. The README's Requirements section lists pandas>=1.5.3. The requirements file at the repository root lists pandas>=2.1.4.

Those two numbers are not weighted equally. The manifest declares dynamic = ["version", "dependencies", "readme"], and the setuptools configuration points dependencies at a file, so the installable dependency set is read from requirements.txt at build time. The README's list is prose. A reader who trusts it and installs against 1.5.3 gets a resolution that differs from what the package declares, and the failure surfaces later as an import error rather than as an install warning.

The other lines did not drift. requests>=2.19.0 appears in both files with the same floor, and lxml appears in both with no version constraint at all. Only the pandas line moved, which is what a floor does over time while the file it is compared against keeps getting updated and the document does not.

The gap matters because of what sits above it. requires-python is >=3.11 and the classifiers name 3.11 through 3.14, so the supported interpreter range is modern, and pairing that with a 1.5.3 pandas floor in the same document describes a combination the project does not actually test. The number to plan against is 2.1.4, and the one to distrust is 1.5.3.

## The README lists distutils as a requirement and installs setuptools

The README's requirement list has four entries: pandas, lxml, requests, and distutils. The fourth carries a parenthetical saying that distutils is not included in the standard library of Python 3.12 and above, and that this can be solved by installing setuptools.

The requirements file answers the same need with a different name. Its four entries are lxml, pandas>=2.1.4, requests>=2.19.0 and setuptools. distutils does not appear there at all, and setuptools does.

So the document lists a module as a dependency and then tells you to install a different distribution to work around that module being missing. The file that feeds the installer skips the intermediate step and names the workaround directly. Both routes land in the same place, which is why this is harmless in practice and confusing to read: the requirement list is not a description of what gets installed, it is a description of a problem and its fix.

The parenthetical is also broken as written. The phrase reads that distutils is not included in the standard library of Python 3.12 and above it, which leaves the sentence unfinished mid-clause.

Scope note: requires-python is >=3.11 and the classifiers stop at 3.14, so the 3.12 cutoff touches only part of the range the project claims. Anyone on 3.11 never meets the problem this line is warning about.

## The index asserts no license while the manifest declares BSD-3-Clause

Three signals describe the licensing here and they are not aligned.

The license field in the repository index is left at NOASSERTION, which is the platform recording that it did not determine a value. The manifest states license = "BSD-3-Clause" and adds license-files = [ "LICENSE.md" ]. The README carries a badge linking to LICENSE.md in the repository, and the root of the tree does contain a file by that name.

So there is a declared SPDX identifier, a named license file, and an index that declines to say which applies. The file itself is the only thing that could settle it, and its contents are not part of what can be read here. Picking BSD-3-Clause because it is the tidier of the two would be an assumption rather than a finding, and a compliance review that reads the index field will come away with nothing at all.

Two details are worth separating from that. The manifest uses the modern license-files mechanism rather than burying the filename in a classifier, so the packaging is explicit about which file travels with the distribution. And the project describes itself as formerly a component of pandas, which is a lineage note rather than a licensing statement, and does not by itself settle which terms carry over.

For anyone shipping this in a compliant environment, the resolution is a single file read, not an inference from metadata.

## Five years separate v0.10.0 from v0.11.0 under a Production/Stable classifier

The release history reads v0.11.1 on 2026-06-24, v0.11.0 on 2026-06-23, and v0.10.0 on 2021-07-13. So the gap between the last 0.10 release and the first 0.11 release is close to five years, and the two releases that ended it landed a day apart.

The classifier in the manifest reads Development Status :: 5 - Production/Stable. A package carrying that label is sitting at version 0.11.1 and has never published a 1.0. Those two facts are not contradictory in the packaging sense, since the version scheme and the maturity claim are independent, but they will read as contradictory to anyone deciding whether to depend on it.

The version is not written down anywhere in the source. The manifest marks version as dynamic, the build requires setuptools_scm, and the configuration writes the result into pandas_datareader/_version.py. The tags are therefore the only record of what shipped, which is why the shape of the release list is the real changelog here. The most recent commit recorded on the default branch is dated 2026-07-21, after both 0.11 releases.

One more bound is visible in the build requirements: setuptools_scm is held at >=8 and below 9, an unusual place to cap a build-time dependency while leaving the runtime requirements open.

## Devel docs are described as tracking master while the branch is main

The README points to a separate set of development documentation and says it covers the latest changes in master. The default branch of this repository is main.

The same document uses main elsewhere: the license badge links to blob/main/LICENSE.md. So master and main appear in one file, describing the same repository, with the branch name wrong in one of the two places.

The stable documentation is duplicated on purpose and the README admits it, saying a second copy of the stable documentation is hosted on read the docs for more details. Both hosts are named, which is more candour than most projects manage. The manifest agrees with the Read the Docs copy, listing it under Documentation, and its Homepage entry points somewhere else entirely, at the main pandas site rather than at anything specific to this package.

The index makes a third choice again, recording the stable documentation URL as the project homepage. So the project has three different notions of where its home is: the pandas site, the pydata.github.io stable docs, and the Read the Docs copy, with the README describing the last two as duplicates of one another.

## Three descriptions of the same package, and the broadest one is the index

Three descriptions, in three different places, with three different scopes.

The index description reads: extract data from a wide range of Internet sources into a pandas DataFrame. The README's own one-line summary reads: macroeconomic and factor-oriented remote data access for pandas. The manifest description reads: pandas-compatible data readers, formerly a component of pandas.

Only one of those promises breadth of sources, and it is the one written by whoever filled in the repository listing. The README is the narrowest and the most specific, and its next sentence backs that up by naming what the focus actually covers: macroeconomic, policy and factor-style sources such as FRED, Fama/French, Bank of Canada, World Bank, OECD and Eurostat, along with a pandas_datareader.macro interface described as new.

Six named sources, all of them official statistics or academic factor libraries. No equity or quote provider appears anywhere in that list, and the usage example is a single macro series, a call to pdr.get_data_fred with the ticker GS10. A reader arriving from the index description and hoping for market data will not find a documented provider for it.

The third description is the odd one out in a different way. Formerly a component of pandas is a lineage claim, not a scope claim, and it sits in the field a consumer is most likely to read first.

## The dev toolchain mixes two coverage services and reserves an import section for compat

Development and testing is listed as eight packages: black, coverage, codecov, coveralls, flake8, pytest, pytest-cov and wrapt.

Seven of those are recognisable tooling. The eighth, wrapt, is a wrapper library rather than a test or coverage tool, and it is the only entry in the list that is not obviously there to run tests or measure them. The list also contains two separate coverage services, codecov and coveralls, backed by a .codecov.yml and a .coveragerc at the repository root, so coverage reporting is wired up twice.

Styling is split three ways and the split is only half documented. black appears in the README list and carries a badge, and the isort configuration in the manifest sets profile="black", so formatting and import ordering are coordinated. flake8 is listed for linting. But isort is configured in the manifest and is not in the README's requirements list at all, so a reader provisioning from the README installs black and flake8 and gets import ordering wrong until they read the manifest.

The isort configuration also reserves a section for compatibility imports, setting known_compat to pandas_datareader.compat.* and listing FUTURE and COMPAT as their own import sections ahead of STDLIB and THIRDPARTY. That is direct evidence of a compat subpackage inside the package, and it is the same compatibility concern the distutils requirement was addressing from the outside.

## Conclusion

pandas-datareader suits an analyst who needs macroeconomic or factor series as a DataFrame and can live with a narrow source list. It does not suit someone reaching for equity quotes, since no quote provider is named among the documented sources. Check three things before adopting it. Which pandas you are on, because the installed floor comes from requirements.txt rather than from the README's list. Whether you need Python 3.12 or newer, where the standard library no longer carries distutils. And which license terms apply, because the index records none while the manifest declares BSD-3-Clause and both cannot be right without the file itself.

## FAQ

### what is pandas datareader

It is a Python library described as macroeconomic and factor-oriented remote data access for pandas, and the manifest adds that it is pandas-compatible and was formerly a component of pandas. Its releases are versioned below 1.0, currently 0.11.1.

### what does pandas datareader do

It provides readers for macroeconomic, policy and factor-style data sources, naming FRED, Fama/French, Bank of Canada, World Bank, OECD and Eurostat, plus a pandas_datareader.macro interface described as new. Results are returned as pandas DataFrames.

### how to install pandas datareader

The README gives pip install pandas-datareader, and a development install is python -m pip install git+https://github.com/pydata/pandas-datareader.git, or a clone followed by python -m pip install -e . The dependencies actually installed come from requirements.txt: lxml, pandas>=2.1.4, requests>=2.19.0 and setuptools.

## Sources

- [Issues](https://github.com/pydata/pandas-datareader/issues)
- [Project website](https://pydata.github.io/pandas-datareader/stable/index.html)
- [pydata/pandas-datareader on GitHub](https://github.com/pydata/pandas-datareader)
- [README](https://github.com/pydata/pandas-datareader/blob/main/README.md)
- [Releases](https://github.com/pydata/pandas-datareader/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pydata-pandas-datareader
