# xlrd: the reader for legacy .xls files that stopped reading .xlsx

> A small, stable Python library that reads Excel files in the historical .xls format, plus the version split that broke a lot of working code, and what to reach for now that it reads nothing else.

**python-excel/xlrd** — Please use openpyxl where you can...

- Repository: https://github.com/python-excel/xlrd
- Website: http://www.python-excel.org/
- Stars: 2,206 · Forks: 436
- Language: Python
- License: NOASSERTION
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/python-excel-xlrd

## A library that deliberately narrowed its scope

The repository description is unusually blunt for a package with a PyPI badge: please use openpyxl where you can. That line is the whole project in miniature. xlrd is a library for reading data and formatting information from Excel files in the historical .xls format, and the README carries a warning that it will no longer read anything other than .xls files, pointing readers to python-excel.org for anything newer.

That narrowing is the interesting part of the project, not an accident. xlrd used to handle both .xls and .xlsx, which meant carrying a zip and XML parser alongside the old binary format reader. The maintainers removed the newer path, and the ecosystem filled the gap with openpyxl. A reader who finds that exception thrown at them, the familiar XLRDError about an Excel xlsx file not being supported, has hit the change rather than a broken install.

What is left is a focused library for one format. If that is the format you have, the surface is small and readable. If it is not, this is the wrong dependency and the error message will tell you so.

## What the reader deliberately drops from a spreadsheet

Beyond the format restriction, the README lists a set of things that are not supported but will be safely and reliably ignored. Charts, macros, pictures, any other embedded object including embedded worksheets, VBA modules, comments, hyperlinks, autofilters, advanced filters, pivot tables, conditional formatting and data validation.

Formulas are the entry worth thinking about: the formulas themselves are not read, but the results of formula calculations are extracted. In practice that means you get cached values, which is usually what a reporting script wants and occasionally not what it needs. A spreadsheet that recalculates on open in Excel may hand you different numbers to the ones stored in the file, and xlrd will report what was stored.

Password protected files are a hard stop. The README states they are not supported and cannot be read by this library, so an encrypted workbook is not a challenge you can work around with a flag.

Formatting information is part of what the library reads, which matters for the ignored list above. If you are extracting a number from a cell, the surrounding styling, column widths and merged regions are available, so a report generator can reproduce a layout rather than just dumping values.

## Installing and opening a workbook

The quick start in the README is a single pip command:

```bash
pip install xlrd
```

Opening a file and walking its sheets takes a handful of calls, and the README's example is short enough to read as the whole API surface:

```python
import xlrd
book = xlrd.open_workbook("myfile.xls")
print("The number of worksheets is {0}".format(book.nsheets))
print("Worksheet name(s): {0}".format(book.sheet_names()))
sh = book.sheet_by_index(0)
print("{0} {1} {2}".format(sh.name, sh.nrows, sh.ncols))
print("Cell D30 is {0}".format(sh.cell_value(rowx=29, colx=3)))
for rx in range(sh.nrows):
    print(sh.row(rx))
```

Two naming details matter more than they look. The cell access is rowx and colx, both zero indexed, so cell D30 is rowx=29 and colx=3. And sheet access is by index or name through methods such as sheet_by_index, rather than by a bracket on the book.

A command line view is included as well, which prints the first, second and last rows of every sheet in every file given to it:

```bash
python PYDIR/scripts/runxlrd.py 3rows *blah*.xls
```

The 3rows argument is the row count for the head and tail of each sheet, and the script is installed as a console entry point through setup.py, so it is available after installation rather than only from a source checkout.

## Python version support and the licensing ambiguity

The setup.py in the repository requires Python 2.7 or any later version, excluding 3.0 through 3.5, and its classifier list runs from Python 2.7 up to 3.9. That is a real constraint for anyone on a newer interpreter, and the README says nothing about it, so the packaging file is the place to check.

The license picture is less tidy, and it is worth naming plainly rather than picking a side. Three things are visible. The setup.py file declares license='BSD' and carries the OSI Approved BSD License classifier. The README points to python-excel.org and describes the project as open source. And GitHub reports no license type for the repository, which usually means the license file in the root was not detected or is not where the tooling expects it.

A LICENSE file is present in the repository tree. Read it directly if the license matters for your use, because the automated metadata and the packaging metadata do not agree on the first question, which is whether GitHub recognized one at all. For internal tooling the difference is academic. For code you intend to redistribute, it is the first thing to settle.

## Where xlrd sits against openpyxl and pandas

The honest comparison is not xlrd against other .xls readers, it is xlrd against the tool you are supposed to have moved to. openpyxl reads .xlsx, the format Excel has written by default since 2007, and it is what the README sends you toward for anything newer. pandas sits above both and has a read_excel function that dispatches on the extension, reaching for xlrd behind the scenes when it sees a .xls and openpyxl when it sees a .xlsx.

That dispatch detail explains a lot of the confusion in the wild. Code that worked for years because read_excel hid which library did the work will now raise on .xlsx input if the environment has a current xlrd, and the traceback points at xlrd even though the caller never named it.

The result is a split that is easy to state. Old binary spreadsheets go to xlrd. Everything else goes to openpyxl, directly or through pandas. The cost of keeping xlrd in a dependency list for a project that only ever sees .xlsx is a package that cannot do the job, and a version constraint on Python that comes along with it.

## What the project does not tell you

The README is accurate and short. It states the format restriction, the ignore list, the password limitation, an install command, a code example and a command line script. It points at http://www.python-excel.org/ for everything else, and the repository publishes no GitHub releases, so there is no changelog attached to a tag and no upgrade notes to read.

The repository itself is more informative than the README in one respect. The tree shows the tests directory, a .coveragerc, a .readthedocs.yml, a CircleCI directory and a scripts directory, which describes a project with an ordinary test and documentation setup rather than an abandoned one. The last push was on 2026-07-15, and the build badges in the README point at CircleCI and Codecov.

What the repository does not settle is why the .xlsx support was removed rather than maintained, or what the migration path is for a caller that needs both formats. The README treats the change as settled. If your workload genuinely spans both, the answer is to branch on the file extension in your own code and load two libraries, which is a small amount of work and a permanently larger dependency tree.

## Conclusion

xlrd remains a good choice for reading .xls, a format that still arrives from decades old systems and from a few enterprise exports nobody has converted. It is not a general Excel library any more, and code that used to open .xlsx through it needs openpyxl instead. Before you commit, check the license question for yourself: the README links to python-excel.org, the setup.py file declares BSD and a BSD classifier, and GitHub itself reports no detected license type. The last push was on 2026-07-15 and the repository publishes no GitHub releases, so version tracking happens on the package index rather than through release tags.

## FAQ

### What are the key differences between openpyxl and xlrd?

They read different file formats. xlrd reads the historical .xls format and, since the change in behaviour after version 2.0, refuses .xlsx files with an error. openpyxl reads .xlsx, the format Excel has written by default since 2007. The README points you to openpyxl for anything newer.

### How do I install xlrd for Python?

The README's quick start gives one command: pip install xlrd. The setup.py file declares support for Python 2.7 and later, excluding 3.0 through 3.5, with classifiers listed up to Python 3.9, so check that range against your interpreter.

### Why does xlrd raise an error when I open an xlsx file?

Because the library no longer reads that format. The README states it will no longer read anything other than .xls files and directs readers to python-excel.org for alternatives that handle newer formats, which means openpyxl for .xlsx.

### Can xlrd read password protected Excel files?

No. The README says password protected files are not supported and cannot be read by the library. The same list also names charts, macros, pictures, VBA modules, comments, hyperlinks, pivot tables and conditional formatting as unsupported but safely ignored, while formula results are still extracted.

## Sources

- [Issues](https://github.com/python-excel/xlrd/issues)
- [Project website](http://www.python-excel.org/)
- [python-excel/xlrd on GitHub](https://github.com/python-excel/xlrd)
- [README](https://github.com/python-excel/xlrd/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/python-excel-xlrd
