Open-source project
JerBouma/FinanceDatabase avatar
JerBouma/FinanceDatabase

FinanceDatabase: a free catalogue of what exists to trade

This is a database of 300.000+ symbols containing Equities, ETFs, Funds, Indices, Currencies, Cryptocurrencies and Money Markets.

9,287 stars961 forksPythonMIT

At a glance

What is it?
A community-maintained CSV and Python package listing over 300,000 equities, ETFs, funds, indices, currencies, cryptocurrencies and money markets, built for classification rather than for prices.
Who is it for?
FinanceDatabase answers a narrower question than most financial data projects answer, and that is its value: what instruments exist, where they are listed, and which sector and industry taxonomy they belong to. The README is explicit that fundamentals and price data are somebody else's job, so pair it with Finance Toolkit rather than expecting it to stand alone, and read the delisted flag before trusting any screen built on it.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 21, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem a catalogue of tickers actually solves

The README opens with an argument rather than a feature list. Well known instruments are easy to find, because everyone has heard of Microsoft or the S&P 500 tracker, but the long tail is not. A private investor looking for a specific sector in a specific country faces millions of candidates and no obvious way to narrow them. FinanceDatabase is built to close that gap, and it does so by concentrating entirely on taxonomy.

The table of contents in the README breaks the offering down by product type. Equities number 112,707 across 11 sectors, 80 industries, 117 countries and 84 exchanges. ETFs number 36,481 across 313 families and 51 exchanges. Funds number 57,853 across 1,540 families and 33 exchanges. Then come 2,556 currencies, 3,367 cryptocurrencies, 91,181 indices spread over 63 exchanges, and 1,367 money markets. That totals a little over 305,000 entries, consistent with the 300,000-plus claim in the package description.

What makes this useful is the grouping layer rather than the row count. A sector such as Information Technology, an industry group such as Software and Services, and an industry such as Software let you slice a market three ways, and ETF families and fund families give you a similar axis for passive products. For screening or sector analysis, having those classifications already resolved by someone is the whole point.

Installing the package and loading the equities table

Installation is a single pip command, and the README asks for an upgrade flag so that you pick up new data releases rather than a cached wheel:

bash
pip install financedatabase -U

Inside Python the package is imported under a short alias, and the README uses `fd` throughout its examples:

python
import financedatabase as fd

Each asset class is constructed separately, which matters for speed once the tables are large:

python
equities = fd.Equities()

equities.select()

The README draws attention to two details that are easy to miss. First, initialization of each asset class is only required once, so the object should be held in a variable and queried repeatedly rather than rebuilt. Second, the documented output is truncated to at most ten entries, and the summary column is dropped for readability. So the example in the README is a shape demonstration, not a result you should expect to reproduce in full.

The wider usage section shows the same pattern applied to ETFs, Funds, Indices, Currencies, Cryptocurrencies and Money Markets. It links to a getting-started notebook for the untruncated version, and the repository carries that notebook in `examples/` as an .ipynb file, so you can read real queries rather than guessing at the API surface.

Why the README says this database holds no fundamentals

This is the sentence that decides whether the project fits your use case. The README states that its aim is explicitly not to provide up-to-date fundamentals or stock data, because those are obtainable with ease once you have this database, and it points at Finance Toolkit as the tool for that step. FinanceDatabase supplies the address book; the toolkit is meant to fill it in.

The example output complicates that slightly. The equities sample shows columns for symbol, name, currency, sector, industry group, industry, exchange, market, country, state, city, zipcode, website, `market_cap`, isin, cusip, figi, composite_figi and shareclass_figi. So a market capitalisation bucket is present, labelled Large Cap or Micro Cap in the two visible rows. Two readings are possible: either it is a coarse static label maintained with the CSV rather than a live figure, or it is a figure that goes stale between releases. The README does not resolve this, and the practical move is to treat it as a category to filter on rather than a number to compute with.

The identifier columns are the more solid part. ISIN, CUSIP, FIGI, composite FIGI and shareclass FIGI are the fields you need to join this catalogue against another vendor's data, and the 2.4.0 release notes describe adding more than 15,000 new ISIN and CUSIP codes after re-evaluating existing mappings. Version 2.4.0 was published on 2026-06-02, and the last push to the repository was on 2026-09-20.

CSV edits as the contribution model

The README carries a call for contributors near the top and the reasoning is architectural. Rather than hosting an API or requiring code contributions, the project keeps its data in CSV files that can be edited by hand, which lowers the barrier to contributing for people who do not write Python. The README routes anyone interested to `CONTRIBUTING.md`.

The repository layout supports that story. A `database/` directory holds the data, a `compression/` directory suggests the shipped artifacts are compressed rather than raw, and `financedatabase/` holds the package that reads them. `tests/` and a `pyproject.toml` configured for pytest, ruff, black and codespell are the ordinary scaffolding around it, and a `uv.lock` indicates the maintainer develops with uv.

Two details in the 2.4.0 notes are worth knowing before you file a pull request. The equities, funds and ETFs datasets were split up by exchange before compression, explicitly so that changes are easier to review, which tells you the size and shape of a single diff. And a delisted flag was added, so the data now distinguishes companies that have left the market from companies that never traded. For any screen built on this catalogue, that flag is the difference between a survivorship-biased universe and a realistic one.

Packaging choices that shape what you can run it on

`pyproject.toml` is short and informative. The project requires Python `>=3.10, <3.16` and lists classifiers for 3.10 through 3.14. It is marked as Development Status 5, Production/Stable, the intended audience is Financial and Insurance Industry, and the license is MIT, which matches the LICENSE file in the repository root and means you can vendor the data into a commercial product as long as you keep the notice.

One structural choice deserves attention: the package declares a runtime dependency on `financetoolkit>=2.0.3,<3.0.0`. That is an unusual direction of dependency for a data catalogue, which you would expect to depend on pandas or nothing at all. It also means installing FinanceDatabase pulls in the toolkit, and a toolkit release that breaks the declared range will block your install or downgrade it. Worth checking before you pin versions in a research environment.

Both build targets exclude `financedatabase/validation`, listed in the wheel and sdist configuration alike. That path is therefore a development-only concern, not something you can import from an installed package. The build backend is hatchling, so there is no compiled extension to worry about and installation should behave the same on Windows, macOS and Linux.

Where FinanceDatabase stops and Finance Toolkit begins

The clean comparison is against the alternatives a reader would otherwise reach for. A commercial terminal such as a Bloomberg or FactSet setup gives richer fundamentals, intraday pricing and coverage guarantees, at a price that puts it out of reach for an individual. Open-source wrappers around public sources can be free but tend to inherit their upstream provider's rate limits and gaps. FinanceDatabase's position is deliberately narrower than both: it is the classification layer, and it is free and MIT licensed because the work being donated is taxonomy maintenance rather than data procurement.

The limits follow from that choice. There are no price series, no ratios and no corporate actions here, so anything resembling a valuation screen needs a second tool. Coverage is also uneven in the way volunteer-maintained lists tend to be: the exchange and country breakdowns are wide, but the depth of correction in any one corner depends on who has looked at it. Version 2.3.1 was a small fix tied to issue 108, which suggests the project does respond to specific reported problems.

One dependency risk deserves a sentence of its own. Because the catalogue depends on the toolkit and the toolkit is a separate project under the same maintainer, a divergence in either direction shows up as breakage here. If you need both, install them together and re-run the README's import and `equities.select()` example after any upgrade. That is the fastest way to find out whether a version bump moved data or moved API.

Editorial conclusion

FinanceDatabase answers a narrower question than most financial data projects answer, and that is its value: what instruments exist, where they are listed, and which sector and industry taxonomy they belong to. The README is explicit that fundamentals and price data are somebody else's job, so pair it with Finance Toolkit rather than expecting it to stand alone, and read the delisted flag before trusting any screen built on it. Install path is `pip install financedatabase -U` on Python 3.10 through 3.15, and the equity identifiers worth checking against your own vendor are the ISIN and CUSIP columns, since the 2.4.0 release added more than 15,000 of them and older mappings were re-evaluated.

Frequently asked questions

What does FinanceDatabase actually contain?

It contains classification data for over 300,000 symbols across equities, ETFs, funds, indices, currencies, cryptocurrencies and money markets. The fields are sector and industry taxonomies, exchange and country, address details, and identifiers such as ISIN, CUSIP and FIGI. Prices and fundamentals are deliberately excluded.

Is FinanceDatabase really free to use in commercial work?

The repository carries an MIT license, which permits commercial reuse including embedding the data in your own product, provided the copyright notice is retained. The README also asks contributors to help maintain the CSVs, and the data itself carries no separate license page beyond that LICENSE file.

How do I screen for delisted companies?

Version 2.4.0 introduced a delisted flag on company entries, along with a filter in the package that removes delisted companies automatically. It was released on 2026-06-02 and is part of the expanded metadata work, so older cached installs will not have it.

Official sources

  1. JerBouma/FinanceDatabase on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jerbouma-financedatabase.svg)](https://hysenlabs.com/projects/jerbouma-financedatabase)