FinMind: open financial datasets and a Python SDK for Taiwan markets
Open Data, more than 50 financial data. 提供超過 50 個金融資料(台股為主),每天更新 https://finmind.github.io/
At a glance
- What is it?
- An Apache-2.0 project that collects Taiwan and international market data daily and hands it to you through a pandas-shaped Python API.
- Who is it for?
- FinMind solves a specific and annoying problem: getting clean, dated, already-joined Taiwan market data without writing collectors for every exchange page. The `DataLoader` returns pandas frames, the `feature` helpers merge institutional and margin data onto a price frame, and the datasets update daily rather than needing a crawl of your own.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 17 days ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A dataset catalogue wrapped in a Python client
FinMind describes itself as open data covering more than 50 financial datasets, weighted toward the Taiwan market, updated daily. The repository is licensed Apache-2.0 and the package installs from PyPI:
pip install FinMindThe README separates what you get into five groups. Technical data covers Taiwan daily stock prices, real-time quotes, historical tick data, PER and PBR ratios, five-second order and trade statistics, the weighted index, and day-trading targets with volumes. Fundamental data is the statements: income, cash flow, balance sheet, dividend policy, ex-dividend results and monthly revenue.
What stands out is the breadth of the less glamorous tables. There is foreign ownership, shareholding distribution, margin financing and securities lending, the three major institutional investors' buy and sell activity, and per-broker daily futures and options trading. For anyone who has tried to assemble institutional investor flows in Taiwan, having that as a single documented endpoint is the actual product.
What the DataLoader call looks like in practice
The plotting example in the README shows the intended shape of use, and it is short enough to read as documentation:
# 取得股價
from FinMind.data import DataLoader
dl = DataLoader()
# 下載台股股價資料
stock_data = dl.taiwan_stock_daily(
stock_id='2609', start_date='2018-01-01', end_date='2021-06-26'
)
# 下載三大法人資料
stock_data = dl.feature.add_kline_institutional_investors(
stock_data
) The pattern is a constructor with no arguments and one method per dataset, keyed by stock ID and a date range. The `feature` namespace then merges the other tables onto the price frame you already have:
# 下載融資券資料
stock_data = dl.feature.add_kline_margin_purchase_short_sale(
stock_data
)
# 繪製k線圖
from FinMind import plotting
plotting.kline(stock_data)A candlestick chart comes out the other end. That last import is worth noting: `plotting` is part of the package and built on pyecharts, so the common case of looking at a downloaded series does not need a separate plotting choice.
Request limits and the token that doubles them
The README states a plain limit: 300 API requests per hour. Registering on the project site and verifying an email address yields a `token` parameter that raises the ceiling to 600 per hour. That is the whole access model, and it is the first thing to plan around if you intend to backfill several years of data, since a single stock's daily series is one call but a universe of a thousand stocks is a thousand calls.
There is also a maintenance window. The contact section notes that Sunday from midnight to 7am is reserved for maintenance and carries no service, which is worth knowing if a scheduled job happens to land there.
The repository links to a stress test page in the documentation, which is an unusual thing to publish and suggests the maintainers have thought about the limits of their own infrastructure. For a dataset that has been collected and normalised daily by someone else, paying attention to their stated request ceiling rather than discovering it through throttling is the sensible approach.
Dependencies that tell you what the SDK is built on
`pyproject.toml` is a good summary of the project's priorities. It requires Python 3.8 or newer and pins a bounded range for every dependency: pandas from 2.0, numpy, requests, pydantic 2, ta for technical indicators, lxml, pyecharts, ipython, loguru, aiohttp, tqdm and pyarrow. Every one uses a lower and an upper bound rather than a hard pin, which is the right call for a client library and shows the maintainers expect long-lived installs.
Python 3.14 support arrived in version 2.0.8 by relaxing the lxml upper bound, and 3.13 installability was addressed in the same run of releases. The classifier list runs from 3.8 through 3.14, which is a wide support surface for a package that started as a personal data collection script.
The dev group is where a decision shows its reasoning. The comment in `pyproject.toml` explains that the ranges exist because pytest 6.2.5 cannot run on Python 3.14, as it uses `ast.Str`, removed in that release, so the resolver needs room to choose a newer pytest on newer interpreters while keeping 3.8 installable. That is the kind of constraint other projects usually leave implicit and discover in CI.
Release notes that read like a changelog of real bugs
The recent releases are more informative than the feature list. Version 2.0.9 contains a cluster of backtest correctness fixes: stock dividends should not increase total cost of holding, OTC exchange traded funds were being charged securities tax three times over, transaction fees were missing from the buy affordability check, and the build was updated to replace deprecated Pydantic dict calls.
Those are not glamorous, and they are exactly what you want to see. A backtest engine that gets dividend or tax handling wrong produces plausible-looking equity curves that mean nothing, so each of these is a fix to trust the output rather than a feature. The same release also reworked CI so that pull requests from forks run the full test suite, with a gate job that makes branch protection actually enforce.
Version 2.0.10 added Taiwan futures k-bar data and a change worth calling out: `get_object` downloads now resume automatically with an HTTP Range request and verify integrity when a transfer is interrupted. For a client whose whole job is pulling multi-hundred-megabyte parquet files, that is a more meaningful improvement than it sounds.
Editorial conclusion
FinMind solves a specific and annoying problem: getting clean, dated, already-joined Taiwan market data without writing collectors for every exchange page. The `DataLoader` returns pandas frames, the `feature` helpers merge institutional and margin data onto a price frame, and the datasets update daily rather than needing a crawl of your own. The honest caveats are in the README itself: the licence section frames the data as educational and non-commercial, and the unauthenticated API is limited to 300 requests per hour with a token raising that to 600. For research and teaching it is hard to beat; for anything commercial, the licence language is the thing to read before you write a line of code.
Frequently asked questions
What is FinMind used for?
It provides open financial datasets for Taiwan and international markets, covering Taiwan stock prices, financial statements, institutional investor flows, margin data, futures and options detail, plus US stock prices, oil, gold, exchange rates and government bond yields. The datasets update daily, and a Python SDK returns them as dataframes.
How do I install FinMind and download stock data?
Install it from PyPI with pip install FinMind, then instantiate DataLoader from FinMind.data and call a dataset method such as taiwan_stock_daily with a stock ID and start and end dates. Feature helpers like add_kline_institutional_investors merge other tables onto the price dataframe you already have.
What are the FinMind API request limits?
The README states a limit of 300 requests per hour. Registering on the project site and verifying your email gives you a token parameter that raises the ceiling to 600 per hour. Service is also paused for maintenance on Sunday from midnight to 7am.
Can I use FinMind data for commercial purposes?
The repository is Apache-2.0, but the README carries a separate notice stating that the provided content is for educational and non-commercial use, that data is for reference only, and that users bear responsibility for trading losses. Read both before relying on it in a product.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/finmind-finmind)