Library / SDK
cuemacro/findatapy avatar
cuemacro/findatapy

findatapy: one Python interface over Bloomberg, Quandl, DukasCopy and Yahoo

Python library to download market data via Bloomberg, Eikon, Quandl, Yahoo etc.

2,131 stars222 forksPythonApache-2.0

At a glance

What is it?
findatapy wraps several market data vendors behind a single MarketDataRequest object, with an FX layer that will synthesise a cross rate from USD legs you did not ask for. Its setup.py lists seventeen install_requires with no version constraints, including one client library per vendor, so installing it to read Yahoo pulls in the Eikon SDK and the Alpha Vantage and Quandl clients too.
Who is it for?
Use findatapy if you pull data from more than one vendor and the switching cost between vendor SDKs is what you are trying to avoid, and if you are already committed to a conda or virtualenv environment where seventeen dependencies are not a problem.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 93 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Seventeen required dependencies, one per data vendor

The dependency list in setup.py is the first thing an evaluator should read, and it is not what you would expect from a library whose selling point is a unified interface.

python
      install_requires=["pandas",
                        "twython",
                        "pytz",
                        "requests",
                        "numpy",
                        "pandas_datareader",
                        "alpha_vantage",
                        "eikon",
                        "yfinance",
                        "quandl",
                        "chartpy",
                        "statsmodels",
                        "multiprocess",
                        "redis",
                        "numba",
                        "pyarrow",
                        "keyring",
                        "openpyxl"],

Seventeen packages, and not one of them has a version constraint. The first problem is the missing pins. A library that resolves pandas, numpy, pyarrow, numba and statsmodels to whatever is newest on install day will eventually install a combination its author never tested, and in a data library the failure will be a shape change or a dtype change rather than an ImportError, so it will surface in your results rather than your build.

The second problem is the structure. Several of those entries exist to serve exactly one data source. twython is for Twitter. alpha_vantage is the Alpha Vantage client. eikon is Refinitiv's Eikon SDK. quandl is the Quandl client. yfinance is the Yahoo client. So a developer who installs findatapy in order to read free Yahoo data gets an Eikon SDK, an Alpha Vantage client, a Quandl client and a Twitter library in their environment. Four of those require an account or a licence to be useful, and one of them is a commercial terminal SDK.

That is the opposite of how a multi-vendor abstraction should be packaged. The convention in this ecosystem is extras: install findatapy for the core, findatapy[bloomberg] for the terminal, findatapy[eikon] for Refinitiv. This project has exactly one extra, and it is the newest source:

python
      extras_require={
          "databento": ["databento"]
      },

databento is correctly optional. Everything else predates that decision. So the pattern the project now uses exists, and the migration to use it for the older vendors has not happened.

Three of the required packages are not market data libraries at all. chartpy is the author's own plotting package, numba is a JIT compiler, and keyring is a system credential store. keyring belongs here because of the API-key story below. numba is presumably there for the numeric hot paths. chartpy is a design choice, and a debatable one, which the next section comes back to.

None of this makes the library wrong. It makes the install heavier and the failure surface wider than the feature list suggests, which is exactly the gap a unified-interface library exists to close.

Four ways to set an API key, one of which is editing the source

The Installation section lists the ways to configure credentials, and the list is the most surprising thing in the README.

You can make sure you edit the dataconstants class for the correct Eikon API, Quandl API and Twitter API keys. Or you can run the set_api_keys.py script to set the API keys via storing in your keyring. Or you can create a datacred.py file which overwrites these keys. Or some of these API keys can be passed via MarketDataRequest on demand.

Four mechanisms for one problem, and they sit at four different levels of quality.

Editing the dataconstants class means editing a file inside an installed package. That is the mechanism that the project lists first. It also means your keys live in site-packages, they are lost on every reinstall, and a careless commit of the site-packages tree or a diff of the installed package leaks them. Nobody should do this in a repository, and the fact that it is listed first is the weakest thing about the library.

The set_api_keys.py route is the right one. It uses keyring, which is why keyring is a hard dependency, and that stores the credential in the operating system's credential store rather than in a file. That is the mechanism to use, and the reason keyring appears in install_requires is legitimate.

The datacred.py route is a local override file. It is the only mechanism that lets a developer keep a per-machine configuration without touching site-packages, and it is presumably expected to sit in .gitignore. The README does not say that, does not show the file, and does not name its location, so a developer who creates one has to work out where the library looks for it.

Passing keys on MarketDataRequest is the best of the four for automation, because the key travels with the request rather than living in process state. The word some matters: the README says some of these API keys can be passed this way, so which ones is undocumented.

The pattern here is a library that grew a credential system organically and never consolidated it. The keyring path is the intended future, the source-edit path is the historical one that should have been removed, and the datacred.py path is the current compromise for anyone who wants a per-machine file.

AUDJPY is computed from USD crosses, and that is not the same data

The headline example in the README contains a sentence that anyone using this for research needs to stop and read.

The text says there is functionality which is particularly useful for those downloading FX market data, and the example below it shows how to download AUDJPY data from Quandl and automatically calculates this via USD crosses.

That is the whole point and the whole hazard. AUDJPY is not necessarily a series that exists in any vendor's feed. What the library does is fetch the two legs it can get, AUDUSD and USDJPY, and divide one by the other to produce the series you asked for. The value you receive in the DataFrame is a derived number.

Derived cross rates differ from quoted ones in ways that matter. The bid and ask you would see on a real AUDJPY quote come from liquidity in that pair. A computed rate inherits the spreads of its two legs, so it is wider than the real thing, and the width varies with which legs are liquid at the time. Timestamps differ: the two legs are fetched independently and can be asynchronous, so the ratio at a given instant is between two slightly different instants. Rounding differs: you are dividing two rounded prices, and the error compounds. And a missing leg in one period produces a hole in a series that would otherwise have been complete.

None of that is a defect. Computing a cross from legs is standard practice and sometimes the only option, and a library that will do it for you is saving you a function. The defect would be not telling you, and the README does tell you, in a subordinate clause in the second paragraph of the file. That is the right information in the wrong place.

The second example shows the interface doing the opposite, which is the useful comparison. The same MarketDataRequest and the same fetch_market call, but with data_source set to dukascopy, freq set to tick, fields set to bid and ask, and a different ticker:

python
md_request = MarketDataRequest(start_date='14 Jun 2016', finish_date='15 Jun 2016',
                                   category='fx', fields=['bid', 'ask'], freq='tick', 
                                   data_source='dukascopy', tickers=['EURUSD'])

df = market.fetch_market(md_request)
print(df.tail(n=10))

Two days of tick data with both sides of the spread, from a broker feed rather than a reference feed. The README describes the change as with the same API calls and minimal changes in the code, and that is accurate: the shape of the request is identical and only the source, the frequency and the fields differ. What a user has to know, and what the README adds later in a single sentence, is that there might be certain limits of the history you can download for intraday data from certain sources, and that you will need to check with individual data providers. A backtest that silently gets six months instead of two years is the failure mode to watch for.

0.1.43 in setup.py, a stray period in a tag, and a coding log dated 2027

Three metadata problems in this repository, all small, all checkable, and together a fair summary of how the project is maintained.

First, the version. The three most recent releases are v.0.1.42 on 2026-03-20, v0.1.41 on 2026-01-02 and v0.1.40 on 2025-03-08. The setup.py says version="0.1.43". The Release Notes list in the README stops at 0.1.42. And the last push was on 2026-07-02. So the manifest is one version ahead of the newest tag, and that version was never cut. If you install from PyPI you get 0.1.42; if you install from git, as the README also documents, you get 0.1.43.

Second, the tag itself. v.0.1.42 has a period after the v, where v0.1.41 and v0.1.40 do not. Git will accept either, and version comparison code that strips a leading v will usually cope, but any tooling that matches tag names against a pattern will not, and a release automation script that derives the next version from the previous tag may produce something unexpected. It is a one-character slip, and it is the kind of slip that survives for years because nothing breaks.

Third, the coding log. The most recent entry is dated 02 Jul 2027, and its content is Added error messages around Parquet reading. The last push to the repository was on 2026-07-02. The same day, one year earlier.

So there is a commit logged to July 2027 in a repository whose most recent push is July 2026. Either the year is a typo, or a commit was made with a wrong system clock. Given the content matches the 2.1.7 and 2.0.0 work in the same log, a typo is the likely explanation. It matters more than a normal typo because a coding log with a future-dated entry is the kind of thing that makes a reader doubt the rest of the log, and this log is one of the few places the project records what it actually does between releases.

What the log does show is a steady, narrow maintenance pattern. Recent entries: switching s3 to pyarrow instead of s3fs so the library can use Python 3.14, adding a UTC default for timezone in get_file_properties in IOEngine, improving caching of Parquet files in IOEngine, speeding up ticker parsing in ConfigManager, and refactoring ConfigManager. Six changes, five of them in the internals, all of them the kind of thing that only matters to someone already using the library.

The Release Notes list runs from 0.1.12 in May 2020 to 0.1.42 in March 2026, which is thirty-one releases over six years, and the gaps are irregular. Three releases landed on 6, 7 and 7 October 2021, two on the same day in March 2025, and one in November 2022 followed by five months of silence. That is the shape of one person maintaining a library between other work, and it is worth knowing before you depend on a fix arriving.

Redis is a hard install dependency that the README calls in-memory

There is a whole README section devoted to an error message, and reading it tells you more about the design than the API documentation does.

The section is titled Couldn't push MarketDataRequest message. It says you might often get an error like the below when downloading market data, and you do not have Redis installed:

code
Couldn't push MarketDataRequest

Then the explanation. findatapy includes an in-memory caching mechanism, which uses Redis a key/value in-memory store. The idea is that if exactly the same data download call is made with the same parameters of a MarketDataRequest, it checks this volatile cache first, before going out to the external data provider, for example Quandl. If Redis is not installed, this caching will fail and you get the error, but all other functionality aside from the caching will be fine.

The description contains a category error that is worth naming because it will confuse anyone setting this up. Redis is not in-memory in the sense implied here. It is a separate server process holding its dataset in memory and, optionally, persisting to disk. findatapy is not starting a Redis instance and it is not holding anything in its own process memory. It is a client talking to a Redis you have to run. The README's own next paragraph makes the point and then contradicts it: Redis is usually set up as volatile cache, so once your computer is turned off, this cache will be lost. A server that loses its data on power off is a specific configuration choice, and calling that volatile and in-memory conflates the deployment with the storage medium.

The practical consequences are clearer than the terminology. redis is a hard entry in install_requires, so every install gets the client library whether or not it is used. At runtime the dependency is optional, which is the right call, because a caching failure should not stop you downloading data. And the failure is visible but unhelpful: the error says a message could not be pushed, not that a cache is unavailable, and it appears on the first cold request rather than at startup, so it reads like a transient network problem rather than a missing service.

The platform note is the other thing worth knowing. Redis is available for Linux, and there is an unsupported, older Windows version which the author says has been found to work fine, although it lacks some functionality of later Redis versions. So the caching path is Linux-first, degraded on Windows, and absent on macOS unless you run Redis yourself.

Whether you want the cache at all is a separate question. Caching by request signature is genuinely valuable for research, where you re-run the same fetch across a parameter sweep, and it is worthless for live work. What the section should have said, and does not, is how long entries live and what invalidates them.

setup.py, Apache 2.0 without a dash, and a LICENCE file

The packaging metadata has the three small problems you would expect from a project that started in 2016 and has not been revisited.

There is a setup.py and a MANIFEST.in at the root and no pyproject.toml. That is a legacy setuptools build, which means the project is not declaring a build backend through PEP 517 and a modern pip resolves it by falling back to the setuptools legacy path. It works, and it is the reason `pip install git+https://github.com/cuemacro/findatapy.git` is documented as a supported route: without a build-system table, pip has to guess. The build itself is straightforward, using setuptools setup and find_packages, with include_package_data=True and zip_safe=False.

The licence is the second problem. setup.py says license="Apache 2.0", which is not a valid SPDX licence expression. The identifier is Apache-2.0, without the space, and the modern form that tools expect is the SPDX string exactly. The file at the root is named LICENCE rather than LICENSE, using the British spelling, so an automated check that looks for LICENSE will not find it. And the README states Uses Apache 2.0 licence in prose rather than as a code span, so the same loose form appears in all three places.

None of this changes what you are legally permitted to do. The licence is Apache 2.0, the terms are the standard ones, and a LICENSE-equivalent file is present under a different name. What it does change is automation. A licence scanner that reads the metadata will see a string it cannot normalise, and a policy check that requires a recognised SPDX identifier will fail or need an exception. For a library whose author also publishes chartpy and finmarketpy and whose whole audience is institutional quant work, where licence scanning is routine, that is a small thing that produces a ticket.

The Python requirement is worth stating too. The README says Required: Python 3.10, and that Python 2 is not supported, with pandas and numpy etc. also required. A floor of 3.10 is reasonable and forward-looking, and the coding log entry about switching s3fs to pyarrow specifically so the library can use Python 3.14 shows the author testing against the far end of the supported range rather than the near end. That is more care than the dependency list deserves.

One last item belongs with the credentials rather than the packaging. blpapi, the Bloomberg Python API, is not on PyPI in a usable form, so the README gives a separate install command with Bloomberg's own index URL. That is correct and necessary, and it is also a reminder that the Bloomberg path is the one source in this library that will never be a plain pip install.

One author, four libraries, and chartpy is not optional

The README places findatapy in a family, and the family explains both the strengths and the dependency list.

The author previously wrote the open source PyThalesians financial library, and says this new findatapy has similar functionality to the market data part of that library, with the API totally rewritten to be cleaner and easier to use. It is also a fully standalone package, so it can be used with whatever libraries you have for analysing market data or doing backtesting, although the author recommends his own finmarketpy for backtesting of trading strategies. The Requirements section adds chartpy for interactive plots, linking the same author's chartpy repository. The contributors section asks for help on finmarketpy, findatapy and chartpy together, and points at PLANNED_FEATURES.md for areas where help is wanted.

So the ecosystem is PyThalesians, then findatapy, finmarketpy and chartpy, all one person, and findatapy is the data layer with the other three arranged around it. That is not a criticism. A coherent set of libraries from one author who understands the domain, sharing conventions and a data model, is often more useful than four unconnected packages. And chartpy being listed as a Recommended requirement rather than required is the correct classification, given it appears in install_requires anyway.

Which is the problem. chartpy is in install_requires, so you cannot install the data library without the plotting library, even though the README frames chartpy as a recommendation for people who want interactive plots. For a container with a slim base image, that is extra weight for a feature many users will never touch. It is the same packaging decision as the per-vendor clients, applied to a library the author owns, and it makes the same point: the boundary between what findatapy needs and what the author's other work needs has not been drawn in the manifest.

Two admissions in the README are worth repeating because they are honest and they should shape your expectations. findatapy is currently a highly experimental alpha project and isn't yet fully documented. And the Gallery section, which is the first thing a visual-library author would fill in, says To appear.

Six years, thirty-one releases, and two blank sections. The code is evidently being used, given the Parquet caching, the IOEngine and the ConfigManager refactors in the coding log. But the project's own assessment, which is the correct one to plan around, is that it is unfinished. If you need a market data layer that is documented, pinned and vendor-isolated, this is not it yet. If you need one vendor today and can live with a 0.1.x version, the abstraction is real and the examples work.

Editorial conclusion

Use findatapy if you pull data from more than one vendor and the switching cost between vendor SDKs is what you are trying to avoid, and if you are already committed to a conda or virtualenv environment where seventeen dependencies are not a problem. Do not adopt it in a clean production image without reading install_requires first, because Eikon, Quandl, Alpha Vantage and yfinance are required rather than optional, and a Bloomberg-adjacent SDK in your dependency list is a conversation with your security team. Do not trust a returned FX series to be quoted data, since the README's own example computes AUDJPY from USD crosses. Verify five things. Confirm the version you get, since setup.py says 0.1.43 while the newest tag is 0.1.42. Decide which of the four API-key mechanisms you will live with, because editing the dataconstants class and creating a datacred.py file are both options the project offers and neither is safe in version control. Check that Redis is available, or accept the Couldn't push MarketDataRequest error on every cold cache miss. Pin your dependencies yourself, because the manifest pins none. And read PLANNED_FEATURES.md, which the README points at as the place where help is wanted, since a project that tells you what it has not built yet is worth more than one that does not. The deciding fact is that the abstraction is genuinely useful and the packaging around it is unfinished, which is the project's own description of itself as a highly experimental alpha.

Frequently asked questions

How do I install findatapy and download market data?

Run pip install findatapy for the latest release, or pip install git+https://github.com/cuemacro/findatapy.git for the newest repository copy. Then construct Market(market_data_generator=MarketDataGenerator()), build a MarketDataRequest with start_date, category, data_source and tickers, and call market.fetch_market with it. Bloomberg's blpapi needs a separate install from Bloomberg's own index URL.

Does findatapy synthesise currency pairs that do not exist?

Yes. The README's own example downloads AUDJPY from Quandl and says the library automatically calculates it via USD crosses, which means the returned series is derived from the two USD legs rather than quoted. Derived crosses inherit the spreads, timestamps and rounding of their legs, which matters if you intend to backtest on them.

Why do I get a Couldn't push MarketDataRequest error?

That error appears when the Redis-backed cache cannot be reached, usually because Redis is not running. findatapy checks a cache for an identical previous MarketDataRequest before going out to the provider, and the caching failure does not stop the download itself. Redis is a hard entry in install_requires, though the README notes it is available for Linux with an older unsupported Windows version.

How do I set the API keys for Eikon, Quandl or Twitter?

The README lists four options: edit the dataconstants class in the installed package, run set_api_keys.py to store keys in the system keyring, create a datacred.py file that overwrites the stored keys, or pass some of them through MarketDataRequest on demand. The keyring route is the one to use, and keyring is a required dependency for that reason.

What version of findatapy should I install?

setup.py declares 0.1.43, while the newest release tag is 0.1.42, tagged 2026-03-20, and the README release notes stop there. Installing from PyPI gives 0.1.42 and installing from git gives the untagged 0.1.43. The last push was on 2026-07-02 and the project describes itself as a highly experimental alpha project that is not yet fully documented.

Official sources

  1. cuemacro/findatapy on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cuemacro-findatapy.svg)](https://hysenlabs.com/projects/cuemacro-findatapy)