Model or dataset
Jon-Becker/prediction-market-analysis avatar
Jon-Becker/prediction-market-analysis

The package description is still the template default

A framework for collecting and analyzing prediction market data, including the largest publicly available dataset of Polymarket and Kalshi market and trade data.

3,858 stars547 forksPythonMIT

At a glance

What is it?
prediction-market-analysis collects Polymarket and Kalshi market and trade data into Parquet, ships a 36GiB compressed dataset, and wraps everything in five make targets. Its manifest still carries the text a scaffolding tool writes, and one make target deletes the data directory after archiving it.
Who is it for?
This repository suits a researcher who wants both venues' trade history in one place and does not want to build the collectors, since the Polymarket blockchain index and the Kalshi client are the parts that take time. Three things to weigh before downloading it.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The description field says add your description here

The project manifest is named prediction-market-data while the repository is prediction-market-analysis, and its description field contains the sentence a project scaffolder writes when it creates a manifest and nobody has edited it since. The version is 0.1.0, and the declared Python floor is 3.9. The dependency list is where the manifest becomes informative again, because it names the two venue clients explicitly, a Kalshi package and a Polymarket package, alongside a general purpose web3 library for the chain side. So the tooling knows exactly which three interfaces it talks to, while the sentence a reader sees first on the package index says nothing at all.

Packaging the data deletes the data directory

One make target has a second effect that is easy to trip over. The README describes packaging as compressing the data directory for storage and distribution, and says the target creates a zstd-compressed tar archive named data.tar.zst and removes the data/ directory. Removing the extracted tree after archiving makes sense for someone about to publish the archive and reclaim the disk, and it does not for anyone who wants to keep working. The download target fetches the same 36GiB archive and extracts it, so a user who packages the data and then realises they need it again is looking at the same download they already performed once. Nothing warns about it, and the two targets sit next to each other in the README.

The blockchain indexer needs an endpoint the example leaves empty

The environment example has two lines and they tell you what the collector needs. One is a Polygon RPC URL variable, and its value is empty in the example. The other is a Polymarket start block number, set to 33605403. So the Polymarket trade history path is a blockchain scan from a fixed block height, and it cannot run until you supply an endpoint of your own, which is a cost line and a rate-limit line rather than a configuration detail. Kalshi has no equivalent variable in the file because it goes through an API client instead. The manifest also depends on a retry library, which is consistent with a collector that talks to a third-party endpoint in a loop and is expected to be interrupted.

A catch-all make rule succeeds without doing anything

The build file ends with a rule that matches any target and runs a shell no-op. The practical effect is that a mistyped target does not error. If you write make anlyze instead of make analyze, make finds the catch-all, runs nothing, prints nothing and exits successfully, so a script wrapping this repository will believe the step ran. The same file does have a useful pass-through target for the other case: it forwards any extra goals after filtering out the target name itself to the analysis command, which is how you run one named analysis without a dedicated target. That is a good design, sitting directly above a rule that hides every mistake in the build file's own vocabulary. The targets the README uses are thin wrappers around one command:

bash
make index
make analyze
make package

with a setup target for the two scripts that install tools and fetch the archive, and lint, format and test targets that call the linter, the formatter and the test runner directly instead.

One paper is listed twice with two identifiers

The research section runs to sixteen references, and one of them appears twice. The paper on the microstructure of wealth transfer in prediction markets is listed once against the author's own site and again against a preprint server, with a different abstract identifier in each case. That is defensible practice, since a reader who cannot reach one can reach the other, but it is not annotated, so a reader counting the bibliography counts one paper twice. Two other links in the list are shaped differently from the rest. One points at a direct PDF delivery path with tracking parameters rather than an abstract page, which is a per-request download link rather than a stable citation, and another carries a long numeric query parameter on a corporate document host that looks like a cache invalidation token. Both will outlive the session that produced them in a citation.

Parquet, DuckDB and three venue clients in one manifest

The storage story is consistent across the README and the manifest. The README says Parquet-based storage with automatic progress saving, and the manifest depends on a columnar library for Arrow plus an analytical database engine, which is the standard pairing for writing Parquet and then querying it without loading it into memory. The schemas are documented separately in a data schemas document for markets and trades. Alongside that sit the three client libraries: one for Kalshi, one for Polymarket, and a web3 library for the chain. So a collection run writes Parquet files through one library, and an analysis run queries them through another, which is why the repository ships schemas as a document rather than inferring them.

The plotting stack says what the analyses draw

Three of the dependencies describe the output of an analysis run without the README ever saying so. A library for skewed axis charts is present, a treemap layout library is present, and an image I/O library is present alongside the plotting library. The README says analyses produce PNG, PDF, CSV and JSON files in an output directory, and the dependency set explains why those formats: skewed axes for distributions like price or calibration curves where the interesting mass sits in a tail, treemaps for volume or share breakdowns where a pie chart fails, and an image library for the animated figures that a chart library cannot write on its own. It also explains the analytics layer: pandas, a numerical library and a progress bar library in the same list as the venue clients.

The dataset arrives through a script, not a link

There is no download link for the data anywhere in the README. Setup is a make target that runs two shell scripts, one to install tools and one to download, and the download script fetches an archive called data.tar.zst from an object store that the README identifies as Cloudflare R2. The host in the link is a domain of the author's own with an s3 prefix, which is the expected shape for an S3-compatible bucket rather than an AWS one, and it means the dataset is hosted on the same person's infrastructure as the analysis code. The archive is described as 36GiB compressed, the indexers write into the same data directory that the archive extracts into, and the project has no GitHub releases, so the dataset version is whatever the archive currently holds rather than anything tied to a commit.

Editorial conclusion

This repository suits a researcher who wants both venues' trade history in one place and does not want to build the collectors, since the Polymarket blockchain index and the Kalshi client are the parts that take time. Three things to weigh before downloading it. The size, at 36GiB compressed, with no partial or per-market download. The fact that the blockchain indexer needs a Polygon RPC endpoint that the example environment file leaves empty. And the fact that repackaging removes the extracted data, so a second archive means a second download unless you move the directory yourself.

Frequently asked questions

What does the prediction-market-analysis repository provide?

A framework for collecting and analysing prediction market data: pre-collected Polymarket and Kalshi datasets, indexers for gathering new data, and an analysis framework that generates figures and statistics. Currently supported are market metadata collection, trade history collection via API and blockchain, Parquet-based storage with automatic progress saving, and an extensible analysis script framework.

How large is the prediction-market-analysis dataset?

The README describes the pre-collected dataset as 36GiB compressed. It is fetched as data.tar.zst by the setup target and extracted into a data directory, and the packaging target recreates that archive with zstd and then removes the extracted directory.

What do I need to run the Polymarket data indexer?

The environment example lists a Polygon RPC URL variable with an empty value and a Polymarket start block number set to 33605403, so the blockchain path needs an endpoint you supply. Kalshi goes through its own Python client dependency instead, and the manifest also includes a retry library.

How are analyses run in prediction-market-analysis?

The analyze target opens an interactive menu to choose one analysis or run all of them, and writes PNG, PDF, CSV and JSON output to an output directory. A pass-through make target forwards a name to the analysis command, and the documentation includes a guide for writing custom scripts plus a document of the Parquet schemas for markets and trades.

What Python and tooling does prediction-market-analysis require?

Python 3.9 or newer, with dependencies installed and run through uv, and a lock file committed. Linting and formatting use ruff with a 120 character line length and a Python 3.9 target, and tests run under pytest with a marker for slow tests and a documented way to deselect them.

Official sources

  1. Issues
  2. Jon-Becker/prediction-market-analysis on GitHub
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jon-becker-prediction-market-analysis.svg)](https://hysenlabs.com/projects/jon-becker-prediction-market-analysis)