# StatsBomb Open Data: Free Football JSON Data for Research and Analysis

> StatsBomb Open Data is a free repository of professional football match data released by StatsBomb for public research and genuine interest in football analytics. It provides event-level, lineup, and StatsBomb 360 tracking data in JSON format, organized by competition, season, and match ID.

**hudl/open-data** — Free football data from StatsBomb

- Repository: https://github.com/hudl/open-data
- Website: https://statsbomb.com/resource-centre/
- Stars: 3,657 · Forks: 985
- Language: Unknown
- License: NOASSERTION
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/hudl-open-data

## What StatsBomb Open Data Provides and Who Uses It

StatsBomb describes itself as committed to sharing data publicly 'to enhance understanding of the game of Football.' The README explicitly states that the goal is to 'actively encourage new research and analysis at all levels' and to 'extend the wider football analytics community and attract new talent to the industry.'

The data is released as a public service for research projects and genuine interest in football analytics. It is not a sample or a demo of a commercial product; it is actual StatsBomb event data for selected leagues and seasons, formatted as JSON files exported directly from the StatsBomb Data API. Analysts, academics, students, and hobbyists are the intended audience.

The repository is hosted under the hudl organization on GitHub, reflecting StatsBomb's parent company, but the content and terms are StatsBomb's own. The last push was on 2026-09-07, confirming the dataset is still being updated.

## Repository Layout and JSON Data Structure

The data follows a hierarchy based on competition and match identifiers. At the root of the data/ directory, competitions.json lists every competition and season included in the release. This is the starting point for any analysis: it maps competition IDs to names and season IDs to names.

Matches for each competition and season are stored in the matches/ directory. Each folder inside matches/ is named for a competition ID, and each file inside that folder is named for a season ID. Opening a match file gives you the metadata for each match in that season: teams, date, score, and the match ID that links to event and lineup files.

Events are stored in the events/ directory, with one JSON file per match named by the match ID. These files contain the detailed play-by-play event log. Lineups follow the same convention in the lineups/ directory. For matches where StatsBomb 360 data is available, the three-sixty/ directory holds those files, also named by match ID.

Documentation about the meaning of different events and the JSON schema lives in the doc/ directory.

## How to Access and Work with the Data

The data is static JSON files in a public GitHub repository. No API key or account is required. The simplest way to access it is to clone the repository or to download specific files directly from GitHub's web interface.

For programmatic work, the StatsBomb community has published open-source libraries in Python and R that parse these JSON files into data frames. The repository itself does not include these libraries, but the README encourages research use of the data, and StatsBomb's wider tooling ecosystem supports it.

The doc/ directory at the root contains documentation on the event schema, which explains what each event type means and how the fields are structured. For any analysis that involves event data, reading the documentation in doc/ is necessary because the event JSON uses numeric IDs and type codes that require the schema to interpret correctly.

For datasets of this size, memory usage can become a constraint when loading all events from a large season into a single data frame. The file-per-match organization allows selective loading by match ID, which is the practical approach for most analyses.

## StatsBomb 360 Data and What It Adds

StatsBomb 360 data is available for selected matches in the three-sixty/ directory. The README does not explain what 360 data contains in its brief introduction, but the format name refers to player tracking data that captures the positions of all visible players at each frame of play, not just the player who received or made an event.

Conventional event data records who did what (pass, shot, tackle) and where the ball was. StatsBomb 360 adds spatial context for all visible players at the moment of each event, enabling analyses of team shape, pressing intensity, and space creation that are not possible with event data alone.

The 360 data is available for selected matches only, not for the entire dataset. This reflects the additional cost of generating tracking data and suggests that coverage will expand over time as StatsBomb processes more matches. The three-sixty/ directory follows the same match-ID naming convention as events/ and lineups/, so linking 360 data to a specific event record requires joining on match ID and event ID.

## Attribution Requirements and Terms of Use

The README specifies a clear attribution requirement: anyone who publishes, shares, or distributes research, analysis, or insights based on this data must state the data source as StatsBomb and use StatsBomb's logo. The logo is available in their media pack, linked from the README.

This is not a standard open-source license. The license is stored as a PDF file (LICENSE.pdf) rather than a standard SPDX license identifier, and the GitHub API reports it as NOASSERTION rather than a recognized SPDX identifier. Teams or individuals who need to verify the exact legal terms for their use case should read the LICENSE.pdf directly.

The terms permit research use broadly, without a formal application process. The attribution requirement is the only condition stated in the README. Using the data internally for a non-published analysis is not addressed in the README, so individuals with stricter compliance requirements should consult the license document.

## Wyscout as a Comparison and Dataset Limitations

Wyscout is a commercial football analytics platform that provides event data, video, and statistical analysis tools for professional clubs and scouts. The fundamental difference from StatsBomb Open Data is access: Wyscout requires a paid subscription and targets professional use, while StatsBomb Open Data is free and publicly available for research.

The practical limitation of StatsBomb Open Data is breadth. The release covers 'certain leagues,' meaning the coverage is a curated selection rather than the complete StatsBomb commercial catalog. Analysts who need data from leagues or seasons not included in the open release cannot access them here. Commercial providers cover broader competition sets, including lower divisions and historical seasons not available publicly.

A second limitation is event schema stability. The JSON format is exported from the StatsBomb API, and the schema can change when StatsBomb updates its data model. The doc/ directory documents the current schema, but older analyses built against a different version may need updating when the schema changes.

The repository does not include pre-built analytical datasets, aggregated statistics, or derived features. All analysis starts from the raw JSON event files.

## Conclusion

Football analysts, data science students, and machine learning practitioners who need real event-level football data for research can use this repository freely. The primary constraint is attribution: publishing, sharing, or distributing any research based on the data requires crediting StatsBomb and using their logo from their media pack. Before building a project around this dataset, check competitions.json to verify the leagues and seasons it currently covers, since the available competitions are a curated selection rather than a comprehensive historical archive.

## FAQ

### How do I use StatsBomb open data?

Clone the repository or download files directly from GitHub. Start with competitions.json to find the competition and season IDs you need, then navigate to matches/, events/, and lineups/ using those IDs. The doc/ directory contains the event schema documentation required to interpret event JSON files.

### Does StatsBomb open data require attribution?

Yes. The README states that anyone who publishes, shares, or distributes research based on this data must credit StatsBomb as the source and use StatsBomb's logo, available from their media pack linked in the README.

### What is StatsBomb 360 data?

StatsBomb 360 data is stored in the three-sixty/ directory and is available for selected matches. It supplements standard event data with spatial context for visible players at each event moment. It follows the same match-ID file naming convention as events/ and lineups/.

## Sources

- [hudl/open-data on GitHub](https://github.com/hudl/open-data)
- [Issues](https://github.com/hudl/open-data/issues)
- [Project website](https://statsbomb.com/resource-centre/)
- [README](https://github.com/hudl/open-data/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hudl-open-data
