# Unsplash Datasets: 7.4M photos, keywords and search logs for research use

> The unsplash/datasets repository packages Unsplash photos with keywords and search queries into two downloadable tiers. The Lite tier is open to commercial use, the Full tier is not, and neither lets you redistribute the images.

**unsplash/datasets** — 🎁  7,400,000+ Unsplash images made available for research and machine learning

- Repository: https://github.com/unsplash/datasets
- Website: https://unsplash.com/data
- Stars: 2,792 · Forks: 142
- Language: Jupyter Notebook
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/unsplash-datasets

## What the Unsplash Dataset actually contains

This repository is not a library you import. It is a distribution point: a README, a DOCS.md field reference, a TERMS.md licence document, a CHANGELOG.md, and a how-to/ directory of loading examples. The payload lives behind unsplash.com/data, not in the Git history. The README describes two tiers. The Lite dataset holds about 25,000 photos, 30,000 keywords and 1 million searches, and the README says it can be used for both commercial and non-commercial purposes under the terms. The Full dataset holds 7.4M+ photos, 1.6M keywords and over 160M searches, and the README restricts it to non-commercial usage, with access granted by request. The stated source is over 415,000 contributing photographers and search data drawn from hundreds of millions of queries. That search-query layer is the part that separates this from a plain image dump: you get the vocabulary people typed alongside the photos they were shown, which is what makes it interesting for retrieval and semantics work rather than just classification.

## Two tiers, two licences, and the trap in the middle

The split is the whole design. Lite is a sample of the same schema, small enough to download without a conversation with anyone, and the README explicitly permits commercial use of it. Full is the same fields at roughly 300 times the photo count, and it is non-commercial only. That asymmetry is deliberate and it is easy to get wrong: a team can prototype happily on Lite, ship, then discover that scaling to Full would breach the terms. The second constraint is harder. The README states plainly that the dataset is made available for research purposes and that it cannot be used to redistribute the images contained within. So even the commercially usable Lite tier does not give you a photo library to embed in a product. If that is what you need, the README redirects you to the Unsplash API. Treat the dataset as training and analysis material, not as an image source.

## Getting the Lite dataset and loading it in Python

The Lite dataset is a direct download from unsplash.com/data/lite/latest, roughly 700MB compressed and about 1GB raw. There is no package to install and no CLI. The repository ships loading examples under how-to/, including a Python path and a PostgreSQL path. The Python example directory is the one to read first, because it shows how the tables are read from their on-disk format rather than leaving you to guess. Start by fetching the archive and unpacking it, then work through the notebook or script in how-to/python against the resulting directory. The README does not document a checksum or a rollback procedure for a partial download, so verify the archive size against the stated ~700MB before you spend time on it.

## Loading the same files into PostgreSQL

The second documented route is how-to/psql, which loads the dataset into a PostgreSQL database. This is the option that matters once you want to join photos to keywords and to search queries, because those relationships are exactly what a relational engine is good at and what a flat CSV read makes awkward. The repository does not publish a schema migration tool or a versioned DDL file in the documentation available; the how-to directory is the reference. If you are working at Full scale, expect the loading step to dominate your setup time, since the README puts the compressed Full dataset at around 20GB and the raw form at roughly 80GB. Plan disk and index build time accordingly, and note that the documentation does not describe an incremental update path from one release to the next.

## Where the dataset is the wrong tool

Three cases. First, any product that needs to serve the images: redistribution is prohibited, full stop, and the API is the sanctioned route. Second, anything requiring per-image licence certainty for commercial output, because the dataset terms govern the dataset, not the individual photographer's licence for a given photo. Third, work that needs current data. Releases are semantically versioned, and the CHANGELOG shows 1.3.0 in April 2025, 1.4.0 in June 2026 and 1.4.1 later that same month, so the cadence is real but not continuous. If your use case depends on what Unsplash looks like this week, a static release is the wrong shape. There is also a practical ceiling: the README does not document how to request an update, how long Full access approval takes, or what happens to your access if the terms change.

## How it compares with generic dataset hubs

The obvious alternative is a general dataset hub, and the difference is not size, it is provenance and structure. A hub typically gives you a file, a description and a licence tag, and you take the licence on trust. Here you get a curated corpus from a single source with a consistent schema across tiers, plus a search-query table that general hubs rarely include, all governed by one TERMS.md file you can read before downloading. The trade is flexibility. A hub lets you pick a dataset that matches your domain exactly; this gives you one domain, photography, at two sizes. If your task is generic image classification, that specificity is a limitation. If your task involves how people search for and describe images, the query layer is the reason to be here and a general hub will not replace it.

## Maintenance, versioning and licence cost

The last push to the repository was on 2026-06-26, and the most recent release, 1.4.1, carries the same timestamp. The README states that updates will add new fields and new images and that each release is semantically versioned, so a major bump signals a breaking schema change and a minor bump signals additions. Upgrading therefore means re-downloading and re-loading, not patching: nothing in the repository describes a delta or migration mechanism. Budget for that as a periodic full refresh rather than a maintenance chore. On licence, note that the repository does not declare a licence in the metadata available, and the governing document is TERMS.md. The README's own attribution instruction is to cite the dataset as Unsplash Lite Dataset 1.4.0 or Unsplash Full Dataset 1.4.0 and link to unsplash.com/data. Read TERMS.md yourself before commercial use; nothing here is legal advice.

## Conclusion

Adopt the Lite dataset if you need a legally usable photo corpus with keyword and search-query fields for a prototype, a retrieval experiment, or a teaching example. Do not adopt either tier if your product needs to display or redistribute the images themselves: the terms forbid redistribution, and the README points product use at the Unsplash API instead. Before you build anything, open DOCS.md and confirm which tables your version actually contains, since the README describes the dataset at a high level and the field-level detail lives in that separate file.

## FAQ

### What does the Unsplash Dataset contain?

It combines Unsplash photos with keywords and search queries. The README describes the Lite dataset as about 25,000 photos, 30,000 keywords and 1 million searches, and the Full dataset as 7.4M+ photos, 1.6M keywords and over 160M searches.

### How do I install the Unsplash Dataset in Python?

There is nothing to install from a package index. You download the Lite dataset from unsplash.com/data/lite/latest and follow the example in the repository's how-to/python directory to load it.

### How do I use the Unsplash Dataset in Python?

The repository provides a Python loading example under how-to/python, alongside a PostgreSQL example under how-to/psql. The README points to those directories as the supported way to load the data.

### Where can I download the Unsplash Dataset?

The Lite dataset downloads directly from unsplash.com/data/lite/latest at roughly 700MB compressed. The Full dataset requires requesting access at unsplash.com/data and is about 20GB compressed.

### How do I install the Unsplash Dataset?

There is no installer. The Lite dataset is a direct download from unsplash.com/data/lite/latest at roughly 700MB compressed, and the repository's how-to/ directory contains the loading examples for Python and PostgreSQL.

## Sources

- [Issues](https://github.com/unsplash/datasets/issues)
- [Project website](https://unsplash.com/data)
- [README](https://github.com/unsplash/datasets/blob/master/README.md)
- [Releases](https://github.com/unsplash/datasets/releases)
- [unsplash/datasets on GitHub](https://github.com/unsplash/datasets)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/unsplash-datasets
