Model or dataset
mlabonne/llm-datasets avatar
mlabonne/llm-datasets

mlabonne/llm-datasets: a curated index of post-training data, and what it leaves out

Curated list of datasets and tools for post-training.

4,795 stars400 forksUnknownLicense varies

At a glance

What is it?
The repository is a README-only catalogue of instruction, math, science and code datasets for SFT and reasoning work, with licence notes attached to individual entries. It is useful as a shortlist generator, not as a data pipeline.
Who is it for?
Use mlabonne/llm-datasets when you need a shortlist of candidate post-training datasets and want the licence caveats visible next to each entry, particularly if you are choosing between instruction mixtures for an SFT run.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 153 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What mlabonne/llm-datasets actually is, and who it is for

This is a list, not a library. The top level of the repository holds CITATION.cff and README.md, and the README describes itself as a "Curated list of datasets and tools for post-training." There is no package to install, no module to import, no CLI. Everything the project offers lives in tables inside the README.

The audience is narrow and specific: people assembling a supervised fine-tuning or reasoning-training mixture who already know what SFT means and need to pick sources. The README opens by arguing that a good dataset targets accuracy, diversity and complexity, and it suggests combining manual review, rule-based filtering and scoring via judge LLMs or reward models. That framing tells you who the author expects: someone running a post-training job, not someone exploring what an LLM is.

If you are looking for a way to download data programmatically, this repository will not help. If you are looking for a place to start comparing candidate datasets before you write your own download script, it will.

How the catalogue is organised, and the metadata each entry carries

The README splits post-training data into categories. Instruction datasets come first, covering general-purpose mixtures, then math, science and code. Each category gets a short paragraph explaining why the domain is hard, followed by a markdown table.

The columns are where the real information sits. Every row has a dataset name linked to its Hugging Face page, a date, a sample count, a Thinking column, and a Notes column. The Thinking column is a three-way flag (Yes, No, or Mixed) indicating whether the samples contain reasoning traces. That single column is the most useful filter in the document, because it separates data you would use for a standard instruction-tuned assistant from data you would use for a reasoning model. Nemotron-Cascade-2-SFT-Data is marked Mixed; SYNTHETIC-2-SFT-verified is marked Yes; open-perfectblend is marked No.

The Notes column carries scale, provenance and licence. It records which teacher models generated the responses, which paper or release post describes the dataset, and which licence applies when it is not permissive. A note near the top states that unless specified otherwise, all listed datasets are under permissive licences such as Apache 2.0, MIT or CC-BY-4.0. The exceptions are flagged inline: Dolci-Instruct-SFT is marked CC-BY-NC-4.0, MegaScience is marked CC-BY-NC-SA-4.0, and Nemotron-Cascade-2-SFT-Data carries the NVIDIA Open Model License. That inline flagging is the design decision that makes the table worth reading rather than skimming.

Using the list: no install, just a clone and a read

There are no install steps in the README because there is nothing to install. The repository is documentation. The practical workflow is to clone it and read the tables, or to open the rendered README on the project's page.

If you want the file locally, a clone gets you the README and the citation file:

bash
git clone https://github.com/mlabonne/llm-datasets.git
cd llm-datasets

After that, the only file worth opening is README.md. There is no build step, no dependency file, and no test suite. What you should see is a single markdown document with the category tables described above.

From there, the first real use is picking a candidate and going to its dataset card. The README links each dataset name directly to its Hugging Face page, so the handoff is one click. For example, the general-purpose table links to the Nemotron-Cascade-2-SFT-Data page, and the math table links to NuminaMath-CoT. The repository does not show you how to load either one; the dataset card does. Treat the README as the index and the dataset card as the manual.

Where the catalogue stops being enough

The list tells you what exists. It does not tell you what is inside. Sample counts and Thinking flags are coarse, and the README does not document per-dataset schema, chat template format, token counts, or deduplication status. Two datasets both marked Yes can differ entirely in how reasoning traces are formatted, and the README will not warn you. You find that out after downloading.

The licence situation is also less tidy than the blanket note suggests. The note says entries are permissive "unless specified otherwise," which puts the burden on you to scan every Notes cell. Non-commercial and share-alike terms do appear, including CC-BY-NC-4.0 on Dolci-Instruct-SFT and CC-BY-SA-4.0 on Nemotron-Math-Proofs-v1. For a commercial fine-tune, that distinction matters more than any other column in the table, and it is buried in prose rather than given its own column.

Finally, there is no freshness guarantee. The README is updated by hand, and the repository gives no release history and no changelog. If a dataset card changes its licence after the row was written, the row will not know. The repository is a snapshot of someone's reading, not a synchronised mirror.

Alternatives, and the difference in approach

The most direct alternative is Hugging Face's own dataset search and filtering. That is a live index with facets for size, task, language and licence, and it covers everything the README lists plus a great deal it does not. The difference is editorial: the README is a small, opinionated shortlist with a Thinking column and hand-written provenance notes, while Hugging Face search returns everything matching your filters with no judgement attached. If you already know what you are looking for, search wins. If you are trying to work out which ten datasets are worth a look, the curated list saves time.

A second comparison point is the datasets the README itself points to as sources of tooling rather than data. The entry for orca-agentinstruct-1M-v1 notes it is a subset "designed for Orca-3-Mistral, using raw text publicly available on the web as seed data." That is a different model of dataset construction: generate from web seed text rather than aggregate existing corpora. If your goal is a bespoke mixture rather than a selection, the README's own framing about combining manual review, rule-based filtering and judge scoring describes the work you would do, and no entry in the list does it for you.

Maintenance, licence posture and what to verify

The repository is not archived, and the last push was on 2026-04-29. There are no retrieved releases, which fits a project whose only artefacts are a README and a citation file. Upgrade cost is close to zero in the software sense: there is nothing to upgrade, and pulling the latest README is the whole maintenance story. The real cost is re-verification, because every row is a claim about someone else's dataset.

The licence implications sit with the datasets, not with this repository. The README's blanket permissive note does not override the individual dataset cards, and it explicitly carves out exceptions. If you plan to ship a model trained on any of these mixtures, the licence of the training data is your problem to check, and the README is only a starting pointer. It does not give legal guidance and should not be read as doing so.

What to verify first: open the linked dataset card for your top candidate, confirm the licence text matches the Notes cell, and confirm the sample count. Then check whether the Thinking flag matches the actual record format, because that determines whether the data fits your training setup at all.

Editorial conclusion

Use mlabonne/llm-datasets when you need a shortlist of candidate post-training datasets and want the licence caveats visible next to each entry, particularly if you are choosing between instruction mixtures for an SFT run. Do not use it if you need a loader, a filtering pipeline, or a guarantee that every licence claim is current: the repository contains CITATION.cff and README.md and nothing executable, and several entries carry non-permissive terms such as CC-BY-NC-4.0 and CC-BY-SA-4.0. Before training on anything listed, open the linked Hugging Face dataset card and confirm the licence and the sample count yourself, because the table is a pointer, not a source of record.

Frequently asked questions

What datasets are used for LLMs?

mlabonne/llm-datasets groups post-training datasets into instruction, math, science and code categories, with general-purpose mixtures like Nemotron-Cascade-2-SFT-Data, SYNTHETIC-2-SFT-verified and open-perfectblend listed first. Each entry links to its Hugging Face page and notes scale, provenance and licence.

What is an LLM in data?

The repository does not define the term. Its scope is post-training data: the README describes SFT as the step that turns a pre-trained model into an assistant that follows instructions, and the listed datasets are the training material for that step.

Can ChatGPT generate datasets?

The README does not discuss ChatGPT. It does describe datasets generated by other models: Nemotron-Cascade-2-SFT-Data notes responses generated from DeepSeek-V3.2, GPT-OSS-120B and Qwen3, and KIMI-K2.5-1000000x is described as reasoning traces distilled from Kimi K2.5.

Is ChatGPT LLM or NLP?

The README does not address this. mlabonne/llm-datasets is a catalogue of post-training datasets and does not cover ChatGPT, NLP terminology, or model classification.

Official sources

  1. Issues
  2. mlabonne/llm-datasets on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mlabonne-llm-datasets.svg)](https://hysenlabs.com/projects/mlabonne-llm-datasets)