# Cookiecutter Data Science v2: What the ccds Command Actually Generates

> Cookiecutter Data Science is a project scaffold for data science work, and v2 ships a separate ccds command-line program instead of the plain cookiecutter utility. The layout is opinionated in ways that matter more than the folder names suggest.

**drivendataorg/cookiecutter-data-science** — A logical, reasonably standardized, but flexible project structure for doing and sharing data science work.

- Repository: https://github.com/drivendataorg/cookiecutter-data-science
- Website: https://cookiecutter-data-science.drivendata.org/
- Stars: 10,057 · Forks: 2,631
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/drivendataorg-cookiecutter-data-science

## The problem CCDS solves is naming, not modelling

Most data science repositories rot in the same way. Raw data ends up next to a cleaned copy, a notebook writes an output into the same folder it reads from, and nobody can say which CSV is canonical three months later. Cookiecutter Data Science addresses that by generating a directory tree with predefined roles for each location, so the argument about where a file belongs happens once, at project creation, rather than in every pull request. The README describes the project as a logical, reasonably standardized but flexible project structure for doing and sharing data science work, and the emphasis on sharing is the real target: a reviewer who has seen one CCDS project can orient themselves in the next one in a minute.

The audience is teams and individuals who start data science projects often enough that repeating the setup is a cost. It is not aimed at library authors, and it is not a modelling framework. Everything it does happens before you write analysis code. If you only ever work in notebooks and never hand a repository to someone else, the structure is overhead you will pay for and not use.

## How the ccds command and the template fit together

Cookiecutter Data Science v2 is a Python package named cookiecutter-data-science that installs a console script called ccds. Underneath, it extends the cookiecutter templating utility, which renders a directory of files with placeholders such as {{ cookiecutter.module_name }} substituted from answers you give interactively. The repository layout reflects that split: the ccds/ directory holds the CLI code, cookiecutter.json holds the questions and their defaults, and the {{ cookiecutter.repo_name }}/ directory is the template itself, including the data, notebooks, models, references and reports folders shown in the README.

The package metadata lists three runtime dependencies: click for the command-line interface, cookiecutter for rendering, and tomlkit. The version coupling is the part worth understanding. By default ccds uses the project template version that corresponds to the installed ccds package version, so installing version 2.3.0 of the package gives you the 2.3.0 template. The -c/--checkout flag overrides that and accepts a branch, tag or commit hash. That means a team can pin the scaffold deliberately, but it also means an unpinned ccds install silently changes what new projects look like when someone upgrades.

## Installing cookiecutter-data-science and generating a first project

The README requires Python 3.9 or newer and recommends pipx, on the grounds that this is a cross-project utility rather than a library your analysis code imports. The three documented install options are pipx, pip, and a conda command that the README marks as coming soon, so treat conda as unavailable for now.

```bash
pipx install cookiecutter-data-science
```

Running the console script with no arguments starts the interactive prompt. The README's example is the bare command, and the resulting tree depends on the settings you choose during that prompt.

```bash
ccds
```

After it finishes you should see a new directory containing LICENSE, Makefile, README.md, pyproject.toml, requirements.txt, setup.cfg, and the module folder with config.py, dataset.py, features.py, modeling/train.py, modeling/predict.py and plots.py. The Makefile is the entry point the README highlights, with convenience targets such as make data and make train. To generate against a different template revision, pass the checkout flag:

```bash
ccds -c master
```

If you need the pre-v2 layout, the README documents a -c v1 option that works with either the ccds program or the plain cookiecutter program, and requires one of the two packages to be installed. The v1 path is a compatibility route, not the default.

## Where the template's opinions get in the way

The data directory is split into raw, interim, processed and external, and the README calls raw the original, immutable data dump. Nothing enforces that. There is no hook documented in the README that blocks a write to data/raw, and the repository's hooks/ directory exists but its behaviour is not described in the README.

The second constraint is the notebook naming convention: a number, the creator's initials, and a hyphen-delimited description, with 1.0-jqp-initial-data-exploration given as the example. That is a filing system, and it works only if the ordering number is meaningful. In a repository where two people both start at 1.0, the convention collapses into a prefix that carries no information.

Finally, the template assumes Python. The generated pyproject.toml, setup.cfg for flake8, requirements.txt and the module layout are all Python-shaped. A team working primarily in R or Julia would be adopting the folder structure and discarding most of the generated files, which is a legitimate use but a much thinner one.

## Cookiecutter Data Science compared with plain cookiecutter

The obvious alternative is the cookiecutter utility on its own, pointed at the v1 template. The difference is not cosmetic. Plain cookiecutter renders a template directory and stops; the ccds program adds the version coupling between package and template, the checkout flag semantics described in the README, and the v1 fallback path. If you already use cookiecutter for other templates and want one tool rather than two, the v1 route is coherent, but you are choosing the older directory layout and giving up the v2 behaviour.

The other real alternative is writing your own template. Cookiecutter templates are just directories with placeholders and a cookiecutter.json, and a team with unusual requirements (a monorepo, a non-Python stack, a fixed cloud layout) will get more from a twenty-file template of their own than from bending this one. The trade-off is that a homegrown template has no upstream: the folder names, the Makefile targets and the pyproject.toml contents are yours to maintain, and there is no release history to diff against. CCDS is worth using when its opinions match yours closely enough that you would not change them.

## Maintenance, licensing and what an upgrade costs

The repository is not archived, and the last push was on 2026-08-07. The most recent release listed is v2.3.0 from 2025-07-24, preceded by v2.2.0 in March 2025 and v2.1.0 earlier the same month. The package metadata declares Development Status :: 5 - Production/Stable and supports Python 3.9 through 3.13.

The upgrade cost is asymmetric, and this is the part teams underestimate. Upgrading the ccds package changes the template that new projects receive, not the projects you already generated. Existing repositories keep their structure until someone regenerates or migrates them by hand, and the README does not document a migration path between template versions. Treat the generated project and the generator as two separate things to maintain.

The project is MIT licensed, with the licence file referenced from pyproject.toml and a LICENSE file generated into each new project if you choose one. MIT is permissive, but the licence you pick for your own generated project is a separate decision from the licence of the tool that generated it, and the template does not make that choice for you beyond offering the file.

## Conclusion

Adopt it if you want a fixed data/notebooks/models layout and a Makefile-driven entry point for a team that keeps re-creating the same folders by hand. Skip it if your work is a single notebook, a library with no data pipeline, or a codebase that already has a structure you are happy with, because the template's value is in agreeing on names before the first commit. Verify first that your Python is 3.9 or newer, that pipx is available if you want the recommended install, and which template version your installed ccds resolves to, since the package version and the template version are tied by default.

## FAQ

### What is cookiecutter in Python?

Cookiecutter is a templating utility that renders a directory of files with placeholders substituted from your answers to interactive prompts. Cookiecutter Data Science v2 is a Python package that extends it and ships a ccds command-line program, so the README tells you to use ccds rather than cookiecutter directly.

### How do I use cookiecutter data science?

Install the cookiecutter-data-science package, then run the ccds command with no arguments to start the interactive prompt. The README states that the resulting directory structure depends on the settings you choose, and that Makefile targets such as make data and make train are the convenience entry points.

### How do I install cookiecutter data science?

The README requires Python 3.9 or newer and recommends pipx install cookiecutter-data-science, with pip install cookiecutter-data-science as the alternative. A conda command is shown but marked as coming soon, so it is not available yet.

### What is cookiecutter data science?

It is a tool for setting up a data science project template that incorporates best practices, described in the README as a logical, reasonably standardized but flexible project structure for doing and sharing data science work. Version 2 requires installing the new cookiecutter-data-science Python package and using the ccds command instead of cookiecutter.

## Sources

- [drivendataorg/cookiecutter-data-science on GitHub](https://github.com/drivendataorg/cookiecutter-data-science)
- [License: MIT](https://github.com/drivendataorg/cookiecutter-data-science/blob/master/LICENSE)
- [Project website](https://cookiecutter-data-science.drivendata.org/)
- [README](https://github.com/drivendataorg/cookiecutter-data-science/blob/master/README.md)
- [Releases](https://github.com/drivendataorg/cookiecutter-data-science/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/drivendataorg-cookiecutter-data-science
