100knocks-preprocess: a 100-exercise SQL, Python and R drill on a fictional supermarket database
データサイエンス100本ノック(構造化データ加工編)
At a glance
- What is it?
- The Japan Data Scientist Society ships 100 structured-data exercises with a Dockerised PostgreSQL instance, notebooks and answer keys. The interesting part is the shared question set across three languages; the awkward part is that everything else has been quiet since v2.1 in 2022.
- Who is it for?
- Adopt it if you are teaching or learning structured data wrangling and want one question set that works in SQL, Python and R against the same fictional supermarket data, with an answer notebook for every exercise.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem 100knocks-preprocess solves, and for whom
Most tabular-data practice material is bound to one language. A pandas tutorial teaches pandas; a SQL workbook teaches SQL; an R course teaches R. The reader ends up fluent in one tool and unable to say how the same join, window or aggregation would be expressed in the other two. This repository attacks that specific gap. Its exercises are written once and answered three times, in SQL, Python and R, so the same question about a customer's purchase history can be attempted through a query, a dataframe and a tibble.
The README is explicit about the trade-off it accepts: some questions are not a natural fit for a given language, but the stated priority is learning how to express the operation in that language anyway. That is a deliberate design choice, not an oversight, and it shapes who the material suits. It is aimed at people who already know one of the three and want to move sideways, or at instructors who need a common exercise bank for a mixed cohort. It is not aimed at someone who wants a production data-processing library.
The data is fictional. The README states that anything resembling personal information is dummy data, and the repository ships a fictional supermarket's purchase records and customer records as CSV. That removes the usual privacy negotiation from classroom use, which is a large part of why the material travels well into universities and companies.
How the Docker Compose setup is wired: PostgreSQL plus a notebook server
The architecture is two services defined in docker-compose.yml. The first, named db, builds from dockerfiles/postgres/Dockerfile and runs as container dss-postgres. It exposes 5432, but only on the loopback address, because the port mapping is written as 127.0.0.1:5432:5432. Two bind mounts matter here: ./docker/db/init is mounted at /docker-entrypoint-initdb.d, which is how the database gets initialised on first start, and ./docker/work/data is mounted at /tmp/data so the CSVs are visible to the loading scripts.
The second service, notebook, builds from dockerfiles/notebook/Dockerfile and runs as dss-notebook. It maps 127.0.0.1:8888:8888 and mounts ./docker/doc, ./docker/work and a Jupyter Lab config file into the container. It declares a dependency on db with condition: service_healthy, so the notebook server waits for the database health check rather than racing it.
The two services use different credentials, and this trips people up. The db service sets POSTGRES_USER=postgres with POSTGRES_PASSWORD=postgres12345 and database dsdojo_db. The notebook service connects as PG_USER=padawan with PG_PASSWORD=padawan12345 to the same PG_DATABASE=dsdojo_db on PG_PORT=5432. If you connect to the database from outside and use the padawan credentials, that is why they fail.
One more detail worth noticing: the notebook command is start-notebook.sh --NotebookApp.token='', so the Jupyter server runs without a token. Combined with the loopback-only port binding, the exposure is limited to your own machine. It would still be a bad idea to rebind that port to 0.0.0.0 without adding authentication back.
Installing it and opening the first notebook
The README gives a three-command install. You need Docker Desktop, and the README notes that on Apple silicon you need Docker Desktop 4.4.2 or later, and that Windows Home Edition works once WSL2 is installed. Clone the repository, change into it, and bring the stack up:
git clone [email protected]:The-Japan-DataScientist-Society/100knocks-preprocess.git
cd 100knocks-preprocess
docker compose up -d --build --waitThe --wait flag makes the command return only once the containers report healthy, which matters because the database has to finish loading the CSVs before the notebooks are useful. When it returns, open http://localhost:8888 in a browser. That is the Jupyter Lab instance, and the exercises live under work, with answer notebooks under work/answer.
If the containers are created but the database cannot be reached, the README points at directory permissions as the most common cause, and notes that on macOS, placing the files outside your home directory requires extra Docker file-sharing configuration. There is also a doc directory with setup material, and the README suggests consulting it.
If you would rather not run Docker at all, the README offers two hosted routes for the Python track only: Amazon SageMaker Studio Lab and Colaboratory, each with a practice notebook and an answer notebook, linked from badge images. The SQL and R tracks are not offered that way. The repository also carries requirements.txt, environment.yml and install.R at the top level, which suggests the Python and R environments can be built outside the container, though the README does not document that path.
Where it stops being the right tool
The most concrete limitation is support. The README states plainly that the society does not handle individual questions about using the material, and that it accepts no liability for any problem arising from its use. For a self-directed learner that is fine. For a company rolling it into an onboarding programme, it means there is no escalation path, no issue triage commitment, and no compatibility guarantee for future Docker or PostgreSQL versions.
Release cadence reinforces this. The newest release listed is v2.1 from 2022-06-30, with v2.0 and v1.1 both from 2022-03-25. The repository has seen pushes since then, the most recent on 2026-09-12, but the exercise content itself has not had a tagged release in years. Treat the question set as stable rather than evolving.
The second limitation is scope. This is structured data processing. There is no machine learning component, no model training, no evaluation harness. If you arrive expecting the kind of content the name might suggest to someone browsing data science material generally, you will find joins, aggregations, window functions and reshaping instead. The README's own subtitle, structured data processing edition, is the honest description.
Third, the notebook server runs with no token and the database password is committed in docker-compose.yml in plain text. That is acceptable for a local teaching sandbox on a loopback interface. It is not acceptable to copy into anything reachable from a network, and the repository does not present it as a template for that.
How it compares with pandas-focused exercise books
The obvious alternative is a single-language exercise collection, of which there are many, typically built around pandas and scikit-learn and distributed as a book plus notebooks. The difference in approach is the axis of practice. A pandas exercise book optimises for depth in one library: you learn the idiomatic pandas way to reshape, merge and group, and you build fluency in that ecosystem's conventions.
100knocks-preprocess optimises for translation between languages. Because the question set is fixed and the answers are given three times, the comparison is forced. You cannot avoid noticing that a grouped aggregation you wrote in SQL as a GROUP BY becomes a dataframe groupby in Python and a different expression again in R. That is a genuinely different learning outcome, and for someone who moves between a warehouse and a notebook it is the more useful one.
The cost is that no single track gets the depth a dedicated book gives. The README acknowledges this directly, saying that some questions are not well suited to some languages but that the goal is to learn how to achieve the result in that language. If your goal is to become expert in pandas specifically, a pandas-first resource will take you further. If your goal is to stop being monolingual about tabular data, the shared question set is the point.
There is also a middle path the README itself provides: run the Python track in Colab or SageMaker Studio Lab with no local install. That removes the Docker requirement entirely, at the cost of losing the SQL and R tracks and the local PostgreSQL instance.
Licence terms and what they mean for reuse
The README splits the licence in two. Everything except one file follows the MIT licence. The exception is docker/doc/100knocks_guide.pdf, which carries the society's logo and other marks and is released under CC-BY-ND. That is the guide document, not the exercise notebooks or the data.
The practical consequence is that the no-derivatives term applies to the guide PDF specifically. If you are building course material and want to adapt the guide, that is the file to check carefully. The rest of the repository, including the exercises and the Docker setup, is MIT, which permits modification and redistribution with the licence notice preserved. This is a description of what the README states, not legal advice; if the distinction matters to your organisation, have someone read the actual licence files rather than this paragraph.
There is a second, softer condition. The README asks that organisations using the material state that they are using the Data Scientist Society's structured data processing edition. It is phrased as a request tied to permission rather than as a licence clause, but for a university or company it is a trivial thing to satisfy and worth doing.
The repository also carries .gitleaks.toml and .pre-commit-config.yaml. The contribution section explains that installing pre-commit, following the instructions at pre-commit.com, causes commits to be checked for credentials based on that configuration. That is a sensible guard for a repository that otherwise ships a hardcoded database password in its Compose file.
Editorial conclusion
Adopt it if you are teaching or learning structured data wrangling and want one question set that works in SQL, Python and R against the same fictional supermarket data, with an answer notebook for every exercise. Do not adopt it if you need a maintained library, a supported product, or anything with a security or upgrade commitment attached; the last push was on 2026-09-12, but the newest release, v2.1, dates from 2022-06-30, and the README states that the society does not handle individual questions and accepts no liability for problems arising from use. Before you commit a group to it, verify two things yourself: that the container starts on your hardware and that port 8888 is free, and that the CC-BY-ND term on docker/doc/100knocks_guide.pdf fits how you intend to redistribute the material, since the rest of the repository follows MIT.
Frequently asked questions
What do I need installed before running 100knocks-preprocess?
Docker Desktop, on Windows 10 or 11 or macOS. The README notes that Apple silicon Macs need Docker Desktop 4.4.2 or later, and that Windows Home Edition works if WSL2 is installed.
Which languages can I solve the 100 exercises in?
SQL, Python and R. The README states that the exercises are common across the three, and that while some questions are not well suited to a given language, the aim is to learn how to achieve the result in that language.
Can I use 100knocks-preprocess without Docker?
For the Python track, yes. The README links practice and answer notebooks for Amazon SageMaker Studio Lab and Colaboratory. The SQL and R tracks are not offered through those hosted routes.
Is the data in 100knocks-preprocess real customer data?
No. The README states that everything resembling personal information is dummy data, and the repository ships a fictional supermarket's purchase and customer records as CSV files.
Where does the 100knocks-preprocess database get its data on first start?
The db service mounts ./docker/db/init at /docker-entrypoint-initdb.d and ./docker/work/data at /tmp/data, so the initialisation scripts in the init directory load the CSVs into the dsdojo_db database when the container is first created.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/the-japan-datascientist-society-100knocks-preprocess)