sklearn==0.0 is in the requirements file, and the schedule samples are PDFs from 2016
The Hitchhiker's Guide to Data Science for Social Good
At a glance
- What is it?
- The open curriculum and manual for the Data Science for Social Good summer fellowship, published as a MkDocs site. The definition of the role it teaches has eleven categories and none of them is deployment, and the dependency file documents its own debt in a comment.
- Who is it for?
- Treat this as a curriculum rather than a software project, because that is what it is: a definition of the role, a manual, and a pile of notebooks, published as a static site. It is worth reading whether or not you apply to the fellowship, because the eleven categories are a defensible statement of what this work demands that most data science curricula omit.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 82 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A MkDocs site split three ways by audience
The top level is six entries: .github/, .gitignore, a pinned Python version file, the README, an mkdocs.yml and a requirements file, plus a sources/ directory holding everything else. So the published site is static, generated by MkDocs, and the material lives in a sources tree rather than in notebooks at the root.
The audience is split three ways, and the split is the navigation.
If you are applying or have been accepted, there is a manual that covers how to prepare before arriving, what orientation and training will cover, and what to expect from the summer. If you are learning at home, there are tutorials and teach-outs written by staff and fellows during the summer, with an invitation to suggest or contribute more. And the third route is the one for anyone who wants to do this work elsewhere: the stated goal is to encourage collaboration, and the file invites people who want to start their own program to build on what is here by using and contributing to these resources.
That third route is the reason the project is public at all, and it is the one that ages best. A curriculum written for one cohort of thirty people in one summer is reusable by anyone who teaches the same material, which is why the licensing question at the end of this matters more than the code question.
Eleven skills, and not one of them is deployment
The guide defines its subject as a data scientist for social good, described as an enigmatic breed combining one part data scientist, one part helper, one part educator and one part bleeding heart idealist, arrived at over many early mornings of trying to define the thing.
What follows is a list of eleven skill areas, and every one has a stated reason rather than a definition. Programming, because you need to tell the computer what to do. Computer science, to understand how your data is and should be structured. Math and stats, with the argument that everything else is applied math and that numerical results are meaningless without a measure of uncertainty, linked to a comic. Machine learning. Social science, for designing experiments in the field and knowing when correlation can suggest causation. Problem and project scoping, going from a vague description to a problem you can solve. Project management. Privacy and security, on the grounds that data is people. Ethics, fairness, bias and transparency. Communications. Social issues.
Read as a whole, the list is entirely about preparation and reasoning. There is no entry for deploying a model, monitoring it, keeping it running, or handing it to somebody else to operate. For a programme whose stated output is a project ready for field trial and implementation at the end of the summer, that is a real gap, and it is the kind of gap that a curriculum aimed at field trials rather than production papers over.
The weighting is also uneven in a way that looks deliberate. Ethics and privacy together get two of the eleven slots, and communications gets its own, which is not how a conventional machine learning curriculum is balanced.
The schedule material is PDFs from 2016 and 2022
The table of contents shows a hierarchy that starts at the manual, goes to a summer overview, and from there into a section on conduct, culture and communications. That last one covers the anti-harassment policy, the goals of the fellowship, what the programme hopes fellows get out of it, and the expectations of the fellows themselves.
Under the summer overview, two of the three items are documents rather than pages. One is a high level plan for the summer, a PDF that sets out the goals for each week. The others are sample orientation schedules, one from 2016 and one from 2022, both PDFs.
Two things follow. The first is that a PDF is not text: a reader searching the site for a session title, a week number, or a room will not find it in the search index, and neither will anyone editing the schedule. A schedule that lives in version control but not as text is a schedule that drifts, because the only cheap way to update it is to upload a new file.
The second is age. The repository was last pushed on 2026-07-16, and the two sample schedules it offers as templates are from 2016 and 2022. They are labelled as samples, which is honest, and they are also the only concrete picture of what a first two weeks look like that the guide offers. Treat them as a record of one programme's shape rather than as what to expect.
The anti-harassment section, by contrast, is a page, which is where policy belongs.
The requirements file pins a placeholder package and documents its own debt
The dependency file is two projects in one, and it says so.
The first half builds the site:
mkdocs==1.0.4
pymdown-extensions==6.0
mkdocs-material==4.2.0
jinja2==2.10.1
markupsafe==1.1.1
markdown==3.1.1
pygments==2.4.2
httpie==1.0.2The second half is a data science stack for the notebooks: Jupyter, IPython, pandas, numpy, scikit-learn, BeautifulSoup, requests, tqdm, matplotlib, statsmodels, seaborn, a Postgres driver, csvkit, nltk, ujson and spaCy. Between the two halves there is a comment saying the packages below are not needed for maintaining the guide itself and belong somewhere else.
So the file states its own problem and does not fix it. A reader installing the site gets the whole notebook stack as well, and anyone who wants only the site has to read the comment and edit the file.
One line in that stack is a real trap: scikit-learn is pinned through the package name sklearn at version 0.0. That name on PyPI is a placeholder that exists only to tell you it is the wrong one, so a clean install of this file fails on that line rather than on anything to do with the guide.
The other thing to notice is the vintage spread. The site half sits around 2020 releases, the notebook half around 2022, and the two are pinned to the same file with no explanation for the gap. Whether those combinations still install cleanly is not something the repository claims either way.
Twenty to forty fellows from about a thousand applicants
The programme description is specific about scale, which is useful context for anyone reading the curriculum.
It is a hands-on, project-based summer programme that launched in 2013 at the University of Chicago, has since expanded to multiple locations globally, and is now coordinated by the Data Science for Social Good Foundation together with Carnegie Mellon University. Fellows are typically graduate students, or senior undergraduates in some cases, drawn from computational and quantitative fields: computer science, statistics, mathematics, engineering, psychology, sociology, economics and public policy.
From a pool of typically around a thousand applicants, 20 to 40 fellows are selected. That ratio is the reason the guide is written the way it is: with thirty people and a dozen distinct disciplinary backgrounds, anything that assumes a shared vocabulary fails, so the material has to define its terms from the beginning.
The work itself happens in small cross-disciplinary teams on projects across education, health, energy, transportation, criminal justice, social services, economic development and international development, with government agencies and non-profits as partners. The stated mentorship model is full-time senior data science mentors plus project and partnership managers with industry or government experience, and the stated outputs are trained fellows, improved capacity at the partner organisation, and a project ready for field trial at the end of the summer.
Alongside the project work there are workshops, tutorials and ethics discussion groups drawn from the programme's own curriculum.
The licence is stated in the file and absent from the repository metadata
One line says all material is licensed under CC-BY 4.0. The repository metadata does not name a licence at all, and there is no licence file among the six top level entries.
That gap matters more here than it would in most repositories, because the whole point of the project is that other people should reuse it. The stated goal is to encourage collaboration and to let anyone who wants to start their own programme build on what is here, and CC-BY 4.0 is exactly the licence that permits that with attribution. But a machine reading the repository, or a person checking before copying a page, sees an unfilled field and no file, and the sentence in the README is the only signal either of them will find.
The dependency file makes the same point a second way. Pinning a specific version of every build and notebook package is an explicit choice about what may be redistributed and under what terms, and none of that is stated. A curriculum with a code component needs the two answers in the same place, and here only one is given.
Neither of these is a reason not to use the material. It is a reason to copy the attribution line with whatever you take, and to decide for yourself what the code is worth.
The stated priority is responsible work, stated before the content list
The first priority line says the job is to train fellows to do responsible data science, machine learning and AI for social good work, and then explains what that means in four clauses: problems with social impact, data science integrated with the social sciences, understanding and discussing the ethical implications, and privacy and confidentiality.
That sentence is the thesis, and the rest of the guide is an expansion of it. What makes it concrete is the following paragraph, which admits the vagueness of the term and then resolves it by splitting the role into eleven areas, each with a reason attached. The reasoning is informal on purpose. The framing is one part data scientist, one part helper, one part educator and one part bleeding heart idealist, arrived at over early mornings of trying to define it.
Two details are worth carrying away. The maths entry is justified by the point that numerical results are meaningless without some measure of uncertainty, which is a stronger claim than most curricula make. And the privacy entry is justified in four words: data is people.
The last paragraph closes the loop by inviting contributions, and the last thing in the project file is the licence line. A curriculum that ends on attribution rather than on a call to action is telling you what it thinks its job is.
Editorial conclusion
Treat this as a curriculum rather than a software project, because that is what it is: a definition of the role, a manual, and a pile of notebooks, published as a static site. It is worth reading whether or not you apply to the fellowship, because the eleven categories are a defensible statement of what this work demands that most data science curricula omit. Two practical notes. The dependency file is not a build you should copy: it pins a placeholder package, mixes 2020 and 2022 era pins, and its own comment says half of it does not belong there. And the schedule material in the manual is frozen PDFs from 2016 and 2022, so treat the process described there as historical rather than current.
Frequently asked questions
What is the Hitchhiker's Guide to Data Science for Social Good?
The open curriculum, manual and tutorials for the Data Science for Social Good summer fellowship, a programme launched in 2013 at the University of Chicago and now coordinated by the Data Science for Social Good Foundation and Carnegie Mellon University. It is published as a static MkDocs site with the material under a sources directory.
Who is the DSSG Hitchhiker's Guide written for?
Fellows first and everyone else second. Applicants and accepted fellows are sent to the manual for preparation and orientation, people learning at home are sent to the tutorials and teach-outs, and anyone who wants to start their own programme is invited to use and contribute to the resources.
What skills does a data scientist for social good need?
Eleven areas, each with a stated reason: programming, computer science, math and stats, machine learning, social science, problem and project scoping, project management, privacy and security, ethics, fairness, bias and transparency, communications, and social issues. Deployment and operations are not on the list.
What licence is the DSSG Hitchhiker's Guide under?
The project file states that all material is licensed under CC-BY 4.0. The repository metadata names no licence, and there is no licence file among the top level entries, which matters for a project whose stated purpose is reuse.
How is the DSSG Hitchhiker's Guide built and pinned?
As a MkDocs site with an mkdocs.yml and content under sources/, pinned to MkDocs 1.0.4 with the Material theme 4.2.0. The same requirements file also carries a Jupyter and data science stack, including a scikit-learn placeholder package pinned as sklearn at 0.0, under a comment saying those packages are not needed to maintain the guide.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dssg-hitchhikers-guide)