Model or dataset
yogeshhk/TeachingDataScience avatar
yogeshhk/TeachingDataScience

Every TeachingDataScience deck compiles twice, to slides and to notes

Open-sourced course notes for Artificial Intelligence and Data Science related topics, prepared in LaTeX

303 stars160 forksJupyter NotebookMIT

At a glance

What is it?
TeachingDataScience is a LaTeX repository of Beamer slide decks and cheat sheets for machine learning, deep learning, NLP, generative AI, maths for machine learning and Python, published at three depths from one hour to forty, plus a Code directory of runnable projects with conda environments and pytest suites. What is unusual about it is the dual output: one source file per topic produces both the talk and a two-column printable handout. What is unusual in the other direction is the References directory, which is deliberately not published.
Who is it for?
TeachingDataScience fits an instructor who needs a full syllabus's worth of material they are allowed to reuse, or a self-taught practitioner who wants a single topic as a one-hour sitting rather than a course commitment. It does not fit someone who wants a PDF without installing anything, because the compiled slides are deliberately not checked in and the build needs a LaTeX distribution that will prompt you for packages.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One topic file, two PDFs: a deck and a handout

The structural decision that distinguishes this repository from a folder of slide decks is that each topic compiles twice. Every seminar and workshop has two driver files built from the same source: a _Presentation.tex for Beamer slides and a _CheatSheet.tex for two-column printable notes. So the handout is not a separate artefact someone maintained by hand, and it cannot drift from the talk, because both are generated from one set of sources. That matters for the audience the project names, which includes anyone who wants to teach from it. The driver naming is fully specified: Main_ followed by Seminar, Workshop or Course, then the subject, then Presentation or CheatSheet, all with the .tex extension. The content hierarchy beneath that runs Course at forty hours, then Workshop at four to sixteen, then Seminar at one, then the raw topic files, which are named subject_maintopic_subtopic, with maths_linearalgebra_matrices given as the example. Nothing is pre-built: the compiled PDFs are not checked in, so you compile locally.

code
cd LaTeX
texify -cp Main_Seminar_ML_Intro_Presentation.tex

Three depths, and the widest coverage sits in the shortest ones

The tiers are given in hours rather than in vague levels, which makes them comparable. A seminar is a focused session of about one to two hours. A workshop is one topic in depth, about one to two days or eight to sixteen hours. A course is a full curriculum of about one to two weeks, roughly forty hours, and there are five of them: Machine Learning, Deep Learning, Generative AI, Maths for ML and Python. Then comes the counterintuitive part. Workshops and seminars span far more ground than the courses above, covering classical machine learning, linear algebra, large language models, retrieval augmented generation, LangGraph, graph neural networks and career advice among others. So the five courses are a spine rather than a ceiling. Each one is assembled from standalone workshops and seminars that you can also take on their own, which is what makes the repository usable at three different scales without maintaining three sets of material.

Five courses plus dozens of standalone sessions, indexed in COURSES.md

The catalogue entry names the five full courses and then immediately undercuts the impression of a tidy set of five. Beyond them sit dozens of standalone workshops and seminars in areas the courses do not cover, including natural language processing, large language models and generative AI, graph machine learning, reinforcement learning, software engineering, data analytics and career preparation. COURSES.md is the index, and its stated job is to carry the full catalogue with links to every driver, so the repository's navigation is one file rather than a directory tree you have to learn. Two details in that sentence matter for anyone automating anything over it. The courses are described as assembled from the standalone sessions, which means the course drivers are compositions over existing files rather than separate bodies of work. And because every driver is linked, you can resolve straight to the _Presentation or _CheatSheet file for a single topic without traversing the hierarchy.

Code/ is organised by area and each project has its own environment

The Code directory is where the repository stops being slideware, and the arrangement is worth reading because it is the part you can execute. Projects are grouped into nine areas. GenAI and agents covers langchain, langgraph, llamaindex, crewai, agents, agno and google-adk. Retrieval augmented generation applications covers chatbot-faqs, chatbot-multimodal, omni-rag, parsing and graphrag. Then fine-tuning, document parsing with docling and opendataloader, deep learning under pytorch, classical machine learning under ml, math and python, natural language processing under nlp, dnlp and spacy, graph neural networks under gnn, and an Indic language grouping of mahamarathi, sarvam and orgpedia. Two commitments make this different from a folder of notebooks. Each project carries its own environment.yml for conda, so the dependency set is per project rather than global. And for most projects there is a test_*.py suite runnable with pytest. A project you can test is a project you can check, which is a different standard from illustrated code.

References/ exists locally and is not uploaded, on purpose

This is the most honest paragraph in the repository and it deserves to be quoted rather than summarised. A References directory holding the papers, code and slides used as source material exists locally but is not uploaded, because most of it belongs to other people's repositories rather than being original work. So the provenance chain is deliberately incomplete in the published artefact, and the reason is a licensing one rather than an oversight. The disclaimer says the notes were built from a lot of publicly available material, that care has been taken to cite sources, and that some citations may be missing, with an invitation to point them out and have them fixed. It closes by saying not to depend on it fully, since there is always more to improve. Read together with the MIT licence and the 2019 copyright, the position is coherent: the notes are yours to reuse, the material underneath them is not, and the author says so in both directions.

Compiling needs a LaTeX distribution and will prompt you for packages

Here is the friction, stated plainly. This is a LaTeX repository, slides compile to PDF, and they are not checked in pre-built. To try one you change into the LaTeX directory and run texify against a driver file. The stated requirement is a LaTeX distribution, with MikTeX or TeX Live named as the two options, and the instruction is to install packages as prompted. That last clause is the honest cost: the first build of a new topic can stop several times while it pulls packages, and on a fresh machine it is not a thirty-second job. It is also a normal LaTeX experience rather than a defect, but it is the difference between this repository and something you can preview in a browser. What you get for it is a deck and a handout generated from the same source, which is a trade most people teaching from scratch will take, and most people browsing will not.

Copyright reads 2019, the branch is master, and there are no releases

Two facts frame the project's age. The copyright line reads 2019, and the last push to the repository was on 2026-09-27, so this is roughly seven years of accumulated material rather than a recent dump. And there are no GitHub releases at all, so nothing here is versioned: a workshop that is broken in a dependency today may be fine tomorrow, and there is no tag to pin a course to. The default branch is master, not main, which dates the repository's conventions. Alongside the LaTeX and Code directories the tree carries a set of governance and agent-facing files that say how the work is organised: AUTHORING.md, TODO.md, CONTRIBUTING.md, COURSES.md, SECURITY.md, CODE_OF_CONDUCT.md, an Admin directory, a CLAUDE.md and a .claude directory. The README frames the intent in three lines, a mission of spreading data science knowledge more widely, a vision of moving from automation to autonomy, and values of giving back and paying it forward.

Editorial conclusion

TeachingDataScience fits an instructor who needs a full syllabus's worth of material they are allowed to reuse, or a self-taught practitioner who wants a single topic as a one-hour sitting rather than a course commitment. It does not fit someone who wants a PDF without installing anything, because the compiled slides are deliberately not checked in and the build needs a LaTeX distribution that will prompt you for packages. Before you redistribute any of it, read the References note and the disclaimer, since the MIT licence covers the notes while the source material behind them belongs to other repositories, and if you plan to cite it, check the citations yourself rather than trusting that the sourcing has been audited.

Frequently asked questions

Can I teach myself data science?

The repository is arranged so that you can start at whatever depth suits you: seminars of about one to two hours, workshops of one topic in depth at about eight to sixteen hours, and five full courses of roughly forty hours. Every session is MIT licensed, and the courses are themselves assembled from sessions you can take standalone.

What does TeachingDataScience produce from one LaTeX file?

Two PDFs. Every seminar and workshop has a _Presentation.tex driver for Beamer slides and a _CheatSheet.tex driver for two-column printable notes, both generated from the same source files. The compiled PDFs are not checked in, so you build them yourself.

How do I compile a TeachingDataScience deck?

Change into the LaTeX directory and run texify against a driver file, for example Main_Seminar_ML_Intro_Presentation.tex. You need a LaTeX distribution, either MikTeX or TeX Live, and you will be prompted to install packages as the build requires them.

Does TeachingDataScience include runnable code?

Yes. The Code directory holds runnable Python and notebook projects grouped by area, including langchain, langgraph, llamaindex, crewai, several RAG applications, fine-tuning, document parsing, pytorch, classical machine learning, NLP and graph neural networks. Each project has its own conda environment.yml and most have a test suite runnable with pytest.

Can I reuse TeachingDataScience material?

The repository is MIT licensed with copyright 2019 Yogesh H Kulkarni. A References directory of source papers, code and slides exists locally but is not uploaded, because most of it belongs to other people's repositories, and the disclaimer notes that some citations may still be missing.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. yogeshhk/TeachingDataScience on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yogeshhk-teachingdatascience.svg)](https://hysenlabs.com/projects/yogeshhk-teachingdatascience)