data-science-interviews: Community Q&A for Data Science Technical Interviews
Data science interview questions and answers
At a glance
- What is it?
- data-science-interviews is a community-maintained GitHub repository that collects data science interview questions with community-written answers, covering machine learning theory, SQL, Python, and probability. It is aimed at data science candidates preparing for technical interviews and at practitioners who want to contribute answers.
- Who is it for?
- data-science-interviews is the right starting point for candidates preparing for data science technical interviews who want a free, community-reviewed resource covering linear models, trees, neural networks, SQL, and Python coding. It is not a structured course; there is no guaranteed answer quality on every question, and the repository leaves gaps that contributors have not yet filled.
- Can I use it commercially?
- Yes, with credit. CC-BY-4.0 allows commercial use as long as you credit the authors and indicate what you changed. It is written for creative content, so check how it applies to any code.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the Repository Solves and Who It Is For
Data science technical interviews typically cover three broad areas: machine learning theory (explaining how algorithms work, their assumptions, and failure modes), technical skills (SQL queries, Python code, and data manipulation), and statistical reasoning (probability distributions, hypothesis testing, and experiment design). Candidates who prepare without a structured resource often discover gaps only in the interview room.
data-science-interviews addresses this by collecting interview questions in each area and pairing them with answers written and reviewed by the community. The primary audience is data science candidates preparing for technical screening rounds at technology companies. Practitioners who have been through interviews and know what questions appear in practice are the intended contributors.
The repository was created by Alexey Grigorev, who is also associated with DataTalks.Club. The project operates on a pull-request model: anyone who knows the answer to an unanswered question, or who can improve an existing answer, is invited to contribute a PR.
Repository Layout: Theory, Technical, Contributed, and Awesome
The repository is organized into a small number of Markdown files, each corresponding to a question category.
theory.md covers theoretical machine learning questions. According to the README, this file includes questions on linear models, trees, neural networks, and other topics. These are the explanatory questions: why does regularization reduce overfitting, what is the bias-variance tradeoff, how does gradient boosting work.
technical.md covers the practical skills tested in coding screens: SQL queries, Python programming, and data manipulation. These are the questions where candidates write or explain code during the interview.
The contrib/ folder holds contributed question sets. The README lists prob as one example: contrib/probability.md. The folder is open for community additions.
awesome.md aggregates links to other data science interview resources. It is a companion reference rather than a Q&A file.
contributors.md lists everyone who has contributed an answer or question. The README notes that the project is a joint effort of many people and links to that list.
The site is built with Jekyll (the repository includes a Gemfile, _config.yml, and _layouts/ directory) and is published at https://alexeygrigorev.com/data-science-interviews/. The README does not document how to build the site locally.
How to Use the Repository as a Study Resource
The most direct way to use data-science-interviews is to clone the repository and read the Markdown files locally, or to browse the published site at https://alexeygrigorev.com/data-science-interviews/. There are no install commands and no software to run; the content is plain text.
A practical study approach is to work through theory.md for conceptual questions first, then move to technical.md for coding practice. Questions without answers are gaps that the community has not yet filled; those gaps are also the best places to identify which topics need deeper research.
For candidates who want to contribute, the README gives a clear workflow: if you know how to answer an unanswered question, create a pull request with the answer. If an existing answer can be improved, create a pull request with the improvement. If you spot an error, create a pull request with a fix. The project uses the GitHub pull request model for all changes.
The awesome.md file is worth reading alongside the question files because it points to external books, courses, and practice sets that complement the Q&A pairs. The README does not specify what those resources are, as the file content is not reproduced in the repository description.
What the Repository Covers and Where the Gaps Are
The confirmed coverage areas are: linear models, decision trees, neural networks, and other machine learning algorithms (theory.md); SQL, Python, and coding questions (technical.md); and probability questions in the contributed section.
Topics that are common in data science interviews but not listed in the README as explicitly covered include: system design (how to build a recommendation system, how to handle class imbalance at scale), behavioral questions, A/B testing statistics beyond basic probability, and domain-specific questions for roles in NLP, computer vision, or time series. The README does not document what the repository leaves out, so candidates preparing for a specific role should check whether the relevant topics appear before relying on this repository alone.
Answer quality varies across questions. Some answers are detailed and well-sourced; others are brief stubs that were contributed as starting points. The README does not document a review process beyond pull requests, so there is no formal quality gate.
License, Maintenance, and Attribution Requirements
The repository is licensed under Creative Commons Attribution 4.0 International (CC-BY-4.0). This means anyone can read, share, and adapt the content, including for commercial purposes, as long as they give appropriate credit to the original source and link to the license. The attribution requirement is enforceable: redistributing the content without crediting the repository and its contributors would violate the license terms.
The last push was on 2026-09-26, confirming the repository is being maintained. The repository has no GitHub releases. Updates are delivered as direct pushes and merged pull requests. The author, Alexey Grigorev, maintains the repository and can be reached through the DataTalks.Club community or on Twitter at @Al_Grigor.
Comparison with Ace the Data Science Interview
Ace the Data Science Interview by Nick Singh and Kevin Feng is a widely known book covering data science interview preparation. It appears as a search term associated with this project category, indicating that candidates often consider both.
The difference in approach is significant. The Singh and Feng book is a single curated work by two authors, reviewed and edited before publication, covering a fixed set of questions and worked answers. It is a paid product. data-science-interviews is a free, community-maintained repository where anyone can contribute questions and answers, content is not uniformly edited, and the question set grows organically based on what contributors know and choose to add.
The tradeoff: a curated book offers consistent quality and editorial coherence; a community repository offers breadth, freshness, and the ability to fix errors through pull requests. The two resources are complementary, not mutually exclusive.
Editorial conclusion
data-science-interviews is the right starting point for candidates preparing for data science technical interviews who want a free, community-reviewed resource covering linear models, trees, neural networks, SQL, and Python coding. It is not a structured course; there is no guaranteed answer quality on every question, and the repository leaves gaps that contributors have not yet filled. Check the CC-BY-4.0 license before using the content in a commercial product, since attribution is required.
Frequently asked questions
What are the most common data science interview questions?
The repository organizes questions into theory (linear models, trees, neural networks), technical (SQL, Python, coding), and contributed topics (probability). The README does not rank questions by frequency, but the theory.md and technical.md files represent the areas that appear most consistently across technical interviews.
What to expect in a data science interview?
Based on the repository structure, data science technical interviews typically include questions on machine learning theory (how algorithms work and when they fail), SQL and Python coding screens, and probability questions. The repository collects questions from these areas with community-written answers.
What are data science interviews like?
The repository covers three main interview formats: conceptual machine learning questions tested in theory.md, hands-on coding questions in technical.md covering SQL and Python, and probability and statistics questions in the contributed section. The content reflects questions that have appeared in real data science screening rounds.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/alexeygrigorev-data-science-interviews)