Model or dataset
sajal2692/data-science-portfolio avatar
sajal2692/data-science-portfolio

sajal2692/data-science-portfolio: A Reference Notebook Collection Refreshed for Python 3.14

Portfolio of data science projects completed by me for academic, self learning, and hobby purposes.

1,425 stars466 forksJupyter NotebookMIT

At a glance

What is it?
It is a set of data science notebooks from around 2016, re-run on a current Python stack in 2026. The value is as a study reference for people building their own portfolio, not as a library you install.
Who is it for?
Adopt it as a reading and reference repository if you are assembling your first data science portfolio and want to see how someone else structured notebooks, ETL steps and model comparisons end to end. Do not adopt it if you need an installable library, a maintained API, or a template for a portfolio website; the README points to sajalsharma.com for current work and this repository is a collection of notebooks, not a package.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 85 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem sajal2692/data-science-portfolio solves, and who it is for

The README is explicit about its audience. These are projects the author built early in his career, around 2016, and the stated hope is that someone "starting out, learning the fundamentals or putting together your own portfolio" finds them a useful reference. That framing matters, because it tells you what the repository is not: it is not a framework, not a package, and not a service. It is a set of worked examples.

The problem it addresses is a real one for beginners. A portfolio is easy to describe and hard to assemble. You need to show that you can load messy data, clean it, pick a model, evaluate it, and explain what the result means. Each notebook here does one of those things in a self-contained way. The Boston housing notebook tunes a decision-tree regressor and evaluates how reliable the predictions are. The finding_donors notebook compares several supervised learning algorithms on a binary income prediction. The customer_segments notebook applies PCA and Gaussian-mixture clustering to wholesale spending. The digit recognition notebook trains a PyTorch CNN with five classification heads on a shared convolutional trunk to read sequences of one to five digits from a single image.

The breadth is the point. Machine learning, NLP, exploratory analysis and visualisation all appear, with a separate set of R analyses published on RPubs. If you are deciding what a portfolio should contain, seeing a working example of each category is more useful than reading a list of categories. The repository also carries a note on the Boston dataset's history and why it was retired from scikit-learn, which is the kind of context a beginner would not think to look up.

How the repository is organised and how a notebook actually runs

The layout is flat and notebook-first. Top-level entries include the standalone notebooks (Titanic exploratory analysis, SMS spam classification, Yelp review classification, 3-way sentiment analysis, cross language information retrieval, stock market analysis, 911 calls, 2016 election polls) and a handful of directories for the larger projects (boston_housing/, customer_segments/, finding_donors/). There is also an ML Micro Projects/ directory described as short, focused scikit-learn exercises, and a data/ directory the README says holds small demonstration datasets.

Data flow is deliberately simple. Most notebooks read from data/ on disk. Two notebooks fetch their data at runtime and cache it: the digit recognition notebook downloads MNIST through torchvision, and the stock market notebook pulls tech-stock prices through yfinance, with a vendored snapshot as a fallback. That fallback is worth understanding before you run the stock notebook, because it is the difference between a notebook that works offline and one that fails on a network error.

The dependency story is the most interesting architectural decision. The pyproject.toml pins requires-python to ">=3.14" and lists the scientific stack (numpy, pandas, scipy, scikit-learn, matplotlib, seaborn), NLP tools (nltk, transformers), yfinance for finance data, and torch plus torchvision for deep learning. One comment in that file is unusually candid: gensim is intentionally omitted because there was no cp314 wheel as of 2026-07 and its source build fails on Python 3.14, so the sentiment notebook's word2vec use is reimplemented in numpy. That is a real constraint, not a preference. It also means the sentiment notebook is not a drop-in word2vec tutorial; it is a numpy reimplementation with a zero-shot twitter-roberta transformer baseline alongside for comparison.

The pyproject.toml also sets package = false under [tool.uv], with the comment that the repository is a collection of notebooks, not an installable package. That single line explains why there is no importable module here and why you should not expect one.

Installing it and running a first notebook on Python 3.14

The README gives two install paths. The primary one uses uv, and the repository ships a uv.lock and a .python-version file, so the environment is reproducible if you use it. From the repository root, the documented commands are:

bash
uv sync
uv run jupyter lab

uv sync resolves the lock file and creates the environment, including the torch and torchvision wheels. uv run jupyter lab then starts JupyterLab inside that environment. If your platform has no matching wheel for one of the pinned packages, uv sync is where you will find out, not at notebook runtime.

The README also documents a pip path, using the autogenerated requirements.txt:

bash
pip install -r requirements.txt

That file carries a header stating it was generated by uv export, and it includes platform markers such as appnope on darwin and cuda-toolkit on linux, so the resolved set differs by operating system. Because the file is generated, editing it by hand is a bad idea; regenerate it instead.

Once the environment exists, open the notebook you want and run it top to bottom. The README states that as of 2026 the whole repository has been refreshed so every notebook still runs top-to-bottom on a current Python 3.14 stack. Start with a notebook that reads only from data/, such as the Titanic exploratory analysis or the SMS spam classification, so your first run does not depend on a network download. If you want the deep learning notebook, expect the first cell to pull MNIST through torchvision, and if you want the stock notebook, expect yfinance to fetch prices before the vendored snapshot is used.

Where the repository breaks down: gensim, network fetches and the absence of tests

The gensim omission is the clearest limitation, and it is documented rather than hidden. The pyproject.toml comment says gensim was left out because there was no cp314 wheel as of 2026-07 and its source build fails on Python 3.14. The consequence is that the word2vec portion of the sentiment notebook is a numpy reimplementation. If your goal is to learn the gensim API, this notebook will not teach it to you. If your goal is to understand the mechanics of word embeddings, the reimplementation may be more instructive than the library call, but that is a different lesson.

The second failure mode is network dependence. Two notebooks fetch data on first run. The digit recognition notebook downloads MNIST through torchvision, and the stock market notebook pulls prices through yfinance. The README notes a vendored snapshot as a fallback for the stock notebook, but no such fallback is mentioned for the MNIST download. In an offline or restricted environment, that notebook is the one most likely to stop early.

The third gap is structural rather than technical. The repository has a .github/ directory, but nothing in the README describes a test suite, a CI workflow that executes the notebooks, or a check that the notebooks still produce the outputs shown. The README's claim that every notebook runs top-to-bottom on Python 3.14 is a statement about the 2026 refresh, and the last push was on 2026-07-07. There is no release history, so there is no changelog to consult when a dependency moves. If a notebook breaks after a scikit-learn or torch upgrade, you will be debugging it yourself.

Finally, the data under data/ is described as being for demonstration purposes only. That is fine for learning and wrong for anything that needs provenance or licensing clarity on the underlying records.

How it compares with a tutorial series or a portfolio site template

The closest alternatives are a structured course or tutorial series, and a portfolio website template. They solve adjacent problems and the differences are concrete.

A tutorial series usually walks one path through one dataset, with the author controlling every step and explaining each one. This repository is a collection of finished notebooks across many problem types, and the explanations live in the notebooks themselves rather than in a separate curriculum. You get breadth and working code, but you do not get a guided sequence that tells you what to learn next. If you already know the fundamentals and want to see how a clustering project or a CLIR system is put together, the repository format is better. If you are starting from zero, a tutorial series will hold your hand in a way this will not.

A portfolio website template solves the presentation problem, not the analysis problem. The related searches around portfolio websites, templates and examples are about how the work is displayed. This repository does not address that at all: there is no site generator, no theme, no deployment configuration. Its homepage field points to sajalsharma.com, which the README describes as the place for the author's current AI engineering work. So the honest comparison is that a template gets you a page and this gets you the analysis behind the page. Most people assembling a portfolio need both, and this repository only supplies one half.

Within the notebook space, the distinguishing choice here is the 2026 refresh to Python 3.14. Many public notebook collections from 2016 will not run on a current interpreter without edits. The refresh, the uv lock file and the explicit note about gensim are what separate this from an abandoned archive.

Maintenance cost, licensing and what to check before you build on it

The repository is not archived, and the last push was on 2026-07-07. That is recent enough that the Python 3.14 refresh described in the README is the current state of the code, but there is no release history, so there is no versioned changelog to track. Upgrades are therefore a manual exercise: change a pin in pyproject.toml, run uv sync, and re-run the affected notebooks. The lock file makes that reproducible, which is the main reason to use the uv path rather than pip.

The licence is MIT, declared both in the LICENSE file at the repository root and in the pyproject.toml license field. MIT is permissive, so reuse and adaptation are broadly allowed, but the notebooks bundle or download third-party datasets. The README says the data under data/ is for demonstration only. That disclaimer covers the repository's own files, not the terms attached to MNIST, the Kaggle 911 calls dataset, the Yelp reviews, the SMS spam corpus or the stock price data. If you plan to republish a notebook or its outputs, check the upstream terms for each dataset separately. Nothing here is legal advice, and the MIT grant on the code does not extend to data the code pulls in.

The practical upgrade risk is concentrated in three places: torch and torchvision for the digit recognition notebook, transformers for the twitter-roberta baseline, and yfinance for the stock notebook. yfinance in particular scrapes a source that the pyproject.toml comment describes as replacing a dead pandas-datareader Yahoo backend, which is a reminder that this dependency class has broken before. The vendored snapshot softens that for the stock notebook. Nothing softens it for MNIST.

Who should take this repository seriously, and who should not

Take it seriously if you are learning and want to see complete, runnable examples of common data science tasks on a current interpreter. The 2026 refresh is the reason to prefer it over the many 2016-era notebook collections that no longer execute. The MIT licence makes it easy to fork and adapt for your own study.

Do not take it seriously as infrastructure. There is no importable package, no API, no test suite described in the README, and no release process. If you need a library, look elsewhere. If you need a portfolio website, this is the wrong half of the problem. If you need word2vec through gensim, the sentiment notebook will not give it to you because gensim is deliberately absent.

One more caveat on the numbers. The README's claim that every notebook runs top-to-bottom on Python 3.14 is a claim about the refresh, and the last push was on 2026-07-07. Dependency pins age. The first thing to verify on your own machine is that uv sync completes on your platform, since the requirements.txt markers show the resolved set differs between darwin, win32 and linux. The second is that the two network-dependent notebooks, digit recognition and stock market analysis, can reach their data sources or fall back cleanly.

Editorial conclusion

Adopt it as a reading and reference repository if you are assembling your first data science portfolio and want to see how someone else structured notebooks, ETL steps and model comparisons end to end. Do not adopt it if you need an installable library, a maintained API, or a template for a portfolio website; the README points to sajalsharma.com for current work and this repository is a collection of notebooks, not a package. Before relying on any notebook, run uv sync on Python 3.14 and open the notebook you care about, because the two notebooks that fetch data on first run depend on torchvision and yfinance downloads succeeding, and the vendored stock snapshot is the only fallback the README mentions.

Frequently asked questions

What is sajal2692/data-science-portfolio?

It is a collection of data science projects presented as Jupyter notebooks, plus a few R analyses published on RPubs. The README describes them as projects built early in the author's career, around 2016, and refreshed in 2026 to run on a current Python 3.14 stack.

How do I build a data science portfolio with sajal2692/data-science-portfolio?

The README documents uv sync followed by uv run jupyter lab, or pip install -r requirements.txt as an alternative. The project requires Python 3.14, and pyproject.toml sets package = false because the repository is a collection of notebooks rather than an installable package.

What should a data science portfolio include, based on sajal2692/data-science-portfolio?

The contents list groups work into machine learning, natural language processing, and data analysis and visualisation, with a separate set of micro projects. That grouping is one concrete answer to the question, since it shows the categories the author chose to cover.

Is sajal2692/data-science-portfolio a good example of a data science portfolio?

It is a reference for notebook structure and problem selection rather than a portfolio website. The README states the author's current work lives at sajalsharma.com, and the repository itself contains no site generator or deployment configuration.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. sajal2692/data-science-portfolio on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sajal2692-data-science-portfolio.svg)](https://hysenlabs.com/projects/sajal2692-data-science-portfolio)