MelihGulum/Comprehensive-Data-Science-AI-Project-Portfolio: A Collection of Project Folders, Not a Platform
A curated collection of AI, data engineering, and DevOps projects featuring real-world applications, advanced techniques, and tutorials—ideal for learners and practitioners exploring data science and machine learning.
At a glance
- What is it?
- This repository gathers separate machine learning, deep learning, data engineering and cloud projects, each in its own folder with its own README. It is useful for reading worked examples, and awkward if you expected one installable package.
- Who is it for?
- Adopt it as a reading and reference collection: pick one folder, such as the NBA Player Stats ETL pipeline or the Terraform Fundamentals guide, and work through that project's own README rather than treating the top-level repository as a package. Do not adopt it if you need a versioned dependency, a supported API, or a single command that runs everything, because no such command is documented.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 122 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What this repository actually is, and who it is for
This is not a library, a framework, or a tool you install and import. It is a monorepo-style collection of independent projects grouped by domain: Machine Learning Projects, Deep Learning Projects, Data Engineering, Data Analysis, AI Systems, Clouds and DevOps, and Tutorials. The top-level README describes it as "a curated collection of projects across Machine Learning, Deep Learning, Data Engineering, Data Analysis, AI Systems, and Cloud/MLOps" and says each project folder contains its own README with setup instructions, methodology, and results.
The intended reader is someone learning or refreshing applied data work: a student building a portfolio, an engineer moving from one domain into another, or a practitioner who wants to see how a full workflow is laid out rather than read a tutorial fragment. The featured table lists ten deep learning projects, seven machine learning projects, two data engineering projects, several EDA projects, a RAG-based retrieval system, and a Terraform guide. Breadth is the point. Depth is per folder, not per repository.
How the collection is organised, and what that means for reuse
The repository is a directory tree with a README at the root that acts as an index. There is no shared package, no common utility module described at the top level, and no single entry point. The root README is navigation: anchor links, collapsible details blocks, and tables mapping each project to a task and a set of tools.
That structure has a direct consequence. Code reuse across projects is manual. If the Medical Cost Prediction project and the Airline Customer Satisfaction project both use Optuna and SHAP, the top-level README lists those tools in both rows but does not claim a shared implementation. You copy what you need. Likewise, the tool list in the header (Python, Scikit-Learn, TensorFlow, SQL, Kafka, Docker, Terraform, Flask, AWS, GCP) describes the union of what appears across folders, not a stack that any one project requires. Treat it as a map of the territory, not a dependency manifest.
The domain grouping is the useful part. Audio work sits together (Urban Sound Classification, Music Genre Classification, the Urban Sound research project). Computer vision sits together (Face Mask Detection, Gender Detection, ASL Recognition, CIFAR-10, Facial Emotion Recognition). If you want to compare how two projects in the same domain handle the same problem, the layout makes that easy.
Installing and running: there is no repository-level install
The top-level README gives no install command, no requirements file reference, and no environment variable. What it does say is that each project folder contains a dedicated README with setup instructions. So the first real use is to clone the repository and open a specific folder.
git clone https://github.com/MelihGulum/Comprehensive-Data-Science-AI-Project-Portfolio.git
cd Comprehensive-Data-Science-AI-Project-Portfolio
lsThe listing should show the domain directories named in the README: AI Systems, Clouds and DevOps, Data Analysis, Data Engineering, Deep Learning Projects, Machine Learning Projects, Tutorials, and README.md. From there, move into one project and read its own README before running anything.
cd "Data Engineering/01. ETL - NBA Player Stats"
cat README.mdThe NBA Player Stats project is described in the root README as an ETL pipeline using web scraping, MSSQL and logging. Its folder README is where setup for that project lives. The same pattern applies to the Kafka pipelines folder, which the root README describes as real-time and batch data pipelines using Kafka, PostgreSQL and Docker. Because the root README does not document a Docker Compose file, a Makefile, or a requirements.txt at the repository root, do not assume one exists. Check the folder.
Where the collection stops being the right tool
The repository has no releases and no versioned artifacts, so there is nothing to pin. If your goal is to depend on code, this is the wrong shape: you would be vendoring files from a moving main branch. There is no package published under this name, and the README does not describe one.
The licence is not stated in the repository. That matters more here than in a typical library, because the repository is a body of source code and notebooks that you may want to copy from. Without a licence file, the default position is that no permission has been granted, and the README does not address reuse terms. If you intend to lift code into a commercial project, that is the first thing to resolve, and it is a question for the repository owner rather than something you can infer.
The third limitation is maintenance. The last push to the default branch was on 2026-05-31. The repository is not archived, but there is no release history and no changelog, so there is no way to tell from the README alone whether individual folders are finished, in progress, or abandoned. The 'Roadmap' anchor in the navigation suggests planned work, but the truncated README does not show what that roadmap contains.
How it compares with a single end-to-end project repository
The obvious alternative is a repository built around one substantial project: a single pipeline with a documented architecture, a test suite, a Dockerfile and a deployment story. That approach trades breadth for depth. You get one thing you can actually run end to end, and you learn the connective tissue between data ingestion, training, evaluation and serving because it is all in one place. This portfolio does not offer that. It offers many smaller pieces, each self-contained, and leaves the integration to you.
A second alternative is a course or a structured curriculum, where the order of projects is deliberate and each builds on the last. The README here groups by domain and marks a set of 'Featured Projects', but it does not present a sequence. The featured table mixes audio AI, computer vision, RAG, ETL and IaC in one list, which is a showcase ordering rather than a learning path.
The trade-off is honest: breadth makes this useful as a reference shelf, and the absence of a spine makes it poor as a single guided course. If you already know what you want to build, the folder structure helps you find a comparable example. If you do not, the repository will not decide for you.
Maintenance, upgrades, and what the licence gap means in practice
There is no dependency file at the repository root, so there is no upgrade path to manage at the repository level. Each folder is on its own. When a project pins a framework version, that pin lives inside that project, and the root README will not tell you about it. Upgrading anything means opening the folder and reading its README.
The last push was on 2026-05-31, which is recent enough that the repository is not dormant, but there are no releases, no tags and no changelog. That combination means you should not expect deprecation notices. If a folder uses an older API, nothing here will flag it.
On licensing: no licence is stated for this repository. The README uses the word 'showcases' and invites contribution through a Contributing section, but it does not grant reuse rights. Practically, if you want to copy a notebook or a pipeline into your own work, ask the maintainer first. This is not legal advice; it is a statement that the permission question is unresolved.
Editorial conclusion
Adopt it as a reading and reference collection: pick one folder, such as the NBA Player Stats ETL pipeline or the Terraform Fundamentals guide, and work through that project's own README rather than treating the top-level repository as a package. Do not adopt it if you need a versioned dependency, a supported API, or a single command that runs everything, because no such command is documented. Before relying on any folder, open it and confirm it has a README with setup instructions, since the top-level README states that each project folder contains one but does not guarantee it for every directory.
Frequently asked questions
What are some good data science portfolio projects?
This repository is itself a list of candidates, grouped by domain. The root README features an Urban Sound Classification project with Flask deployment, a real-time Face Mask Detection project using MobileNetV2 and transfer learning, a University Information Retrieval System built on RAG with a vector database, and an NBA Player Stats ETL pipeline using web scraping and MSSQL.
What is an AI portfolio?
In this repository the term refers to a curated collection of projects across machine learning, deep learning, data engineering, data analysis, AI systems and cloud/MLOps, each in its own folder with a dedicated README covering setup, methodology and results.
What are 7 real-world AI projects I can build in 2026?
The repository does not answer this directly, but the Machine Learning Projects section lists seven projects: Diabetes Classification, Heart Attack Prediction, Medical Cost Prediction, Melbourne House Price Prediction, Clustering Techniques, Airline Customer Satisfaction, and Forecasting USD-TRY Exchange Rates.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/melihgulum-comprehensive-data-science-ai-project-portfolio)