Model or dataset
patchy631/ai-engineering-hub avatar
patchy631/ai-engineering-hub

AI Engineering Hub: 93 Jupyter Projects on LLMs, RAG and Agents

In-depth tutorials on LLMs, RAGs and real-world AI agent applications.

38,109 stars6,260 forksJupyter NotebookMIT

At a glance

What is it?
AI Engineering Hub is a MIT-licensed collection of over 90 self-contained Jupyter Notebook projects on LLM application engineering, retrieval-augmented generation and multi-agent workflows. Projects span three difficulty tiers, from single-component beginner exercises to production-grade fine-tuning and evaluation pipelines, and a roadmap directory guides first-time learners through a suggested sequence.
Who is it for?
AI Engineering Hub suits engineers who want a broad catalogue of working code to study and adapt across RAG, agents, MCP and multimodal tasks. It is not a structured course with assessments or feedback, so learners who need guided progression will find it insufficient as a sole resource.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What the Hub contains and who it is for

AI Engineering Hub is a GitHub repository that collects over 90 self-contained project directories, each addressing a distinct AI engineering task. The scope runs from optical character recognition with vision models, through retrieval-augmented generation over structured and unstructured documents, to multi-agent systems that coordinate across roles and tools.

The primary audience is engineers who learn by reading and modifying working code. Projects are not lecture notes or slide decks. Each directory contains one or more Jupyter Notebooks, and the intention is that a reader can clone the repository, navigate to a directory, and run it with minimal setup. The README describes the repository as targeting beginners, practitioners and researchers, though the advanced tier assumes solid familiarity with the underlying frameworks.

The homepage link in the repository points to the Daily Dose of Data Science newsletter, indicating the project also serves as a companion resource to that publication. The MIT licence allows free use and modification.

Three difficulty tiers and what each covers

The README divides the 93 projects into three tiers. Twenty-two are labelled Beginner, covering single-component implementations. Examples include building a local OCR app with Llama 3.2 and Streamlit, running a simple RAG workflow with LlamaIndex and Ollama, and creating a chatbot that streams responses in real time using the Motia framework.

Forty-eight projects occupy the Intermediate tier, introducing multi-component systems and agentic workflows. Representative examples are an AutoGen-powered stock analyst, a CrewAI content planner, a real-time voice agent using AssemblyAI and Cartesia for speech-to-text and synthesis, and an agentic RAG pipeline that falls back to a web search when the local document store returns no useful result.

Twenty-three Advanced projects address fine-tuning, production deployment and research-adjacent techniques. The README lists a DeepSeek fine-tuning directory, a ColBERT-based RAG implementation and an LLM evaluation and observability pipeline at this tier. Setup documentation is thinner at this level, and some directories assume prior experience with the frameworks involved.

Model Context Protocol projects and multimodal coverage

A dedicated MCP section groups projects that connect agents to external data sources through the Model Context Protocol. Listed examples include integrating Cursor with the Linkup search API via an MCP server and deploying a cloud-hosted agent that uses an MCP server for memory via the Graphiti framework.

The multimodal section goes beyond standard text-over-PDF retrieval. One project applies IBM's Docling library to parse and retrieve over Excel spreadsheets. Another builds a website RAG system that treats web pages as images rather than extracted text, which is a different approach from the standard HTML-to-markdown pipeline used by most web retrieval tools. A third combines audio transcription, vector storage and a CrewAI agent into a single pipeline.

A context-engineering-pipeline and a context-engineering-workflow directory appear in the repository listing, as does an ai-engineering-roadmap directory. The README does not document what the roadmap directory contains in detail.

How to start using the repository

The README does not supply a single install script or a root-level requirements.txt. The suggested approach is to pick a Beginner project, such as the simple RAG workflow or one of the OCR apps, and read its directory-level README before installing anything.

The repository includes a .gitmodules file, which indicates that some subdirectories are git submodules. Cloning with a plain git clone may leave those directories empty. The README does not document which directories are submodules or how to initialise them.

Because each project specifies its own dependencies, the practical first step is to enter the chosen subdirectory and look for a requirements.txt or pyproject.toml. That file tells you exactly which Python packages and versions that specific project needs.

Where the project falls short

The repository carries no test suite and no CI configuration at the root level. A notebook that ran correctly at the time it was added can break silently when an API client library releases a new version. Several intermediate and advanced projects depend on external paid services, including SambaNova for fast inference, GroundX for enterprise document retrieval, Cartesia for voice synthesis and BrightData for web scraping. Switching to a different provider typically requires rewriting the integration.

Quality is not uniform across all 90 directories. Some contain multiple annotated notebooks with setup instructions; others supply a single notebook and a brief README paragraph. There is no stated acceptance standard for contributions.

Jupyter Notebooks are the primary format throughout. Notebooks are harder to review and version-control cleanly than plain Python scripts, and converting a notebook into a production-ready module requires additional work beyond what the repository addresses.

Comparing AI Engineering Hub to deeplearning.ai short courses

The deeplearning.ai short course catalogue covers overlapping ground, including RAG, agents and fine-tuning. Those courses provide video lectures, autograded exercises and a defined learning sequence. AI Engineering Hub provides none of these. Access to the Hub requires no registration or subscription, while deeplearning.ai courses require an account and some are behind a paywall.

The trade-off is breadth against depth. A deeplearning.ai course on RAG walks through a single well-tested implementation with explanation at each step. The Hub provides dozens of RAG implementations built on different stacks, but deciding which matches a given use case, and understanding why particular design decisions were made, is left to the reader.

Maintenance status and licence

The last push to the repository was on 2026-09-10. The repository carries no formal GitHub releases, so there is no changelog and no semantic version history. New project directories are added directly to the main branch without a release tag.

The MIT licence permits free use, modification and redistribution of the code in the repository. It does not cover the third-party services referenced in individual projects; AssemblyAI, Cartesia, GroundX, SambaNova and others each carry their own pricing, usage limits and terms of service.

Editorial conclusion

AI Engineering Hub suits engineers who want a broad catalogue of working code to study and adapt across RAG, agents, MCP and multimodal tasks. It is not a structured course with assessments or feedback, so learners who need guided progression will find it insufficient as a sole resource. Before adopting any project, confirm that its third-party API dependencies, such as AssemblyAI, GroundX or Cartesia, match your budget and data requirements, since the MIT licence covers only the code the repository contains, not those external services.

Frequently asked questions

Is AI Engineering Hub free to use?

The repository is MIT-licensed, so the code is free to use, modify and distribute. Several projects require paid third-party API keys from services such as AssemblyAI, GroundX or Cartesia, which the MIT licence does not cover.

What skill level do I need to start with AI Engineering Hub?

The README explicitly targets beginners, practitioners and researchers with its three-tier structure. The 22 Beginner projects focus on single-component implementations. The Advanced tier assumes familiarity with frameworks such as AutoGen, CrewAI and model fine-tuning.

Does AI Engineering Hub include a structured learning path?

The repository includes an ai-engineering-roadmap directory and groups projects by difficulty level. The README does not document the full contents of the roadmap directory, and there are no assessments, quizzes or guided exercises.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/patchy631-ai-engineering-hub.svg)](https://hysenlabs.com/projects/patchy631-ai-engineering-hub)
Community notes

Community notes