LLM Zoomcamp: Free 10-Week Course on RAG, Vector Search, and LLM Applications
LLM Zoomcamp - a free online course about real-life applications of LLMs. In 10 weeks you will learn how to build an AI system that answers questions about your knowledge base. Register here 👇🏼
At a glance
- What is it?
- LLM Zoomcamp is a free, hands-on online course from DataTalks.Club that teaches production-ready LLM applications in 10 weeks. Covers retrieval-augmented generation, vector search, embeddings, agents, evaluation, monitoring, and a capstone project. Available as live cohorts with peer review and certificates, or self-paced.
- Who is it for?
- Join LLM Zoomcamp if you are a software engineer, data engineer, or ML practitioner who learns best by building projects. Do not enroll if you need certification for immediate hiring or if you are not comfortable with Python and the command line.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Who LLM Zoomcamp is for and its structure
LLM Zoomcamp teaches people who learn by doing. The course is split into two parallel tracks: live cohorts and self-paced. Live cohorts have pre-recorded lectures (not live classes), graded homework, peer review of three projects, a leaderboard, and certificates upon completion. Self-paced learners access the same materials on GitHub, skip grading and certificates, but keep all homework and capstone code. Both modes are free and use identical video lectures and course materials.
The course targets software engineers adding LLMs and RAG to production, data engineers learning vector search and retrieval pipelines, and ML practitioners building evaluation and monitoring systems. Prerequisites are intentionally low. Students need confident Python coding (enough to write a function), familiarity with the command line, basic Docker knowledge, and about USD 1-5 in API credits for calling external LLM APIs. ML expertise and GPU hardware are not required; any laptop or PC works. The README explicitly states that if you can write a Python function and have heard of ChatGPT, you have enough to get started. No machine learning background is necessary; the course teaches LLM concepts from the ground up.
Live cohort members register at courses.datatalks.club and work with announced deadlines, receive scored homework and project feedback, participate in peer review, and earn certificates. Live cohorts run on a schedule with specific start dates and cohort deadlines published in advance. Self-paced learners follow the GitHub materials at their own pace, can ask questions in the Slack community (#course-llm-zoomcamp), and complete homeworks without scoring. The repository is at https://github.com/DataTalksClub/llm-zoomcamp with all materials in the main branch. Neither mode charges tuition. The last push to the repository was on 2026-09-15.
Curriculum and hands-on modules
The course spans 7 core modules plus a capstone. Module 1 (Agentic RAG) covers building a retrieval-augmented generation pipeline with keyword search, then adding function calling to make it agentic. RAG is the core technique for making LLMs aware of your data. Module 2 (Vector Search) teaches semantic search with embeddings using three vector storage backends: minsearch (in-memory), sqlitesearch (SQLite), and PGVector (PostgreSQL). This module is critical for semantic retrieval at scale.
Module 3 (Orchestration) covers AI orchestration with Kestra, a platform for building and scheduling complex data pipelines and workflows. A workshop on Data Ingestion teaches dlt pipelines for ingesting and analyzing LLM traces, covering Filesystem and REST API sources, DuckDB for analytics, and marimo dashboards for visualization. Module 4 (Evaluation) measures retrieval and answer quality through offline evaluation (metrics like Mean Reciprocal Rank, NDCG) and online evaluation (user feedback signals). Module 5 (Monitoring) covers monitoring user feedback and system health with live dashboards to track performance in production. Module 6 (Best Practices) teaches hybrid search combining vector and keyword search, reranking results with cross-encoders for higher precision, and using LangChain for building chains. Module 7 (End-to-End Project) provides a complete fitness assistant example showing all concepts integrated.
The capstone project and community projects
The capstone is where students apply everything end-to-end. Build a searchable knowledge base by choosing a dataset, ingesting it, cleaning it, and storing it for retrieval. This might be documents, Q&A pairs, textbooks, or any structured data. Implement a full RAG pipeline: retrieve relevant context from the knowledge base, assemble a prompt combining the context and user question, call an LLM (OpenAI, Anthropic, or another provider), and return grounded answers sourced from your data. Create an evaluation process measuring how well your system retrieves relevant documents and whether the final answers are accurate. This might use search metrics (Mean Reciprocal Rank, NDCG) or LLM-as-a-judge evaluation. Build a user-facing interface using Streamlit, FastAPI, or similar frameworks so others can interact with your application. Finally, add monitoring and feedback loops to track user queries, collect feedback on answer quality, and monitor system performance over time.
Past community projects include fitness and nutrition assistants, study companions for textbooks or course notes, medical FAQ assistants for healthcare data, codebase Q&A bots for documentation, and news summarization and retrieval tools. Live cohorts submit projects within deadlines; self-paced learners build for their portfolio. Browsing the course website shows 2024 and 2025 cohort submissions with working projects for inspiration. The project.md file in the repository contains full capstone guidelines.
Obtaining certification and peer review
Live cohort members earn certificates by completing the final project demonstrating all course concepts, reviewing three peers' projects with written feedback, and meeting cohort deadlines. Certificates are issued after all peer reviews are completed. Self-paced learners are not eligible for certification but can build and share portfolio projects freely, which serve the same purpose for demonstrating skills.
Certificates are issued through DataTalks.Club's platform and include verification on LinkedIn. The certificate guide (on datatalks.club/docs/courses/zoomcamp-logistics/certification/) explains how to add the certificate to LinkedIn profiles. Live cohort schedules and deadlines are listed at courses.datatalks.club/llm-zoomcamp-2026 (the current cohort). The community uses Slack (#course-llm-zoomcamp channel in the DataTalks.Club workspace) for support and questions, with responses from instructors and peers. Telegram (@llm_zoomcamp) is used for announcements about new modules and cohorts. Video lectures are available on a YouTube playlist (PL3MmuxUbc_hLZFNgSad56pDBKK8KO0XIv), and all materials live in the GitHub repository.
Course materials and support infrastructure
All course materials are free, open-source, and available on GitHub at DataTalksClub/llm-zoomcamp. Video lectures are on a YouTube playlist covering each module. Documentation is split into general zoomcamp logistics (datatalks.club/docs/courses/zoomcamp-logistics/) and LLM Zoomcamp-specific guides (datatalks.club/docs/courses/llm-zoomcamp/). The repository contains notebooks and code examples for each module, stored in 01-agentic-rag/, 02-vector-search/, and so on. The course.yaml file defines the curriculum. Examples sometimes include implementations for different frameworks (JavaScript with Next.js, Python with Streamlit, etc.).
The cohort schedules, leaderboards for live cohorts, and project submissions are on courses.datatalks.club. The leaderboard shows how many assignments live cohort members have completed. Instructors listed in the README teach and contribute. The Slack community (link on datatalks.club/slack.html) is active with hundreds of members sharing projects, answering questions, and collaborating. The Telegram channel shares announcements about new cohort dates, module releases, and platform updates. Self-paced learners can follow GitHub materials directly, ask questions in Slack, and build projects without time pressure.
Editorial conclusion
Join LLM Zoomcamp if you are a software engineer, data engineer, or ML practitioner who learns best by building projects. Do not enroll if you need certification for immediate hiring or if you are not comfortable with Python and the command line. Before starting, ensure you can write Python functions confidently, have Docker installed, and are willing to spend approximately USD 1-5 on API credits. The last push was on 2026-09-15, and the course runs in live cohorts and self-paced modes simultaneously.
Frequently asked questions
Is there a free LLM certification available?
Yes, through LLM Zoomcamp's live cohort track. Complete the final project, review three peers' projects, and meet deadlines to earn a certificate.
Do I need a GPU to take LLM Zoomcamp?
No. The course requires any laptop or PC but no GPU hardware. You may call external APIs rather than run models locally.
What is the best LLM engineering course?
LLM Zoomcamp teaches production-ready LLM applications through a 10-week hands-on curriculum covering RAG, vector search, evaluation, monitoring, and a capstone project. It offers both live cohorts with peer review and self-paced learning.
What modules are covered in LLM Zoomcamp?
Modules cover Agentic RAG, Vector Search, Orchestration, Evaluation, Monitoring, and Best Practices. There is also a workshop on Data Ingestion and a capstone project.
Can I access LLM Zoomcamp materials offline or self-paced?
Yes. Self-paced learners can follow all materials on GitHub at their own pace. They do not receive scoring or certificates but keep all code and can build portfolio projects.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/datatalksclub-llm-zoomcamp)