Model or dataset
calmrocks/ai-engineer-notebooks avatar
calmrocks/ai-engineer-notebooks

calmrocks/ai-engineer-notebooks: a framework-free applied-LLM track on the free Groq API

Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, and customer craft. Runs on the free Groq API.

642 stars52 forksJupyter NotebookMIT

At a glance

What is it?
The repository is a set of runnable Colab notebooks that build model APIs, RAG, evals and agents from raw API calls before any framework appears. It is aimed at backend engineers moving into AI Engineer or Forward Deployed Engineer roles, and its most useful habit is measuring before tuning.
Who is it for?
Adopt it if you already ship production code and want the applied-model layer demonstrated through raw API calls, with evals introduced before anything you would need to tune. Skip it if you want a maintained framework to build a product on, or if you need the LoRA and serving notebooks to be fully runnable on the free tier, since the README describes those two as concept-first with optional GPU appendices.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 34 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap these notebooks fill for engineers moving into AI Engineer or FDE roles

Most LLM tutorials hand you a framework and a result. You copy the cell, it prints an answer, and you have learned the shape of a library rather than the shape of a system. This repository takes the opposite position: you write the agent loop, the retrieval path and the evaluation harness from raw API calls first, so that when you later pick up LangChain or LlamaIndex you can tell which part of your problem the wrapper is solving and which part it is hiding. The README states this directly, calling the project framework-free on purpose and arguing that patterns are durable while wrappers churn.

The intended reader is specific. According to the README, it is for backend or full-stack engineers moving into AI Engineer, FDE, Applied AI or Solutions Engineer (AI) roles, people who can already ship production code and want the applied-model layer on top. That is a narrower audience than a general machine-learning course. There is no treatment of training transformers, no math prerequisites, and no dataset work. The subject is the layer between a foundation model and a deployed feature: how you call it, how you constrain it, how you know it got worse, and how you keep it running.

The cost constraint is part of the design, not an afterthought. Everything runs on the free Groq API, which the README notes requires no credit card. Two topics Groq cannot host, LoRA fine-tuning in section 06 and self-hosted serving in section 09, are described as concept-first with optional fenced Colab-GPU appendices that the README says were verified on a real Colab T4. That split is honest about what a free tier can demonstrate, and it is also the first thing to check before you plan a study schedule around the later sections.

How the track is organised, from environment setup to a capstone

The repository layout is the curriculum. Top-level directories run from 00-setup through 12-case-studies-and-capstone, with numbered sections for model APIs, evals, RAG, agents, adaptation, security, operations, serving and inference, ML system design, and customer craft. Each notebook is self-contained: the README says each installs its own dependencies and reads API keys from Colab secrets, and each ends with exercises. That self-containment is what makes the notebooks usable out of order, though the README recommends working top to bottom.

The ordering carries an argument. Section 01 covers prompting fundamentals, structured output, tool calling, streaming, and context and caching. Section 02 is a single notebook, measuring outputs, and the README calls evals the spine of the project, explaining that the measure-before-you-tune habit is installed before you build anything you would need to tune and returns in every section after. In practice that means you meet golden sets and metrics on the section-01 tasks rather than on a toy dataset invented for the occasion. Section 03 then covers RAG across five notebooks: what RAG is, embeddings and retrieval, hybrid and reranking, chunking, and why RAG fails. The README's guidance on chunking is worth quoting in spirit: it is revisited last, once you can judge chunking strategies against retrieval results rather than in the abstract.

Two structural choices stand out. First, the case studies in section 12 are described as end-to-end rather than illustrative, including a support assistant debugged in production, a pipeline-versus-agent cost comparison, and a red-team robustness benchmark. Second, the API surface is OpenAI-compatible throughout, so the README claims every pattern transfers to OpenAI directly and to Anthropic with small changes. That compatibility is the reason the framework-free approach stays practical: the code you write is not a Groq dialect.

Installing the shared helpers and running a first notebook

The notebooks are designed to run in Colab, and the README's badge points at the environment notebook as the entry point. There is also a small Python package in the repository, named aien, described in pyproject.toml as shared setup helpers for credential loading and a Groq client. Its only dependency is groq>=0.11, and it requires Python 3.9 or later. If you want those helpers outside Colab, install the package from the repository root.

bash
pip install -e .

The top-level requirements.txt is broader than the package, listing groq>=0.11, sentence-transformers>=3.0, tiktoken>=0.7, numpy, requests, jupyter and mlflow>=3.1. That file reflects what the notebooks pull in across sections rather than what the helper package needs, so installing it globally is heavier than installing aien alone.

The README states that notebooks read API keys from Colab secrets. Create a key at the Groq console, then open the environment notebook in Colab and add the secret before running cells that call the API. The README does not document the exact secret name beyond describing keys as coming from Colab secrets, so if a notebook raises an authentication error, confirm the name you used against the cell that reads it. The environment notebook also covers spend guards and model picking, which is where you should start if you are worried about accidentally leaving a loop running.

Where the framework-free approach costs you time

Writing your own agent loop and retrieval path is the point of the project, but it is also the main limitation. If your goal is to ship a document-question-answering feature this month, you will spend time reimplementing chunking, retrieval and orchestration that a framework already provides, and the notebooks do not pretend otherwise. The README frames the trade explicitly: you learn what LangChain and LlamaIndex do before reaching for them, and you learn when not to. That is a learning outcome, not a delivery schedule.

The dependency situation is the second constraint. The helper package pins only groq>=0.11, but requirements.txt lists sentence-transformers, tiktoken, numpy, requests, jupyter and mlflow with lower bounds rather than exact versions. Nothing in the repository pins a lockfile, so a notebook run months from now may resolve different library versions than the ones the author used. For notebooks that call a hosted API this is usually tolerable. For the embedding and retrieval notebooks it is the kind of drift that produces confusing similarity results.

The third constraint is coverage. LoRA fine-tuning and self-hosted serving are concept-first, with GPU appendices that the README says were verified on a Colab T4. If you need a working fine-tuning pipeline rather than an understanding of when fine-tuning beats prompting, this is the wrong resource. The same applies to anything requiring a persistent service: the notebooks run in Colab, and the repository does not present itself as a deployment toolkit. Section 09 covers serving as a topic, not as an operational runbook you would hand to an on-call engineer.

How it compares with LangChain-oriented tutorials and hosted course platforms

The closest alternative is a framework-first tutorial, typically built around LangChain or LlamaIndex. The difference is where abstraction enters. A LangChain tutorial starts with a chain or an agent executor and shows you the configuration that makes it work; this repository starts with the HTTP call and shows you the loop that the executor implements. If you already know the framework, the notebooks will feel slower and more verbose. If you have used the framework and could not explain why retrieval returned the wrong chunk, the raw version is the one that answers the question.

A second alternative is a hosted course platform with graded exercises and a certificate. Those usually provide a managed environment, video instruction and a fixed syllabus. This project provides notebooks, a README and exercises at the end of each notebook, with no grading and no schedule. The trade is control: you run the code on your own Groq key, in your own Colab, and you can read every line. The cost is that nothing tells you whether you got the exercise right except the evals you wrote yourself, which is arguably the lesson.

A third comparison is the companion career plan linked from the README, titled Plan: Transitioning to Forward Deployed Engineer / AI Engineer. The README describes the relationship plainly: the plan explains what to learn and why, and the notebooks are where you run it. If you only want the roadmap, the plan page is the shorter read. The notebooks are for the part where you type the code.

Maintenance, licence and what upgrading actually involves

The repository is not archived, and the last push was on 2026-08-28. The only release listed is v0.1.0, published on 2026-08-27 and described as the full applied-LLM track. That is a young project with a single tagged release, so treat the notebooks as a coherent first version rather than a long-stabilised surface. There is no changelog in the repository, and the README does not document a migration path between versions.

Upgrade cost is low in one sense and open-ended in another. Because the notebooks are self-contained and install their own dependencies, there is no application to redeploy and no schema to migrate. You pull the branch and rerun cells. The open-ended part is dependency drift: with lower-bound pins in requirements.txt and no lockfile, a notebook that worked can start behaving differently when sentence-transformers or mlflow resolves to a newer release. If you build anything on top of the aien helpers, the pyproject.toml constraint of groq>=0.11 is the only version guarantee you have.

The project is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That matters if you plan to lift patterns or helper code into a work project. It is a permissive licence, not a copyleft one, so derivative work does not have to be released under the same terms. This is a description of the licence text, not legal advice; check the LICENSE file and your organisation's policy before copying code into a product.

Editorial conclusion

Adopt it if you already ship production code and want the applied-model layer demonstrated through raw API calls, with evals introduced before anything you would need to tune. Skip it if you want a maintained framework to build a product on, or if you need the LoRA and serving notebooks to be fully runnable on the free tier, since the README describes those two as concept-first with optional GPU appendices. Before starting, verify that a Groq key works from Colab secrets and that the notebook you want is still on the main branch, because the project has published only v0.1.0.

Frequently asked questions

What are AI engineer notebooks?

In this project they are runnable Colab notebooks covering the applied-LLM stack for AI Engineer and Forward Deployed Engineer roles, from model APIs and RAG through evals, agents, adaptation and serving. The README describes them as framework-free and built on raw API calls rather than a wrapper library.

What are engineering notebooks used for?

In this repository they are used to build working systems on top of foundation models section by section, with each notebook installing its own dependencies, reading API keys from Colab secrets, and ending with exercises. The README recommends working through them top to bottom.

What does an AI engineer do exactly?

The repository does not define the role in general terms, but its scope shows the working surface: calling model APIs, constraining output, building retrieval, writing evals, designing tools and agents, and handling prompt injection and operations. The README frames the target reader as someone who can already ship production code and wants the applied-model layer on top.

What topics are covered in the AI engineer notebooks track?

The top-level directories run from environment setup through model APIs, evals, RAG, agents, adaptation, security, operations, serving and inference, ML system design, customer craft, and a case-studies-and-capstone section. LoRA fine-tuning and self-hosted serving are concept-first with optional Colab-GPU appendices.

Official sources

  1. calmrocks/ai-engineer-notebooks on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/calmrocks-ai-engineer-notebooks.svg)](https://hysenlabs.com/projects/calmrocks-ai-engineer-notebooks)