Model or dataset
iusztinpaul/designing-real-world-ai-agents-workshop avatar
iusztinpaul/designing-real-world-ai-agents-workshop

designing-real-world-ai-agents-workshop: A Hands-On Build of Two MCP-Served Agent Systems

Hands-on workshop: Build a multi-agent AI system from scratch — Deep Research Agent + Writing Workflow served as MCP servers. Includes code, slides, and video

510 stars140 forksPythonMIT

At a glance

What is it?
iusztinpaul/designing-real-world-ai-agents-workshop is a Python repository from the AI Engineering Conference Europe that walks through building a multi-agent system with two FastMCP servers: a Gemini-grounded Deep Research Agent and a LinkedIn Writing Workflow with an evaluator-optimizer loop.
Who is it for?
This workshop is a practical fit for a Python developer who understands LLM basics and wants to move from prompt experimentation to shipping a multi-agent system with real infrastructure patterns. The code is complete, the three modes let you choose how deeply to engage, and the implement_yourself skeleton provides a structured build path.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 120 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the Workshop Builds and Who It Is For

The workshop builds a two-server multi-agent system where each server exposes tools, resources, and prompts via the Model Context Protocol (MCP). An MCP-compatible harness such as Claude Code or Cursor orchestrates both servers as a single pipeline. The first server handles deep research: it accepts a topic, runs multiple Gemini-grounded searches, analyzes referenced YouTube videos for transcripts, fills gaps with follow-up searches, and compiles a structured research brief (research.md). The second server handles LinkedIn post generation: it reads the research brief, generates a draft post, runs it through an evaluator-optimizer loop, and produces a final post with an AI-generated image.

The intended audience is engineers who have used LLMs in application code but have not yet built a multi-agent system with persistent tools, structured output schemas, and automated quality evaluation. The workshop was presented at AI Engineering Conference Europe. A full recording is available on YouTube and slides are linked from the README.

Deep Research Agent: Gemini Grounding and YouTube Analysis

The research server uses Google Gemini with Google Search grounding to retrieve factual information for a given topic. Grounding means that Gemini's generation is anchored to live search results rather than relying solely on training data. The agent runs multiple search queries in sequence, each targeted at a different angle of the topic.

If the research seed includes YouTube video URLs, the server calls an analyze_youtube_video tool to extract and analyze the transcript. After the initial passes, the agent identifies coverage gaps and runs additional grounded queries to fill them. The final output is a structured research brief in Markdown format. The README shows an example brief covering AI agent architecture spanning roughly 20,000 tokens of compiled content from two queries and one video transcript.

The research flow is: topic seed with key questions and reference links, then deep_research passes, then gap-fill passes, then compile_research. This pattern is documented in the repository as an example of tool-use agents where the LLM decides when to call which tool and how many passes to run based on coverage quality.

LinkedIn Writing Workflow and the Evaluator-Optimizer Loop

The writing server accepts the research.md file and a guideline document, generates a LinkedIn post draft, then runs it through a review-edit cycle. The evaluator step scores the draft against the guideline using an LLM-as-judge approach: a separate LLM call evaluates whether the draft meets quality criteria and returns structured feedback. The optimizer step applies that feedback and regenerates the post. The loop runs N times, where N is configurable.

This pattern, called an evaluator-optimizer loop in the workshop, models the human writing process of drafting and revising. The Opik library handles observability and evaluation scoring. Quality scores at each loop iteration are tracked, giving a visible record of whether each revision cycle improved the output.

The final output is a post.md file and an AI-generated image. The README shows a complete example post produced by this pipeline, including the seed topic, the research brief summary, and the final post text.

Three Modes: Watch, Run, or Implement Yourself

The README presents three ways to engage with the repository. The first is to watch the two-hour YouTube recording and review the slides to build a mental model of the architecture before touching any code. The second is to clone the repository and run the finished code against a real topic, watching the system produce a research brief and a LinkedIn post end to end. The README estimates this takes about 30 minutes.

The third mode is the implement_yourself/ directory, which contains a stripped-down skeleton with 25 pre-groomed implementation tickets and a custom /implement Claude Code skill. The skill orchestrates SWE and Tester agents in a ticket-by-ticket loop until the skeleton matches the reference implementation in src/. The skeleton is a self-contained project: opening the agent harness directly in implement_yourself/ scopes its working directory so the agents cannot read the reference implementation in ../src/.

This three-mode structure means the workshop works for different amounts of available time, from a two-hour overview to a two-to-four-hour full build.

Running the MCP Servers Locally

The project uses uv for Python environment management and requires Python 3.12 or later. Start by copying the environment file and filling in the required API key:

bash
cp .env.example .env

The .env.example shows two keys: GOOGLE_API_KEY (mandatory for the research server) and OPIK_API_KEY (optional, for observability and evaluation scoring). To run the research server:

bash
uv run fastmcp run src/research/server.py

To run the writing server:

bash
uv run fastmcp run src/writing/server.py

Both servers run on stdio transport, which is the MCP transport mode that Claude Code and Cursor use to communicate with local servers. A Streamlit UI is also available for running the full pipeline interactively:

bash
uv run streamlit run streamlit_app.py

The Makefile documents additional commands for testing individual workflows against the included datasets.

Coupling to Google Gemini and the Scope of What Is Built

The research server depends on Google Gemini via the google-genai package. Gemini provides both the generation and the Google Search grounding that makes the research factual rather than hallucinated. There is no documented path to swap in a different search-grounded provider: changing the research server to use a different LLM would require modifying the tool implementations. The GOOGLE_API_KEY is therefore a hard requirement for running the research component.

The writing workflow example is specific to LinkedIn posts. The evaluator-optimizer loop pattern it demonstrates is reusable: the core structure of generate, score, revise, repeat applies to any text generation task with an evaluable quality criterion. But the workshop does not generalize beyond the LinkedIn use case. Teams who want to apply the pattern to a different output type need to adapt the guideline schema and the evaluator prompt.

The implement_yourself path requires Claude Code or a compatible agent harness to drive the /implement skill. Developers using a different tool chain would need to manually work through the 25 tickets rather than having the agent orchestrate them.

Compared to Framework-Driven Agent Starters

Agent frameworks like LangGraph or CrewAI provide declarative abstractions for building multi-agent pipelines: graph nodes, edge conditions, and team configurations map to a DSL that the framework executes. The trade-off is that the framework controls the execution engine and the communication pattern. Debugging an agent that misbehaves requires understanding both your application code and the framework's internal state management.

This workshop takes the opposite approach: both MCP servers are plain FastMCP implementations with explicit Python tool definitions. The MCP protocol handles communication between servers and the harness, but there is no framework-level abstraction managing agent roles or conversation flow. The Coordinator-to-worker routing happens in the harness (Claude Code or Cursor), not in a Python orchestration layer. The result is that every line of agent logic is visible in src/, but building more complex multi-agent topologies requires more manual wiring.

Editorial conclusion

This workshop is a practical fit for a Python developer who understands LLM basics and wants to move from prompt experimentation to shipping a multi-agent system with real infrastructure patterns. The code is complete, the three modes let you choose how deeply to engage, and the implement_yourself skeleton provides a structured build path. The coupling to Google Gemini for the research server means GOOGLE_API_KEY is a hard requirement for that component. The LinkedIn Writing Workflow is a concrete artifact that only applies directly to teams producing LinkedIn content; the underlying evaluator-optimizer pattern is reusable in other contexts, but the workshop does not generalize it beyond that example.

Frequently asked questions

What does the implement_yourself directory in the workshop provide?

It contains a skeleton project with 25 pre-groomed implementation tickets and a custom /implement Claude Code skill that orchestrates SWE and Tester agents ticket by ticket. The agents cannot see the reference implementation in ../src/, so the build is genuine rather than a guided copy.

What API key is required to run the designing-real-world-ai-agents-workshop code?

GOOGLE_API_KEY is mandatory for the Deep Research Agent, which uses Gemini with Google Search grounding. OPIK_API_KEY is optional and enables observability and evaluation scoring through the Opik platform.

What evaluation framework does the workshop use to score agent output quality?

The workshop uses Opik for LLM-as-judge evaluation. The writing server scores each draft post against a guideline document using a separate LLM call, and Opik tracks the quality scores across the evaluator-optimizer loop iterations.

Official sources

  1. Issues
  2. iusztinpaul/designing-real-world-ai-agents-workshop on GitHub
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/iusztinpaul-designing-real-world-ai-agents-workshop.svg)](https://hysenlabs.com/projects/iusztinpaul-designing-real-world-ai-agents-workshop)