Datapizza AI: A Python Framework for GenAI Agents and RAG Pipelines
Build reliable Gen AI solutions without overhead 🍕
At a glance
- What is it?
- Datapizza AI is a Python framework that provides typed, multi-provider LLM clients, a tool-calling agent, RAG pipeline primitives, and built-in OpenTelemetry tracing. It targets engineers who want explicit control over their GenAI application without the abstraction overhead common in larger frameworks.
- Who is it for?
- Datapizza AI suits Python engineers who need a structured, observable base for building LLM-powered agents and RAG pipelines and want to swap providers without rewriting application logic. It is a poor fit for teams already invested in LangChain or LlamaIndex, for projects that need mature community packages and long-tail integrations, or for anything targeting Python versions below 3.10.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 135 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem Datapizza AI addresses
Building LLM applications in Python today often means choosing between large frameworks that abstract away provider details or writing raw API calls by hand. The first path gives you convenience but couples your code to a specific abstraction layer. The second path gives you control but forces you to re-implement tracing, tool-calling conventions, and provider switching.
Datapizza AI positions itself between those two points. The README describes it as a framework designed for speed with less abstraction and more control, with an API-first design that keeps provider differences behind a common interface. The core claim is that swapping from OpenAI to Anthropic or Gemini should not require rewriting business logic, only changing which client class you instantiate.
The target user is a Python engineer shipping an agent or RAG application to production who wants to observe what the LLM is doing at each step without integrating a separate tracing stack.
Architecture: core package plus optional sub-packages
The repository is organized as a Python workspace with multiple packages. The base install, datapizza-ai, depends on datapizza-ai-core and brings in the OpenAI client, an OpenAI embedder, and a Qdrant vectorstore as default components. Other providers and modules are installed separately:
pip install datapizza-ai
pip install datapizza-ai-clients-openai
pip install datapizza-ai-clients-google
pip install datapizza-ai-clients-anthropicThe pyproject.toml lists workspace members covering clients for OpenAI, Google, Anthropic, Mistral, Azure, and Bedrock; embedders for Cohere, Mistral, Google, and FastEmbed; rerankers for Cohere and Together; and vectorstores (the Qdrant member is visible, the Milvus member is excluded from the default test run according to the Makefile).
This modular design means a small project can install only the pieces it needs. The trade-off is that the version constraints between sub-packages must be managed manually if you assemble a custom set, and not all workspace members appear to have public PyPI releases yet.
Building agents with tools and multi-agent systems
The agent abstraction wraps a client with a set of tools and a system prompt. Defining a tool uses the @tool decorator, which infers the function signature for the LLM automatically:
from datapizza.agents import Agent
from datapizza.clients.openai import OpenAIClient
from datapizza.tools import tool
@tool
def get_weather(city: str) -> str:
return f"The weather in {city} is sunny"
client = OpenAIClient(api_key="YOUR_API_KEY")
agent = Agent(name="assistant", client=client, tools = [get_weather])
response = agent.run("What is the weather in Rome?")Multi-agent systems follow the same pattern: instantiate multiple Agent objects with different names, system prompts, and tool sets, then wire them together in application code. The README shows a trip-planning example where separate agents handle weather queries and web search, though the README does not show the full orchestration code for that example.
For web search, a DuckDuckGo tool is available as a separate package:
pip install datapizza-ai-tools-duckduckgoOnce installed, DuckDuckGoSearchTool() is passed into an agent's tools list the same way as a custom @tool function.
OpenTelemetry tracing and the ContextTracing API
The README gives observability its own section, describing it as a key requirement for principled development of LLM applications. Datapizza AI instruments agent runs with OpenTelemetry and exposes a ContextTracing context manager:
from datapizza.tracing import ContextTracing
with ContextTracing().trace("my_ai_operation"):
response = agent.run("Tell me some news about Bitcoin")The trace output shows a span summary with total span count. The README also mentions client I/O tracing as an optional toggle that logs inputs, outputs, and in-memory context, and custom spans for tracing sub-steps within an operation.
Using OpenTelemetry as the standard means the trace data can be exported to any compatible backend, though the README does not document which exporters are configured by default or how to point traces at a specific endpoint. That integration detail requires consulting the documentation at docs.datapizza.ai.
RAG pipeline construction with DagPipeline
For retrieval-augmented generation, Datapizza AI provides a DagPipeline that accepts named modules chained together. A typical RAG flow involves document parsing, chunking, embedding, vectorstore ingestion, and then a query path that rewrites the user question, retrieves chunks, and passes them to the LLM.
Document ingestion uses an optional parser package. The Docling-based parser is installed separately:
pip install datapizza-ai-parsers-doclingThe ingestion pipeline code in the README combines QdrantVectorstore, ChunkEmbedder, OpenAIEmbedder, DoclingParser, and NodeSplitter, though the README does not reproduce the full pipeline code. The Makefile excludes the Docling parser from the default test run with --ignore=datapizza-ai-modules/parsers/docling, which suggests that component has additional system dependencies.
On the query side, a ToolRewriter module rewrites user queries before retrieval. The DagPipeline.add_module call names each step so the execution graph can be traced at the module level.
Limitations and cases where it is the wrong tool
Datapizza AI requires Python 3.10 or later. Projects locked to older Python versions cannot use it.
The framework is at version 0.1.0, and the version history visible in the releases shows rapid iteration (0.0.7 in October 2025, 0.0.9 in November 2025, 0.1.0 in March 2026). Early minor versions carry a higher risk of breaking changes between releases. The README does not document a migration guide between the pre-1.0 versions, and the pyproject.toml shows datapizza-ai-core pinned to a narrow range (>=0.1.0,<0.2.0), which means a future core bump will require an explicit update.
The optional sub-package ecosystem is useful when it covers your needs, but it also creates a dependency management burden. Not every workspace member visible in pyproject.toml is confirmed available on PyPI, and third-party integrations that exist in LangChain or LlamaIndex as community plugins do not exist here yet. Teams that need long-tail integrations, such as specific vector database clients or document loaders beyond PDF, will have to implement them as custom @tool functions or contribute to the project.
Comparison with LangChain
LangChain is a general-purpose LLM application framework with a large ecosystem of integrations, a chain-based composition model, and community tooling for evaluation and deployment. It covers more providers and more document loaders out of the box.
Datapizza AI takes a narrower approach. The README explicitly positions it around less abstraction and more control. Where LangChain's chain model wraps LLM calls in multiple layers of composable objects, Datapizza AI exposes the client directly: you call client.invoke() or agent.run() and get a typed result back. The tracing is built into the framework rather than delegated to a separate LangSmith account.
The practical difference is that Datapizza AI is faster to understand for a new engineer on the team but smaller in what it covers today. LangChain is the right choice when community-contributed integrations and examples outweigh the desire for a minimal surface area.
Maintenance and license
The repository is not archived. The last push was on 2026-05-19. The latest release is v0.1.0, published on 2026-03-13. The time between the v0.1.0 release and the last push suggests ongoing development beyond what the release tags capture.
The license is MIT for the main package. The individual sub-packages are separate Python packages and may carry their own license files, which should be checked before bundling them into a commercial product. The README does not mention any paid tier, hosted service, or license restriction on production use.
Editorial conclusion
Datapizza AI suits Python engineers who need a structured, observable base for building LLM-powered agents and RAG pipelines and want to swap providers without rewriting application logic. It is a poor fit for teams already invested in LangChain or LlamaIndex, for projects that need mature community packages and long-tail integrations, or for anything targeting Python versions below 3.10. Before adopting it, check the datapizza-ai-core changelog for breaking changes between 0.0.x and 0.1.0 and confirm that the optional sub-packages you need, such as datapizza-ai-clients-anthropic or datapizza-ai-parsers-docling, are available on PyPI at the versions your environment requires.
Frequently asked questions
Which LLM providers does Datapizza AI support?
The README lists OpenAI, Google Gemini, Anthropic, Mistral, and Azure as supported providers, with a Bedrock client also visible in the pyproject.toml workspace. Each provider requires its own sub-package, installed separately with pip.
Does Datapizza AI support custom tools for agents?
Yes. The @tool decorator converts any Python function into a tool that the Agent class can pass to the LLM. The framework also ships a DuckDuckGoSearchTool and document processing components as optional installable packages.
What Python version does Datapizza AI require?
Datapizza AI requires Python 3.10 or later, as specified in the pyproject.toml requires-python field. Python versions below 3.10 are not supported.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/datapizza-labs-datapizza-ai)