Model or dataset
generative-computing/mellea avatar
generative-computing/mellea

Mellea: typed Python functions as the unit of an LLM call

Mellea is a library for writing generative programs.

1,818 stars162 forksPythonApache-2.0

At a glance

What is it?
Mellea is an Apache-2.0 Python library from IBM Research that turns type-annotated functions into structured LLM calls with requirements and automatic retries. It is a good fit when you need parseable output from a model you do not control, and the wrong tool when a single prompt already works.
Who is it for?
Adopt Mellea if you are already writing Pydantic models and want the model call to be validated against them, with retries when a natural-language requirement is not met, and you are on Python 3.11 or newer. Do not adopt it for a one-off prompt where the output is read by a human, or if you cannot accept a validation step that re-invokes the model and therefore costs more tokens.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The LLM call is the part of the pipeline with no type

Most Python pipelines are typed everywhere except at the boundary where text goes to a model and text comes back. Everything downstream then has to treat that return value as a string, which means a parser, a fallback branch, and a test that only sometimes passes. Mellea's claim is narrow: make that boundary a typed function call, so the same tooling you use for the rest of the program applies to it.

The library is aimed at people building agent-style workflows and retrieval pipelines who have already been burned by silent failures. The README frames the problem directly, saying that inside every AI-powered pipeline the unreliable part is the same, the LLM call itself, and lists silent failures, untestable outputs and no guarantees as the symptoms. That is a fair description of the failure mode, and it explains the design: the return value of a generative function is a Pydantic model, not a string you have to parse.

It is not a framework for building chat interfaces, and it does not try to be an agent runtime with a planner and a memory store. The README lists agents among the topics and the examples directory includes agent code, but the library's own primitives are generation, validation and repair.

How @generative turns a docstring into a schema

The mechanism is the decorator. A function annotated with @generative takes a session as its first argument, and its remaining parameters and return type are read by the library. According to the README, docstrings become prompts and type hints become schemas, so there are no templates and no parsers to maintain. When the return annotation is a Pydantic model, that model is what the caller receives.

The session object is the second half. start_session() is described as the convenience entry point that returns a MelleaSession with defaults you can override. The session is what holds the backend connection, so the same generative function can in principle run against a different model by changing how the session was created rather than by editing the function.

On top of that sit requirements. The README says natural-language requirements can be attached to any call, and that Mellea validates and retries automatically. This is the part worth understanding before adopting: a requirement is checked after generation, and a failed check leads to another generation attempt. That is a loop with a cost, not a guarantee. Sampling strategies are the third layer, where a generation runs multiple times and the best result is picked, with rejection sampling and majority voting named as options that are swapped with a parameter change.

Backends are pluggable. The README lists Ollama, OpenAI, HuggingFace, WatsonX, LiteLLM and Bedrock. The optional dependency groups in pyproject.toml split along the same lines, with hf, litellm and watsonx as separate extras, which tells you the base install does not carry the heavy machine-learning dependencies.

Installing Mellea and extracting a typed record

The README gives one install command. The project uses uv, and the base package is installed from PyPI:

bash
uv pip install mellea

If you need every optional feature, the README points to the installation docs and gives the extras form, which is worth quoting because it is easy to get wrong:

bash
uv pip install 'mellea[all]'

For an install from a source checkout, the README defers to CONTRIBUTING.md rather than spelling out the steps, so that file is where to look if you are working against main.

The first real use is the example from the README. A Pydantic model defines the shape, the decorator marks the function, and the session provides the backend:

python
from pydantic import BaseModel
from mellea import generative, start_session

class UserProfile(BaseModel):
    name: str
    age: int

@generative
def extract_user(text: str) -> UserProfile:
    """Extract the user's name and age from the text."""

m = start_session()
user = extract_user(m, text="User log 42: Alice is 31 years old.")
print(user.name)  # Alice
print(user.age)   # 31

What you should see is a UserProfile instance rather than a string. The README notes that age is always an int, guaranteed by the schema, which is the whole point: the parsing step has moved inside the library. Note that the session is passed as the first positional argument to the decorated function, and that the docstring is doing real work as the instruction to the model, so it should read like a precise task description rather than a comment for other developers.

Requirements and retries are a loop, not a guarantee

The most important limitation is in the word the README chooses. Requirements are validated and the call is retried automatically. A retry can fail again. A schema constrains the shape of the output, but a requirement written in natural language is checked by some mechanism the README does not describe in detail, and there is no claim that a requirement is always satisfiable. If your requirement is impossible for the model to meet, the sensible outcomes are a bounded number of attempts and then an error, or a loop that burns tokens until something stops it. The README does not state which, and it does not document a retry budget or a rollback path. That is the first thing to verify in the docs before you put this in a request path.

The second limitation is cost. Structured output, validation and retries all add model calls relative to a single prompt. Sampling strategies multiply that again, since running a generation several times to pick the best result is by definition more than one call. Majority voting also pulls in dependencies: pyproject.toml lists math_verify, rouge_score and nltk as required by majority voting sampling strategies and granite citation parsing, so those packages arrive with the base install whether or not you use those features.

The third is scope. If your task is to summarize a paragraph for a human to read, a typed return value buys you nothing and the validation step is pure overhead. Mellea is for the case where the output feeds another program. The README's own framing, replacing brittle prompts and flaky agents with structured, testable AI workflows, is a statement about pipelines, not about chat.

Mellea compared with calling the OpenAI SDK directly

The obvious alternative is the provider SDK you already have. The openai package is a direct dependency of Mellea, so this is not an either-or at the transport level: Mellea sits above a client library and adds the typed layer. If you call the SDK directly and want a Pydantic object back, you write the schema, pass it as a response format, parse the result, handle the case where parsing fails, and decide whether to try again. That is a small amount of code, and for one endpoint it is less code than learning a new abstraction.

The difference shows up when there is more than one endpoint and more than one backend. With the SDK you re-implement the retry and validation logic per call site, and swapping from a hosted API to a local Ollama model means rewriting the call. Mellea's session is the seam: the same decorated function runs against a different backend, and the README lists six of them. The requirements layer is the other difference, since a natural-language constraint attached to a call has no direct equivalent in a provider SDK's response format parameter.

A second alternative is LiteLLM, which Mellea lists as one of its backends and as an optional extra. LiteLLM's focus is a uniform interface across providers. Mellea's focus is what happens to the output after the call. They overlap on backend portability and differ on everything else, which is why using them together is a reasonable configuration rather than a redundant one.

Maintenance, licence and the cost of upgrading

The repository is not archived. The last push was on 2026-09-10, eight days before this writing, and the most recent release in the list is v0.7.0 from 2026-07-13, with v0.6.0 in May and v0.5.0 earlier the same month. The pyproject.toml on main declares version 0.8.0.dev0, so development is happening between releases rather than only at tagged points.

The version numbers matter for adoption. Before 1.0, minor releases are where breaking changes land, and the gap between 0.5.0 and 0.6.0 was about two weeks while 0.6.0 to 0.7.0 was roughly two months. A pinned dependency is the safe posture, and CHANGELOG.md is the file to read before bumping it, since the README does not describe a deprecation policy.

The licence is Apache-2.0, stated in the README and in the classifier in pyproject.toml. Apache-2.0 includes an express patent grant, which is usually the reason teams pick it over MIT for infrastructure code, and it permits commercial use. That is a description of the licence text, not advice about your situation; if you redistribute Mellea inside a product, the NOTICE and attribution requirements are the parts to read with whoever handles licensing. One detail worth noting from pyproject.toml: the litellm extra carries a comment tying a langchain-core version floor to specific CVE identifiers, so the dependency pins in this project are not arbitrary and should not be loosened casually.

Editorial conclusion

Adopt Mellea if you are already writing Pydantic models and want the model call to be validated against them, with retries when a natural-language requirement is not met, and you are on Python 3.11 or newer. Do not adopt it for a one-off prompt where the output is read by a human, or if you cannot accept a validation step that re-invokes the model and therefore costs more tokens. Before committing, check the installation page at docs.mellea.ai for the backend you actually use, and read the Requirements and repair section of the docs to see how many retries a failing call performs by default.

Frequently asked questions

What is Mellea AI?

Mellea is a Python library for writing generative programs, started by IBM Research in Cambridge, MA, and licensed under Apache-2.0. Its core idea is that a type-annotated function decorated with @generative becomes a structured LLM call, with docstrings used as prompts and type hints used as schemas.

What is Genai?

The project's README and pyproject.toml do not define this term. It appears only as a search question about the project, so there is nothing here to answer with.

How do I install Mellea in Python?

The README gives uv pip install mellea for the base package, and uv pip install 'mellea[all]' for all extras, with further options on the installation page at docs.mellea.ai. Python 3.11 or newer is required according to pyproject.toml.

Which LLM backends does Mellea support?

The README lists Ollama, OpenAI, HuggingFace, WatsonX, LiteLLM and Bedrock. The optional dependency groups in pyproject.toml split these into hf, litellm and watsonx extras, so the backend you choose determines which extra you install.

Does Mellea work with Ollama?

Yes. Ollama is one of the backends named in the README, and ollama>=0.5.1 is listed as a base dependency in pyproject.toml, so the client is installed with the base package rather than as an extra.

Official sources

  1. generative-computing/mellea on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/generative-computing-mellea.svg)](https://hysenlabs.com/projects/generative-computing-mellea)