The Temporal AI Agent demo names its own production limits, and its Compose file asks postgres12 for a postgres14 database
This demo shows a multi-turn conversation with an AI agent running inside a Temporal workflow.
At a glance
- What is it?
- temporal-ai-agent is a multi-turn agent demo that runs the whole conversation inside a Temporal workflow, with goals as files in a directory and tools that include MCP servers. The valuable parts are its own production notes about event history limits and a single workflow ID, plus the packaging details: pyproject says 0.2.0 while the tags are at 0.4.1, and the development container declares a database driver version its database image does not match.
- Who is it for?
- Treat this as the demo its own description calls it. It shows one pattern clearly, an agent loop made durable by putting the conversation in a workflow and the tool calls in activities, and the code is small enough to read in an afternoon, with goals as files, tools in one directory, and MCP servers configured in one module.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The project lists three production problems it has not solved
The most useful section of this repository is the one headed for productionalization, because it is written as a list of things the demo cannot yet do.
The first is event history growth. The recommendation is to store payload data separately, in S3 or a NoSQL database, which is the claim-check pattern, or otherwise garbage collect it. Without that, long conversations fill the workflow's conversation history and start to breach Temporal event history payload limits. This is a hard platform limit rather than a tuning problem, and it lands on exactly the use case the demo is built for, which is a multi-turn conversation that keeps going.
The second is concurrency. A single worker can support many agent workflows at once, but the workflow ID is the same every time, so only one agent actually runs. The stated fix is a distinct workflow ID per agent, for example a UUID or a timestamp, which is one line of work in a system that otherwise needs a routing table.
The third is visibility: the UI should show when an LLM response is being retried, that is, when an activity is retried because the model produced bad output. Retries are invisible today, so a stalled conversation gives the user nothing to look at.
Compose asks for postgres12 while the database service runs postgres:14
The compose file is where this demo stops being a demo, and it has a version mismatch worth knowing before you build on it. The database service is `postgres:14`, while the Temporal auto-setup service is configured with `DB=postgres12`.
Everything else in the compose file is a stock local Temporal stack: the auto-setup image at 1.27.2 publishing 7233, an admin tools container at 1.27, a web UI at 2.37.2 on 8080 with CORS origins limited to `http://localhost:8080`, and three more values all set to `temporal`, being the database user, password and name. Those are the well known local development credentials, fine on a laptop and worth changing the moment a port is reachable from anywhere else.
Two application services sit on top: an API built from the `Dockerfile` and published on 8000, and a worker whose command is `uv run scripts/run_worker.py`. Both read `env_file: .env` and both are told `TEMPORAL_ADDRESS=temporal:7233`, and all services share one network called temporal-network.
The `Dockerfile` is a Python 3.10 slim base with build-essential and curl, uv installed from its own install script, `uv sync --frozen` against the committed `uv.lock`, and a default command that runs only the API. Its comment names worker and train-api as separate compose services, so the service count in the compose excerpt is not the whole picture.
pyproject says 0.2.0 while the tags are at 0.4.1
Two different version numbers live in this repository and they do not match. The project metadata declares `name = "temporal_AI_agent"` and `version = "0.2.0"`, while the release tags reached 0.4.1 on 2025-06-17, after 0.4.0 on 2025-06-09 and 0.3.0 on 2025-06-04.
That gap matters for anything that pins on the package rather than the repository, since nothing in the metadata tells you which tag it corresponds to. The default branch also took a push on 2026-03-27, several months after the newest tag, so the branch carries work that no release describes.
The dependency pins are unusually specific in a demo. `temporalio>=1.8.0,<2`, `fastapi>=0.115.6,<0.116`, `stripe>=11.4.1,<12`, `fastmcp>=2.7.0,<3` and `gtfs-kit>=10.1.1,<11`, which is a transit data library and explains the train goal. The `litellm` requirement is the most telling line, since it excludes two exact versions: `litellm!=1.82.7,!=1.82.8,>=1.70.0,<2`. Something in that pair broke the agent loop, and the pin is the scar.
The build backend is hatchling, and the package list is a directory list: activities, api, goals, models, prompts, shared, tools and workflows. Tests, scripts, frontend, thirdparty and enterprise are not in it.
Two variables configure the model, and the rest of the tools fall back to mocks
The setup claim is that basic configuration needs two values:
LLM_MODEL=openai/gpt-4o # or any other model supported by LiteLLM
LLM_KEY=your-api-key-hereBecause the model is addressed through LiteLLM, the same two variables cover OpenAI, Anthropic Claude, Google Gemini, Deepseek and a local Ollama. The example environment file shows the alternates as commented pairs, for instance an Anthropic model with `LLM_KEY=${ANTHROPIC_API_KEY}` and a Gemini model with `${GOOGLE_API_KEY}`.
The third group is where the demo protects itself. Tool API keys are documented as optional, and each one has a defined behaviour when absent. Without `RAPIDAPI_KEY`, flight search generates realistic mock data. An empty `FOOTBALL_DATA_API_KEY` falls back to a built-in mock fixtures generator. Without `STRIPE_API_KEY`, a mock invoice is created, which is why the key is only needed for the flight invoice goal.
So you can run the whole loop with one LLM key and no third party credentials at all, and the failure mode when you do add keys is a different system failing rather than the agent silently returning fiction. That is the right default for a demo, and it is also a reason not to read demo output as evidence about the tools.
A goal is a file, and one environment variable picks the starting one
The agent's task space is the filesystem. Goals live in the `/goals/` directory organized by category, finance, HR, travel, ecommerce and others, and each goal can reach both native tools and MCP tools.
The starting point is a variable, and the example file names three values with three different intents. `AGENT_GOAL=goal_event_flight_invoice` is the default single agent mode. `AGENT_GOAL=goal_choose_agent_type` switches to the multi goal mode that the page calls experimental, where the user picks an agent type and can switch between types during a conversation. `AGENT_GOAL=goal_match_train_invoice` is described as the replay goal.
Which goals the picker offers is also a variable. `GOAL_CATEGORIES` accepts system, which is always included, plus hr, travel-flights, travel-trains, fin, ecommerce, mcp-integrations and food, and the shipped value is all.
Tools split the same way. Native tools are implemented in the codebase under `/tools/`, and MCP tools are external services reached over the Model Context Protocol, with server configuration in `shared/mcp_config.py`. The worked example is `AGENT_GOAL=goal_food_ordering` with `SHOW_CONFIRM=False`, a goal that calls MCP tools such as Stripe. That flag is the hook for the approval step.
Time-skipping is how the workflow tests stay cheap
Testing a durable agent conversation is slow if the tests wait on real time, so the test commands are built around Temporal's time-skipping environment:
uv sync
uv run pytest
uv run pytest --workflow-environment=time-skippingThe three coverage areas are named: workflow tests for AgentGoalWorkflow signals, queries and state management, activity tests for ToolActivities with the LLM integration mocked and environment configuration covered, and integration tests that run the workflow and its activities end to end. Mocking the LLM is what makes these deterministic, which is the right choice, and it is also why the tests say nothing about model behaviour.
The test tooling is a mix of conventions. Pytest runs with `asyncio_mode = "auto"` and `log_cli_level = "INFO"`, and `norecursedirs` excludes a directory named `vibe`, a leftover that suggests a tool directory not in the current tree. Formatting and linting run through poethepoet rather than npm style scripts, with `poe format` running black and isort, and `poe lint` adding a mypy pass configured with `--check-untyped-defs --namespace-packages`.
There is also a `tests/README.md` and a `docs/testing.md`, so the testing conventions are documented twice at two levels of detail.
The project defines agentic by its failure modes
Rather than asserting that agents are useful, the page enumerates what an agentic framework has to contain, and the list is built around failure handling. Goals made of tools that can execute individual steps. Agent loops that execute an LLM, execute tools and elicit input from an external source such as a human, repeated until the goals are done. Support for tool calls that require input and approval. An LLM used to check human input for relevance before calling the real LLM. An LLM used to summarize and compact the conversation history. Prompt construction assembled from system prompts, conversation history and tool metadata. And, ideally, high durability, which the page says is done in this system with Temporal workflows and activities.
Four of those seven items are about conversation length and untrusted input, which connects directly to the event history limit in the production notes. Compacting history is listed as a feature and large history is listed as a risk, and the difference between the two is the claim-check pattern the demo does not implement.
The Temporal case itself is short: reliability, state management, a code-first approach, built-in observability and easy error handling, with the reasoning deferred to an architecture decisions document and the mechanics to an architecture guide.
An enterprise directory, a .sln file and a frontend the overview never mentions
The tree carries items the page does not discuss, and they are worth a look before you decide how big this project is. There is an `enterprise/` directory, a `thirdparty/` directory, a `frontend/` directory, a `temporal-ai-agent.sln` file, which is a Visual Studio solution file in a Python repository, an `AGENTS.md`, a `Makefile` alongside the poethepoet tasks in `pyproject.toml`, a `.devcontainer/` directory and a `docker-compose.override.yml` sitting next to the main compose file. A `__init__.py` at the repository root suggests it was also used as a package at some point.
None of that is explained. The hatch package list covers eight directories and none of the extras, and the Dockerfile comment mentions a `train-api` compose service that the compose excerpt does not show. So either the repository has moved past its README, or the README was written for a smaller subset of it.
Two more items. The enablement guide for this project is an internal resource for Temporal employees, linked as Google Slides and Google Docs rather than in the repository, so it is not part of what you can read here. And the maintenance picture is dated: newest tag 0.4.1 on 2025-06-17, last push to the default branch on 2026-03-27, three videos on YouTube, one for the interaction and one for multi agent execution, and a note that watching the five minute demo video is the fastest way to understand how the interaction works.
Editorial conclusion
Treat this as the demo its own description calls it. It shows one pattern clearly, an agent loop made durable by putting the conversation in a workflow and the tool calls in activities, and the code is small enough to read in an afternoon, with goals as files, tools in one directory, and MCP servers configured in one module. It is not a production agent, and the repository says so in its own words: long conversations will breach Temporal event history payload limits without a claim-check pattern, the workflow ID is constant so only one agent runs at a time, and the UI does not show when an LLM call is being retried. If you take that pattern anywhere, those three items are your work, not the demo's. On versions, do not trust the package number, since `pyproject.toml` still reads 0.2.0 while the tags reached 0.4.1 on 2025-06-17 and the branch took a push on 2026-03-27. And read the tool configuration before you wire anything real, because the example environment deliberately leaves several keys unset so that flights, football fixtures and invoices come from mock generators.
Frequently asked questions
What does the temporal-ai-agent demo actually demonstrate?
It demonstrates a multi-turn conversation with an AI agent running inside a Temporal workflow, where the agent collects information toward a goal and runs tools along the way. Goals are files in the /goals/ directory organized by category, and the agent uses both native tools implemented in /tools/ and external MCP tool servers configured in shared/mcp_config.py.
How do I choose the model and the goal in temporal-ai-agent?
Two variables cover the model, LLM_MODEL and LLM_KEY, with any model LiteLLM supports, including OpenAI, Anthropic, Google, Deepseek and local Ollama. The starting goal is AGENT_GOAL, whose example values are goal_event_flight_invoice for single agent mode, goal_choose_agent_type for the experimental multi goal mode and goal_match_train_invoice for replay. GOAL_CATEGORIES selects which categories the goal picker lists.
Can temporal-ai-agent run without third party API keys?
Yes for the mockable tools. The example environment documents the fallbacks: without RAPIDAPI_KEY flight search generates realistic mock data, an empty FOOTBALL_DATA_API_KEY uses a built-in mock fixtures generator, and without STRIPE_API_KEY a mock invoice is created. You still need an LLM key, since LLM_MODEL and LLM_KEY are the two required settings.
What are the known production limits of the temporal-ai-agent demo?
The page names three. Long conversations will fill the workflow conversation history and breach Temporal event history payload limits unless payload data is moved out with a claim-check pattern or garbage collected. The workflow ID is the same every time, so only one agent runs at a time unless you use a distinct ID such as a UUID. And the UI does not show when an LLM call is being retried after bad output.
How do I run the tests for temporal-ai-agent?
Run uv sync to install dependencies including test extras, then uv run pytest, or uv run pytest --workflow-environment=time-skipping for the faster time-skipping run. Coverage spans workflow signals, queries and state management, activities with the LLM integration mocked, and end to end integration tests. Formatting and linting run through poethepoet tasks that wrap black, isort and mypy.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/temporal-community-temporal-ai-agent)