TinyTroupe: Microsoft's LLM persona simulation library, and what it is not
LLM-powered multiagent persona simulation for imagination enhancement and business insights.
At a glance
- What is it?
- TinyTroupe simulates people, not tasks. It builds TinyPerson agents inside TinyWorld environments so you can run focus groups, ad reactions and synthetic data generation against personas you define, and it is still an experimental research project with an unstable API.
- Who is it for?
- Adopt TinyTroupe if you need simulated audiences, persona-driven feedback or synthetic data and you can absorb an unstable API and per-run LLM cost. Do not adopt it if you need a stable interface, deterministic output, or a system that supports real users rather than modelling them.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 89 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem TinyTroupe solves, and the audience it assumes
Most LLM tooling is built to do work for a person: answer the question, draft the email, call the API. TinyTroupe inverts that. The README states the focus is on understanding human behavior and not on directly supporting it, which is why the library contains mechanisms that only make sense in a simulation. A TinyPerson is not an assistant. It has a personality, interests and goals, it can listen to you and to other TinyPersons, and it can reply.
The audience follows from that. The README lists advertisement evaluation, software testing input, synthetic data generation, product and project management feedback, and simulated focus groups. In each case the user is someone who needs a plausible reaction before spending real money or real user attention: a market researcher testing a concept, a product manager reading a proposal back through the eyes of a physician or a lawyer, a team that needs synthetic data to train or to analyse. The library is Python, requires Python 3.10 or later per pyproject.toml, and ships its examples as Jupyter notebooks, so the expected workflow is exploratory rather than production-service shaped.
One boundary is stated plainly. The README says TinyTroupe is for research and simulation only and that you are fully responsible for any use of the generated outputs. That is not boilerplate buried at the bottom; it is repeated in a caution block near the top and backed by a full legal disclaimer section. If you need agents that act on the world, this is the wrong library.
How TinyWorld, TinyPerson and the LLM client fit together
The architecture visible in the repository is a three-layer simulation. TinyPerson objects carry persona definitions, which the examples build either inline or from fragments (the Political Compass notebook is explicitly about customizing agents with fragments). TinyWorld objects are the environments those personas inhabit and interact within. The LLM sits underneath as the generator of behavior, reached through the OpenAI client, with Azure OpenAI as an alternative per .env.example.
What makes this more than a chat wrapper is the state the simulation keeps. Agents listen, reply, and go about their lives in the simulated environment, and the library tracks what happened. Release 0.6.0 added cost tracking utilities at client, environment and agent levels, which tells you the maintainers expect long multi-agent runs where token spend is a real variable rather than an afterthought. The same release introduced AgentChatJupyterWidget for interactive conversations with agents inside notebooks, and SimulationExperimentEmpiricalValidator, which compares simulation results against real-world empirical data using t-test and KS-test. That validator is the most interesting design decision in the project: it treats a simulation as a hypothesis to be tested, not as an output to be believed.
Caching has also changed shape. Release 0.7.0 moved LLM API caching from pickle to JSON. If you have cached artifacts from earlier versions, expect them not to carry over.
The model layer is configurable and has moved repeatedly. Defaults went from GPT-4o-mini to GPT-4.1-mini in 0.5.2, then to gpt-5-mini in 0.6.0, with gpt-4.1-mini and gpt-4o-mini still supported. The 0.6.0 notes warn that the GPT-5 series uses different parameters than the GPT-4 series, so config.ini settings may need adjusting. Treat the model choice as part of your experiment design, not as a constant.
Installing TinyTroupe and running a first simulation
The repository does not publish to PyPI under a documented install command; it ships Windows batch scripts at the top level for building and installing from the repo. On Linux or macOS you would follow the same path through pip against a local checkout. The scripts listed are install_package_from_repo.bat, reinstall_package_from_repo.bat, uninstall_package.bat, build_package.bat and build_and_install_package_from_repo.bat.
python -m venv .venv
. .venv/bin/activate
pip install -e .That installs the package in editable mode from pyproject.toml, which pulls the dependency list including openai >= 1.65, llama-index, pydantic>=2.5.0, scipy, pandas and jupyter. Expect a heavy install; llama-index and the Jupyter stack dominate it.
Next, credentials. Copy .env.example and fill in either the OpenAI key or the Azure OpenAI pair. Entra ID authentication is supported by leaving the API key unset.
cp .env.example .env
# then edit .env and set one of:
# OPENAI_API_KEY=...
# or AZURE_OPENAI_API_KEY=... and AZURE_OPENAI_ENDPOINT=...Configuration lives in config.ini at the repository root. The 0.6.0 release notes are explicit that GPT-5 parameters differ from GPT-4 parameters and that this file is where you adjust them, so open it before your first run rather than after your first confusing result.
For a first real use, the repository's own entry point is the examples directory. Simple Chat.ipynb is the smallest starting point; Interview with Customer.ipynb and the Bottled Gazpacho Market Research notebooks show the focus-group pattern at increasing length. Open one in Jupyter, confirm the kernel is the environment you just created, and run the cells in order. What you should see is agent turns printed as the simulation proceeds, with the notebook driving the world rather than you typing into a chat box. If the first cell fails on an import, the editable install did not take; reinstall_package_from_repo.bat exists for exactly that loop on Windows.
The API is explicitly unstable, and that is the main adoption risk
The README carries a work-in-progress note that says the API is still subject to frequent changes and that TinyTroupe is under very significant development requiring further tidying up. This is not a hedge in a changelog; it is a statement that upgrading can break your notebooks. The release history supports the warning. Between 0.5.2 and 0.6.0 the default model changed twice in effect, new classes appeared (SimulationExperimentEmpiricalValidator, AgentChatJupyterWidget, cost tracking), and Ollama support arrived as experimental and limited. Between 0.6.0 and 0.7.0 the cache format changed from pickle to JSON.
There is a second, quieter failure mode: model drift. The 0.5.2 notes say GPT-4.1-mini can have significant differences in behavior with respect to the previous default of GPT-4o-mini and ask you to retest important scenarios. The 0.6.0 notes repeat the request for the GPT-5 transition. A simulation whose behavior depends on a hosted model you do not control is not reproducible in the way a fixed-seed test is. If your use case requires a stable, auditable result, this library is the wrong tool, and the maintainers say so themselves by labelling it experimental.
The cost profile is the third constraint. Multi-agent simulations multiply LLM calls by the number of agents and turns. The cost tracking utilities added in 0.6.0 exist because that multiplication is easy to underestimate. Local models through Ollama are described as experimental and limited, so they are not yet a way to make long runs free.
TinyTroupe compared with CrewAI and other agent frameworks
CrewAI appears in the searches around this project, and the comparison is worth making precisely because the two look similar from a distance and are not. CrewAI-style frameworks assemble agents that cooperate to complete a task: research this, write that, hand the result to the next agent. The output is work product. TinyTroupe assembles agents that represent people with personalities and goals, and the output is behavior to be observed. The README draws this line itself, contrasting TinyTroupe with game-like LLM simulation approaches and stating an aim toward productivity and business scenarios.
The practical difference shows up in what you measure. In a task framework you check whether the deliverable is correct. In TinyTroupe you check whether the simulated reactions resemble real ones, which is why SimulationExperimentEmpiricalValidator exists and why the repository includes notebooks demonstrating validation against real survey data. If you drop TinyTroupe into a pipeline that expects a deliverable, you will spend your time parsing transcripts. If you drop a task framework into a market research question, you will get a confident answer from a single perspective rather than a distribution of reactions across personas.
The trade-off is maturity. Task-oriented frameworks have settled interfaces because their contracts are simpler. TinyTroupe's interface is still moving because the research question is still open, and the README says the maintainers are looking for feedback and contributions to steer development. Adopting it means accepting that you are an early user of a research artifact.
Licence, maintenance and the cost of upgrading
TinyTroupe is MIT licensed per pyproject.toml and the repository classifiers, which is permissive and places few constraints on reuse. The licence is not the whole story. The README's caution block points to a separate legal disclaimer section and states that various important additional legal considerations apply and constrain its use, so the MIT grant on the code should not be read as a grant covering what you generate or how you use it. Read that section rather than assuming the licence classifier settles the question.
The last push to the default branch was on 2026-07-03, and the most recent release is v0.7.0 from 2026-03-28. The repository is not archived. Development is clearly ongoing, but the README's own work-in-progress note is the better guide to what that means for you: frequent changes, an API still being shaped, and a stated intent to stabilize it over time rather than a promise that it is stable now.
Upgrade cost is concentrated in three places. config.ini, because model parameter conventions have changed between model generations. Your cached LLM responses, because 0.7.0 switched the cache format to JSON. And your own scenario code, because classes and defaults have been added or changed across 0.5.2, 0.6.0 and 0.7.0. Budget for retesting your important scenarios on each upgrade, which is what the release notes ask for.
Editorial conclusion
Adopt TinyTroupe if you need simulated audiences, persona-driven feedback or synthetic data and you can absorb an unstable API and per-run LLM cost. Do not adopt it if you need a stable interface, deterministic output, or a system that supports real users rather than modelling them. Before you commit, verify three things: that your config.ini matches the parameter style of the model you intend to call, that tinytroupe is importable after the reinstall script, and that your empirical validation plan uses SimulationExperimentEmpiricalValidator rather than eyeballing transcripts.
Frequently asked questions
What is AI simulation?
In TinyTroupe's terms it is the use of LLM-generated agents to reproduce human behavior inside a controlled environment, so you can observe reactions without recruiting real participants. The README frames the goal as understanding human behavior rather than supporting it, and the library ships an empirical validator for checking simulated results against real-world data.
How do I install TinyTroupe and run a first example?
Install from the repository rather than a documented package index: the top level ships install_package_from_repo.bat and reinstall_package_from_repo.bat for Windows, and pyproject.toml declares the tinytroupe package with Python 3.10 or later. Then copy .env.example to .env, set the OpenAI or Azure OpenAI credentials, and open a notebook such as Simple Chat.ipynb in the examples directory.
Which LLM does TinyTroupe use by default?
Release 0.6.0 changed the default model to gpt-5-mini, with gpt-4.1-mini and gpt-4o-mini still supported. The same release notes warn that GPT-5 uses different parameters than the GPT-4 series, so config.ini may need adjusting.
Can TinyTroupe run local models instead of OpenAI?
Release 0.6.0 added experimental and limited Ollama support for local models, documented in docs/guides/ollama.md. The release notes describe it as experimental, so it is not presented as a drop-in replacement for the hosted models.
How do I check that a TinyTroupe simulation matches real behavior?
Release 0.6.0 introduced SimulationExperimentEmpiricalValidator, which compares simulation results against real-world empirical data using t-test and KS-test. The repository also includes example notebooks demonstrating empirical validation against real survey data.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/microsoft-tinytroupe)