Framework
plurai-ai/intellagent avatar
plurai-ai/intellagent

intellagent simulates a thousand conversations to find an agent's blind spots, and the part that fixes them is behind a paid tier

A framework for comprehensive diagnosis and optimization of agents using simulated, realistic synthetic interactions

1,260 stars153 forksPythonApache-2.0

At a glance

What is it?
intellagent is a Python framework that decomposes a prompt into a policy graph, samples policies the way they co-occur in real conversations, drives a user agent against a chatbot under test, and critiques the transcript. Its dependencies are pinned to exact patch versions, its own README and its package metadata disagree about the minimum Python, and the optimisation layer built on its diagnostics is not in the MIT package.
Who is it for?
intellagent suits a team with a LangChain-era chatbot that needs coverage of awkward conversation shapes rather than a unit test suite, because the policy-graph sampling is a better model of real traffic than hand-written cases, and the visualizer makes the output readable.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The README says Python 3.9 and the metadata requires 3.10

The quick start opens with a requirement, and the build metadata states a different one. The README says intellagent requires `python >= 3.9`. The `requires-python` field in pyproject.toml says `>=3.10`. There is no note explaining the discrepancy, so a reader on 3.9 will find out at install time rather than at read time. The version story has a second mismatch of the same shape: the package declares `version = "0.1.0"`, while the repository's only release is 0.0.1, published on 2025-01-22. That release is also the only one. The default branch took commits as recently as 2026-09-14, roughly twenty months after the sole tagged release, so the package you install from an index is not the code that has been worked on.

Every dependency is pinned to an exact patch version

The dependency list has fourteen entries and not one of them has a range. `google-cloud-aiplatform==1.75.0`, `anthropic[vertex]==0.42.0`, `pandas==2.2.3`, `networkx==3.2.1`, `streamlit==1.39.0`, `plotly==5.24.1`, and six packages from the LangChain family at fixed patches: `langchain==0.3.7`, `langchain_core==0.3.15`, `langchain_openai==0.2.5`, `langchain_community==0.3.5`, `langchain_ollama==0.2.0`, plus `langchain_google_vertexai==2.0.9` and `langchain_google_genai==2.0.7` and `langgraph==0.2.44`. The consequence is that a transitive conflict in any of those trees is a hard failure with no resolution path, and you inherit an Anthropic SDK from early 2025 rather than a maintained one. requirements.txt is the same fourteen entries with `pytest==9.1.1` appended, and that same pytest pin is the entire `dev` extra in pyproject.toml, so the two install paths converge on one frozen set.

Installing the package does not give you the command the docs tell you to run

The packaging configuration includes one package pattern: `include = ["simulator*"]`. The repository root holds `run.py`, `config/`, `tests/`, `docs/`, and `examples/` beside the `simulator/` directory, and the quick start runs the tool as `python run.py --output_path results/education --config_path ./config/config_education.yml`. None of that is inside the installed distribution. So `pip install intellagent` gives you an importable `simulator` package and a pinned dependency tree, and the entry point the documentation walks you through is not something the install produced. The two documented runs are:

bash
python run.py --output_path results/education --config_path ./config/config_education.yml
python run.py --output_path results/airline --config_path ./config/config_airline.yml

The education config runs fast with no database, while the airline config is described as more complex and slower because it has one. A third example environment, retail, exists in the repository and is not in either command.

Azure's jailbreak filter has to be disabled before the simulator runs

One operational instruction sits directly in front of the run commands: if you are using Azure OpenAI for the `llm_intellagent` role, you must disable the default `jailbreak` filter before running the simulator. That is not a workaround for a bug, it is a consequence of what the tool does. The framework's job is to generate edge-case and adversarial scenarios, and a content filter tuned for production traffic will reject the prompts the simulator needs to send. So the documented path asks the user to switch off a safety control on an account that may also serve real users. The README does not say which deployment to isolate that on, or whether the same applies to Vertex, OpenAI, or Anthropic keys.

Ten cents a sample, with a cost limit in the config

The token section is unusually candid. The stated goal is minimising the total cost of running the simulator, and the number given is approximately $0.10 per sample under the default parameters. Two controls come with it: a `cost_limit` parameter in the config file, and `num_samples` in the dataset block, which the example sets to 30. So the default configuration is a few dollars of API spend for one run, before retries. The stated plan for reducing that is not a cheaper model but more data: the team says they are working on leveraging user data to significantly reduce the cost per sample. That is a reasonable direction and it is also a dependency on customer conversations arriving, so the cheap version of this tool may not be the one you get.

Troubleshooting points at a config file the quick start never names

The troubleshooting notes are two lines and they both name `config_default`. Rate limit messages are addressed by decreasing the `num_workers` variables in `config_default`; frequent timeout errors by increasing the `timeout` values in `config_default`. Neither `config_default` nor either of those keys appears in the run commands, which name `config_education.yml` and `config_airline.yml`, nor in the configuration examples, which show an LLM type block and a dataset block. So the two fixes you are most likely to need point at a file the rest of the document does not tell you where to edit. The visualization step is clearer: `streamlit run simulator/visualization/Simulator_Visualizer.py` opens a dashboard, and that path does at least sit inside the packaged `simulator` package.

The diagnostic is open and the optimiser that consumes it is paid

The roadmap is where the business model shows. Beta Release is ticked. Agent platform integrations are an open parent with LangGraph ticked underneath it and CrewAI and AutoGen still open, which reads as one integration shipped inside an item that is not finished. Open items include event generation from existing databases, an API integration for external conversational agents, and personality dimensions for user agents. Then the line that matters: optimizing conversational agent performance using simulator diagnostics is marked as available now with premium access, with three sub-items under it for system prompt optimization, tools optimization, and graph structure optimization. So the framework that produces the diagnosis is Apache-2.0 and the part that acts on it is not. The README also states that basic usage metrics are collected and that nothing identifying you or your company is tracked, with the metrics said to be reviewable in the code.

Editorial conclusion

intellagent suits a team with a LangChain-era chatbot that needs coverage of awkward conversation shapes rather than a unit test suite, because the policy-graph sampling is a better model of real traffic than hand-written cases, and the visualizer makes the output readable. It is a poor fit if you need the framework to be self-contained, because the diagnostic-driven optimisation is offered as a paid tier and because the package installs only the simulator, not the `run.py` the documentation tells you to run. Before you start, fix which file your keys go in, check the Python floor against the one in the metadata, and decide whether a tenth of a dollar per generated conversation fits the evaluation budget you had in mind.

Frequently asked questions

What does intellagent do and how does it work?

It runs in three steps. Given a user prompt plus optional tools and database schema, it decomposes the prompt into a policy graph, samples a subset of policies based on how they co-occur in real conversation distributions, and generates a user-chatbot interaction covering that subset. A user agent then simulates the conversation, and the transcript is critiqued with feedback on the tested policies.

What does intellagent cost to run?

Approximately $0.10 per sample with the default parameters, and the example config sets `num_samples: 30`. Spending is controlled by a `cost_limit` parameter in the config file. The stated plan to reduce the cost per sample is to use customer data rather than to change models.

What Python version does intellagent need?

The README says `python >= 3.9`, while the package metadata declares `requires-python = ">=3.10"`. The build metadata is what an installer enforces. The package also pins all fourteen of its dependencies to exact patch versions, including six LangChain-family packages.

How do I run the intellagent simulator?

Clone the repository, install with `pip install -r requirements.txt`, put your LLM key in `config/llm_env.yml`, then run `python run.py --output_path results/education --config_path ./config/config_education.yml` for the fast no-database case, or the airline config for the slower database-backed one. Visualize the results with `streamlit run simulator/visualization/Simulator_Visualizer.py`.

Does intellagent work with Azure OpenAI without changes?

Not out of the box. The README instructs you to disable the default `jailbreak` content filter on Azure OpenAI before running the simulator, because the tool generates edge-case and adversarial scenarios that the production filter rejects. The documentation does not say whether the same applies to the Vertex, OpenAI, or Anthropic configurations.

Official sources

  1. License: Apache-2.0
  2. plurai-ai/intellagent on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/plurai-ai-intellagent.svg)](https://hysenlabs.com/projects/plurai-ai-intellagent)