Model or dataset
takahirom/arbigent avatar
takahirom/arbigent

arbigent: the agent under test fails in ways assertions cannot catch

AI Agent for testing Android, iOS, and Web apps. Get Started in 5 Minutes. Arbigent's intuitive UI and powerful code interface make it accessible to everyone, while its scenario breakdown feature ensures scalability for even the most complex tasks.

651 stars62 forksKotlinApache-2.0

At a glance

What is it?
Arbigent is an Apache-licensed Kotlin framework that tests AI agents driving Android, iOS, web and TV interfaces. Its answer to an agent that opens the wrong app is decomposition into dependent scenarios, its provider support is a module boundary rather than a setting, and its readme carries an impersonation warning naming the maintainer's real accounts.
Who is it for?
arbigent is worth a look if you are shipping an agent that operates an app and need evidence beyond a screenshot, because the decomposition into dependent scenarios is the part that addresses the actual failure mode rather than the symptoms around it. Three cautions.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Kotlin, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The failure mode is an agent that opens the wrong app

The motivation section is unusually honest about what is broken on both sides of the testing problem. Traditional interface testing is brittle in ways that have nothing to do with your code: an A/B test, an updated tutorial, an unexpected dialog, dynamic advertising or user-generated content that changes will all fail a test that was correct yesterday. Agents emerged as an answer to that, and then introduced their own failure mode. Agents often do not work as intended; the readme's examples are an agent opening other applications and clicking the wrong button because the task was too complex. The framework's response is decomposition. A complex goal is broken into smaller scenarios that depend on each other, with login leading into search given as the example, and the framework sits between them as a mediator managing execution flow across the whole set. That is a different architecture from one agent given one long instruction, and it is the reason the project claims scalability for complex tasks.

Three releases inside one morning

The release cadence is worth knowing before you pin anything. Three versions were published on a single day, a few hours apart, and each version number is roughly one higher than the last, so the project is releasing continuously rather than on a schedule. The last commit to the repository came a week after those releases, which means the default branch is already ahead of the newest tag again. The licence is Apache, which is the permissive one and removes most questions about embedding the tool in a commercial pipeline. Taken together with the fact that this is a project iterating on the shape of its own scenario model, the practical reading is that the data format of a scenario is the thing most likely to change. If you keep your scenarios in version control as the readme's YAML-based execution implies, a migration is a diff you can read; if you have built automation around them, expect to spend time on upgrades.

Provider support is a module boundary, not a setting

The second motivation is about why agent testing frameworks have not been adopted, and the three barriers named are worth quoting because they are the framework's design targets. Limited provider support, where a framework is locked to one vendor and excludes the model a company already uses internally. Slow operating system adoption, where support for one platform arrives long after the other. And delayed form factor support, where getting beyond phones takes years. The answer is an extension mechanism borrowed from a familiar place: the interceptor pattern from a well-known HTTP client, applied to the agent's execution rather than to requests. Concretely, that shows up in the repository as one Gradle module per provider, with separate modules for two named vendors, so adding a third is a matter of implementing the interface rather than patching a dispatch table. The web reporting and CLI are also separate modules, which means the framework is assembled from parts rather than being one artifact.

Scenarios are designed in a UI and executed from YAML

The workflow is deliberately split between two audiences. Non-programmers, and the readme names quality assurance engineers specifically, design scenarios visually in a user-friendly interface. Software engineers then execute the saved scenarios programmatically from YAML files, which is what lets the whole thing attach to existing testing infrastructure and a continuous integration pipeline. That division is the reason for the four reference documents sitting at the repository root, covering the command line, the YAML format, the protocol specification and the reusable scenario specification; they are the contract between the interface and the pipeline. A reusable-scenario specification is a stronger commitment than most projects make, because it implies scenarios are meant to be shared between projects rather than kept inside one. If you are evaluating this as a test framework rather than an agent, that is the document to read first, and the interface is a convenience rather than the product.

Image assertions double-check the model and a stuck screen has its own recovery

Two robustness features address the ways an agent run goes wrong at run time. The first is a second opinion on the agent's own decision, implemented by integrating an image-assertion feature from another project the author maintains: the decision is verified with an image-based prompt and the agent is allowed to re-evaluate if the check disagrees. That is a meaningful design choice, because it treats the agent's judgement as a proposal rather than as a verdict, and it uses pixels rather than the UI tree as the independent channel. The second is stuck screen detection: when the agent lands on the same screen repeatedly, the framework notices and prompts it to reconsider its actions instead of letting it loop. Supporting both of these is a data-preparation problem, and the answer given is two-fold. The UI tree is simplified and filtered before it reaches the model, and for interfaces that expose no accessibility information at all, annotated screenshots are provided instead, so the agent is not left with nothing to reason about.

Maestro flows run before the agent takes over

The integration with another test automation system is the feature that will save the most setup time, and it is framed as reuse rather than replacement. Existing Maestro YAML flows can be executed as initialisation methods inside an Arbigent scenario, which means your deterministic test assets become the setup step for an agent run rather than being thrown away. The named use cases are all about state: running a login flow before the agent starts, putting the application into a specific state such as a completed onboarding, executing setup sequences that need precise timing, and generally reusing assets you already maintain. The distinction is important for how you adopt this. Agent runs are non-deterministic, so you cannot use them for setup steps where a later assertion depends on exact state; deterministic flows can. The framework leans on that rather than trying to make the agent reliable enough to do the job itself.

MCP servers are set per project and overridden per scenario

Initial support for the Model Context Protocol is described as extending testing beyond direct interface interaction, and it is configured as a JSON string in the project settings rather than in a separate file. The shape is the familiar one, a map of server names to a command and arguments, with one addition that is the interesting part: an enabled field that lets a project disable a server by default.

json
          {
            "mcpServers": {
              "filesystem": { "command": "npx", "args": ["..."] },
              "github": { "command": "npx", "args": ["..."], "enabled": false }
            }
          }
          

That default-off flag matters for a testing tool, because a project that configures a server once would otherwise run it for every scenario, including the many scenarios that do not need it. On top of it sits a per-scenario override list in the YAML that can flip a server back on, so the effective configuration for a run is the project default with the scenario's exceptions applied. The listed use cases are the ones you would expect from a mobile test tool: installing and launching applications, retrieving debug logs, and inspecting server-side logs with external tools.

The readme warns about an account impersonating the project

The most unusual thing in the documentation is a warning, placed before anything else, about an account that appears to be impersonating the project and is not connected to the author. It names the accounts that are genuinely the creator's so readers can tell the difference. That is worth taking seriously rather than skimming past, because agent testing tools are a plausible target for exactly this kind of impersonation: a fake account can collect install commands from people who are already primed to trust the name. The practical response is narrow. Install from the repository the releases are published from, and check the publisher on any package you fetch. Two other things in the tree suggest a real project underneath the warning: a benchmark directory and a sample test, and three separate tool configuration directories for different assistants, which is a maintenance cost a hobby project carries voluntarily.

Editorial conclusion

arbigent is worth a look if you are shipping an agent that operates an app and need evidence beyond a screenshot, because the decomposition into dependent scenarios is the part that addresses the actual failure mode rather than the symptoms around it. Three cautions. Three releases landed within eight hours of each other, so pin a version and expect the API to move. The default model is an OpenAI one even though an Anthropic module exists in the tree, which is worth checking against your own account constraints before you run a suite. And the project has had an impersonation problem it warns about publicly, so install from the repository named on the release rather than from an account that merely looks like it.

Frequently asked questions

What problem does arbigent solve for agent testing?

Agents often open the wrong application or click the wrong button when given one long complex instruction. Arbigent breaks the task into dependent scenarios, such as login leading into search, and mediates execution flow across them.

How does arbigent support different AI providers?

Through an extension mechanism inspired by the interceptor pattern, implemented as one module per provider in the Gradle build, with separate modules for two named vendors. Adding a provider is a matter of implementing the interface.

Which platforms and form factors does arbigent test?

Mobile and TV, covering iOS, Android, web and TV interfaces, including D-pad navigation for televisions. Form factor support beyond phones is one of the stated motivations for the project.

How does arbigent work with Maestro?

Existing Maestro YAML flows run as initialisation methods inside a scenario, so deterministic setup such as a login flow or a completed onboarding happens before the agent starts, reusing automation you already maintain.

What happens when an agent gets stuck in arbigent?

Stuck screen detection notices the agent landing on the same screen repeatedly and prompts it to reconsider. Separately, decisions are double-checked with an image-based assertion that can send the agent back to re-evaluate.

How are MCP servers configured in arbigent?

As a JSON string in the project settings, where each server can be disabled by default with an enabled field, and overridden per scenario from the YAML so a run uses the project defaults plus that scenario's exceptions.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. takahirom/arbigent on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/takahirom-arbigent.svg)](https://hysenlabs.com/projects/takahirom-arbigent)