TestZeus Hercules: a Gherkin-driven testing agent for UI, API and security checks
Hercules is the world’s first open-source testing agent, enabling UI, API, Security, Accessibility, and Visual validations – all without code or maintenance. Automate testing effortlessly and let Hercules handle the heavy lifting! ⚡
At a glance
- What is it?
- Hercules turns plain Gherkin steps into autonomous browser and API test runs using an LLM agent stack. It is worth a look if your team writes BDD features and wants fewer selectors to maintain, but the install pulls in a large dependency tree and the README is thin on rollback and failure handling.
- Who is it for?
- Adopt Hercules if your team already writes Gherkin and you want an agent to drive Playwright without hand-maintained selectors; skip it if you need a deterministic, fully offline test runner or a small dependency footprint.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 58 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Hercules targets: Gherkin steps that no one has to translate into selectors
Most end-to-end suites fail for a boring reason. A button gets a new class name, and a test that passed yesterday cannot find it today. Hercules attacks that by putting an LLM agent between the Gherkin step and the browser. You write the step in the same business language your team already uses, and the agent decides how to carry it out against the live page.
The README describes the project as the first open-source testing agent and lists UI, API, security, accessibility and visual validation as the areas it covers. That scope is the interesting part: the same feature file is meant to drive a browser, hit an endpoint, or run a security check, rather than three separate tools with three separate reporting formats.
The intended user is a QA engineer or a developer who can write Gherkin but does not want to maintain a page-object layer. The README points at enterprise platforms and CI/CD pipelines as the target settings, and it leans on the argument that Gherkin needs no coding skills. That framing is fair for step authoring, but it hides the real work: someone still has to configure a model provider and read the run output.
How the agent actually executes a feature file
The repository layout shows the shape of the system. The Python package lives in testzeus_hercules/, the feature and script inputs sit under opt/, and the dependency list in pyproject.toml names langgraph, langchain-core and langchain-openai alongside playwright. So the runtime is a LangGraph-orchestrated agent loop that calls an OpenAI-compatible model and drives Playwright as its tool.
The Dockerfile confirms the browser side. It installs Playwright browsers with uv run playwright install at image build time, and the system packages it adds (libgbm1, libxkbcommon0, libasound2 and similar) are the shared libraries Chromium needs on a slim Debian base. If you run Hercules outside Docker, those libraries are your responsibility.
Two mechanisms are worth calling out because they change how you write tests. The first is the Python sandbox described in the README, which lets a Gherkin step call a function from a local script file with the Playwright page object already injected. The second is the MCP support visible in the repository files mcp_hercules.example.json and mcp_servers.example.json, which suggests the agent can be extended with external tool servers. The README does not document the MCP wiring in detail, so treat those example files as the starting point rather than a specification.
The sandbox has a tenant model. The README states that SANDBOX_TENANT_ID controls module access, with executor_agent described as full access to requests, pandas, numpy and BeautifulSoup. That is a real design decision: the same feature file can execute arbitrary Python, so the tenant setting is the boundary between a test and a script that can reach the network.
Installing testzeus-hercules from PyPI and running a first feature
The README gives two install paths. The PyPI route is the shorter one and is what most people will try first. Hercules requires Python 3.11 or newer and below 3.14, according to pyproject.toml.
pip install testzeus-herculesAfter the package installs, Playwright itself still needs browsers and their system dependencies. The README gives this as a separate command, and skipping it is the most common reason a first run fails with a missing-browser error.
playwright install --with-depsIf you prefer a container, the repository ships a Dockerfile that builds on python:3.11-slim, installs uv, runs uv sync --frozen --no-dev and then installs the Playwright browsers. The image entrypoint is entrypoint.sh, which the Dockerfile marks executable.
Before any run, the agent needs a model. The repository includes agents_llm_config-example.json.txt as a template, and .env-example for environment variables. The README does not spell out the full key list in the section reproduced here, so open both files and fill them in rather than guessing at variable names.
The README also shows a CLI invocation. The sandbox tenant flag appears in this example, and the input file is a .feature file.
testzeus-hercules --sandbox-tenant-id executor_agent --input-file test.featureWhat you should see is a run against the target application, with results written out. The README references junitparser and junit2html in the dependency list, which points to JUnit XML as the reporting format, and the Makefile's test target writes tests/test_output.xml. The README does not document where the CLI writes its own report by default, so check the run output rather than assuming a path.
A feature step can call into a local Python script. The README gives this example, which passes a filter value into a function defined elsewhere.
And execute the apply_filter function from script at "scripts/apply_filter.py" with filter_type as "Turtle Neck"The script side receives the page object and logger without you importing them, per the README's example, and returns a dictionary the agent can read. That is the escape hatch for anything the agent gets wrong on its own: a selector strategy with fallbacks, written as normal Playwright code.
Where Hercules is the wrong tool
The dependency list is the first honest signal. pyproject.toml pulls in torch, transformers, unstructured[all-docs], numba and a full LangChain stack. That is not a test runner you drop into a small CI container without thinking about image size and cold-start time. If your suite needs to finish in under a minute, an agent loop that calls a model for each step is the wrong shape for the job.
Non-determinism is the second issue. The README's own pitch is that the agent adapts to page changes, and adaptation means the same feature file can take a different path on two runs. For a smoke test that guards a checkout flow, that variability is a liability. Hercules fits exploratory and regression coverage better than it fits a hard release gate, unless you are prepared to inspect failures that come from the model rather than from the application.
Cost and network access are the third. The runtime depends on a hosted model through langchain-openai or the Anthropic client in the dependency list. Air-gapped environments and teams with strict data-egress rules will find that the agent cannot run as shipped. The Python sandbox makes this sharper: the README describes executor_agent as having access to requests, so a feature file can send data outward. The README does not describe an audit trail for sandbox execution, which is worth knowing before you point Hercules at a production-like environment.
Finally, the documentation has gaps. There is a docs/run_guide.md and a docs/Migration/ directory with MIGRATION.md and ARCHITECTURE.md, but the README section reproduced here does not cover rollback, retry policy, or what happens when the model returns an unusable action. Plan to read the source for those answers.
How this differs from Playwright and Selenium suites you write by hand
The obvious alternative is Playwright with pytest, which is what Hercules itself uses underneath. The difference is where the intelligence sits. In a hand-written Playwright suite, you encode the locator and the assertion, and a page change breaks the test loudly at a known line. In Hercules, you encode the intent in Gherkin and the agent chooses the locator at run time.
That trade is real in both directions. Hand-written suites are deterministic, cheap to run thousands of times, and debuggable with a stack trace. Hercules absorbs small UI changes without a code edit, which is exactly the maintenance cost that makes large Selenium estates expensive. But when Hercules fails, the failure is a model decision, and the fix may be a prompt or a config change rather than a line of code.
Selenium is the other comparison the search data keeps raising. Selenium's WebDriver protocol and its grid have a decade of tooling around them, and every CI provider knows how to run them. Hercules is at version 1.0.2 as of 2026-08-03, with 1.0.0 released on 2026-07-13. That is a young release line. If your organisation needs vendor support contracts or a large hiring pool, the ecosystem around Selenium still wins.
A more useful framing: Hercules is not a replacement for a unit test suite or for contract tests. It is a layer above the browser that trades determinism for resilience to UI churn. Teams that already have a stable Playwright suite should treat Hercules as a second layer for flows that change often, not as a migration target.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-08-04, so the project is being worked on. The release cadence is three releases in about three weeks: 1.0.0 on 2026-07-13, 1.0.1 on 2026-07-20, and 1.0.2 on 2026-08-03. That pace is a signal about both activity and churn. Expect to re-read the migration notes in docs/Migration/MIGRATION.md before jumping a minor version.
Hercules is licensed under AGPL-3.0. For internal testing use this is usually uncontroversial, but AGPL's network clause matters if you ever expose a modified Hercules as a service to users outside your organisation. That is a question for your legal team, not something to settle from a README. The repository also carries a LICENSE file and a CodeofConduct.md, and CONTRIBUTING.md for anyone planning to send patches.
Upgrade cost has two components. The Python dependencies are pinned with upper bounds in pyproject.toml, including playwright<=1.49.0, which means a Playwright upgrade waits on Hercules. The model side is the other component: because the agent's behaviour depends on the LLM you configure, a provider-side model change can alter test outcomes without any version bump in your lockfile. Pinning a model version in your config is the practical mitigation, and the example config file in the repository root is where that setting lives.
Editorial conclusion
Adopt Hercules if your team already writes Gherkin and you want an agent to drive Playwright without hand-maintained selectors; skip it if you need a deterministic, fully offline test runner or a small dependency footprint. Before committing, verify the LLM configuration keys in agents_llm_config-example.json.txt against your provider, confirm Python 3.11 to 3.13 is available, and check that the Gherkin steps you rely on are covered in docs/run_guide.md rather than only in the demo videos.
Frequently asked questions
What is TestZeus Hercules?
It is an open-source testing agent from TestZeus that turns Gherkin feature files into automated end-to-end tests, covering UI, API, security, accessibility and visual checks according to the README. It is written in Python and drives Playwright through an LLM agent stack.
How do I install TestZeus Hercules?
The README gives pip install testzeus-hercules as the PyPI route, followed by playwright install --with-deps to fetch the browsers. A Dockerfile is also provided that builds on python:3.11-slim and installs the browsers during the image build.
Which Python versions does TestZeus Hercules support?
The pyproject.toml file declares requires-python as >=3.11,<3.14. Note that the classifiers list Python 3.10 alongside 3.11 and 3.12, which does not match the requires-python bound, so treat 3.11 through 3.13 as the supported range.
Will AI replace Selenium for TestZeus Hercules users?
Hercules does not remove Selenium or Playwright; it drives Playwright from an agent loop, and the README positions it as an alternative to writing selectors by hand. The trade is resilience to UI changes in exchange for deterministic, cheap-to-rerun tests, so hand-written suites still fit stable flows.
Can TestZeus Hercules run custom Python inside a test?
Yes. The README documents a Python sandbox where a Gherkin step calls a function from a local script file, with the Playwright page object and logger injected automatically. SANDBOX_TENANT_ID controls which modules the script can import, with executor_agent granting access to requests, pandas, numpy and BeautifulSoup.
What licence does TestZeus Hercules use?
The repository states AGPL-3.0. That matters if you plan to expose a modified Hercules over a network to users outside your organisation, which is a question for your legal team rather than something the README settles.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/test-zeus-ai-testzeus-hercules)