WHartTest: a Django and Vue platform that generates test cases from requirements
WHartTest 是一款AI驱动的测试自动化平台,实现从需求到可执行测试用例的自动化生成与管理,帮助测试团队提升效率与覆盖率。 (WHartTest is an AI-driven test automation platform that automates the generation and management of executable test cases from requirements, helping testing teams improve efficiency and coverage.)
At a glance
- What is it?
- An open source test automation platform wiring LangChain and LangGraph to a RAG knowledge base, with API, UI and functional test management across a six-service monorepo.
- Who is it for?
- WHartTest is worth evaluating if your team already writes tests by hand and wants a first pass generated from requirements documents rather than a blank editor, since the value is in the context it feeds the model, not in the model call itself. It is the wrong tool for a team without a requirements corpus, because the RAG layer has nothing to retrieve and generation degrades to a generic guess.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A monorepo of separate services rather than one test runner
WHartTest is a front and back end separated monorepo. The README describes it as built on Django 5.2 with DRF and a Vue front end, aggregating natural language understanding, knowledge base retrieval and embedding search, then combining LangChain and LangGraph with MCP tool calling to cover the path from a requirement to an executable test case.
The repository tree shows the parts as directories: `WHartTest_Django/`, `WHartTest_Vue/`, `WHartTest_Actuator/`, `WHartTest_MCP/`, `WHartTest_Skills/` and `WHartTest_WeixinPluginHost/`. The Actuator is the UI automation executor, the MCP directory is the tool service, and the Skills directory is an agent skill library.
One discrepancy is worth knowing before you plan a deployment. The README states the monorepo is made of 5 subprojects and then enumerates six of them, listing the Django back end, the Vue front end, the UI automation executor, the MCP tool service, the agent skill library and the online documentation editor. That matters because the tree also contains `WHartTest_WeixinPluginHost/`, a seventh directory not mentioned in either list. Work out which services you actually need rather than assuming the README's count is authoritative.
The project is licensed NOASSERTION with a `LICENSE` file present in the tree, which means GitHub cannot classify it and you will need to read that file yourself before redistribution.
Where the generation quality actually comes from
The feature list is long, but the mechanism underneath the AI claims is the same across all of it: requirements in, knowledge base context in, LangChain and LangGraph orchestrating multiple reasoning rounds, test case out.
The knowledge base section describes the retrieval layer in plain terms. You create a knowledge base, upload documents, split them into chunks, vectorize, then retrieve. You configure an embedding model and a reranker, with global knowledge base settings and a connection test. Product documentation, API documentation and business rules go in as project knowledge, and that becomes the retrieval context for both generation and review.
This is the part that determines output quality, and it is the part a demo will not show you. A requirements-to-test-cases feature with an empty knowledge base produces plausible generic cases. The same feature with your business rules, your API contracts and your previous review comments produces cases that name your actual error codes. Anyone evaluating this platform should spend their trial time uploading real documents rather than trying prompt variations against an empty index.
The README also mentions prompt templates, test case templates and generation strategy configuration, which is the mechanism for codifying team conventions. It also mentions multi-round orchestration for context completion, case optimization and result tracking, which is what LangGraph is doing here: holding state across the sub-steps rather than making one large call.
Review is a separate feature from generation. The requirements management section covers upload, parsing, viewing and online editing, plus requirement splitting, dedicated analysis, context checks and versioned report viewing, aimed at surfacing gaps, ambiguity and missing test concerns in the requirement itself.
Interface automation and the multi-format import layer
Interface testing is treated as asset management rather than as a collection of scripts. The platform holds interface modules, interface definitions, environment variables, global headers, database configuration, functions and tags as managed assets, and then lets you orchestrate cases, create tasks, run them and read result detail and reports.
The AI role here is narrow and concrete. Given an interface definition, request parameters, response structure and business context, it generates test steps, assertions, variable extraction and pre and post scripts. The editing and repair path covers request parameters, headers, environment variables, database assertions, function scripts, assertion rules and dependent variables, and the failure path analyses results, execution logs and response bodies to produce a repair proposal.
The substantial addition in release v2.7.0, published 2026-08-20, is the interface documentation import and export layer. Import covers OpenAPI 3.0, Swagger 2.0, Postman, HAR, curl, Markdown, Insomnia, Apidoc, Apifox, Apipost and YApi. The release notes describe a unified parsing and conversion layer that detects the source format automatically rather than requiring you to specify it, fills in request parameters, request bodies and response examples, and preserves the detailed descriptions in the original document. Export goes the other way to OpenAPI in JSON or YAML, plus Apifox, Apipost and YApi, with JSON and YAML detected automatically on the way out.
The compatibility work described in that release, handling field differences between tools to raise the success rate on heterogeneous documents, is the honest indicator of what such a layer costs. Every exporter disagrees with every other exporter about something, and you will find the ones that matter to you.
UI automation with Redis slot leasing and executor capability matching
The UI automation side has the most engineering detail in the release notes and is the part most likely to matter if you are running browser tests at scale.
v2.7.0 added executor editing directly from the executor list, covering browser type, persistent mode, start timeout and operation timeout, applied immediately without restarting the service. It added visibility of executor capabilities in that list: supported browsers, maximum slot count, occupied slot count, and whether the executor is running inside a container. It added per-run overrides, so a test run can temporarily replace browser type and headless mode through run options rather than requiring a global change.
The interesting piece is a Redis-backed slot lease mechanism for UI automation, supporting slot reservation, reclamation and release, so concurrent runs allocate executor resources more sensibly. That is the answer to a problem every shared browser test fleet hits: two runs that both believe they own the one Chrome instance. The release notes state the lease mechanism ships with complete unit test coverage, which is the right thing to attach to logic that decides who gets which browser.
There is also executor registration and online discovery, where a task automatically resolves the runtime environment and filters for online executors whose capabilities match. WebSocket and HTTP batch execution share this path. The Actuation directory in the tree is where the executor lives, and the release notes mention capturing logs, screenshots, video and traces during runs so a failure can be replayed.
The MCP and LangGraph side adds a chat entry point scoped to project context, remote MCP configuration so an agent can call external tool services, tool approval, system prompts, token usage display and multi-model configuration.
Deploying the stack and the two file-handling limits you will hit
Deployment is Docker Compose based. The root of the tree holds `docker-compose.yml`, `docker-compose.local.yml`, `run_compose.sh`, `ai_install.sh` and a `deploy-scripts/` directory, and the compose file states its own usage at the top: `./run_compose.sh remote` pulls prebuilt images from GitHub Container Registry, and locally building means commenting out the `image` line and enabling the `build` section.
The compose file runs PostgreSQL and Redis. PostgreSQL is `postgres:16-alpine` with `max_connections=500`, a named volume for the data directory, and a `pg_isready` healthcheck on a 10 second interval. Redis is `redis:7-alpine` and is described as the broker and backend for Celery, which tells you the async work in this platform runs through Celery tasks rather than in request threads. Host port mappings are 8919 for PostgreSQL and 8911 for Redis, both nonstandard, so check they are free before you start.
The README documents two environment variables under file management, and they are the ones you will hit first. The README says to copy `.env.example` to `.env` and adjust as needed:
FILE_MANAGEMENT_MAX_FILE_SIZE=104857600
FILE_MANAGEMENT_LLM_MAX_CHARS_PER_FILE=0The first caps a single uploaded file at 104857600 bytes, a 100MB default, which matters because requirements documents and API specs are the input to everything downstream. The second caps the characters parsed from a single file when it is attached to an LLM call, where 0 means no limit. The stated behaviour when a file exceeds the cap is that it errors rather than truncating silently, which is the correct choice, but it means a large document produces a hard failure instead of a partial result and you will want to size it deliberately.
The security warning, and the credentials that ship by default
The README carries a section headed as an important security warning about Skills permissions and deployment safety in v1.4.0 and later. The substance is that the Skills module holds high system execution permissions, and the deployment advice is to run it only on an internal network or a trusted private network. It states plainly that you should never expose the service directly to the public internet, or grant access to unauthenticated or untrusted people. The disclaimer says the project is for study and research, that users bear responsibility for improper deployment such as opening it publicly or skipping authentication, and that the team accepts no liability for data leaks or server compromise resulting from misconfiguration.
That warning is unusually direct, and it should be read as a design fact rather than boilerplate. A skills system whose purpose is to extend what an agent can do is by construction a system that executes things, and the platform also ships Playwright skills that drive a real browser. An agent that can drive a browser and holds filesystem and system permissions is close to a remote code execution surface if it is reachable by anyone.
The second thing to check is `.env.example` itself. It ships `DJANGO_DEBUG=False`, which is the right default, and it ships an administrator account block with `DJANGO_ADMIN_USERNAME=admin`, `[email protected]` and `DJANGO_ADMIN_PASSWORD=admin123456`, with a note that the account is created automatically on first start. It also sets `DJANGO_ALLOWED_HOSTS=*`, which is fine for a trial and wrong for anything else. Change the password and narrow the allowed hosts before the first person outside your team reaches the service.
For the record on licensing and maintenance: the last push to the repository was on 2026-09-18, the default branch is `master`, and releases v2.7.0, v2.7.1 and v2.7.2 all landed in August 2026 with the latter two described only as fixing some issues.
Editorial conclusion
WHartTest is worth evaluating if your team already writes tests by hand and wants a first pass generated from requirements documents rather than a blank editor, since the value is in the context it feeds the model, not in the model call itself. It is the wrong tool for a team without a requirements corpus, because the RAG layer has nothing to retrieve and generation degrades to a generic guess. Before deploying it anywhere, change the default administrator credentials that ship in `.env.example`, keep it off the public internet as the project itself insists, and decide whether the Skills module belongs in your environment at all given the system permissions it holds. Check that the monorepo you clone matches the services your deployment needs, since the README counts five subprojects and then lists six, and confirm which of the two compose files you are meant to use before you wire it into CI.
Frequently asked questions
What does WHartTest generate test cases from?
From requirement documents, business descriptions and knowledge base context. The knowledge base layer accepts uploaded documents, splits them into chunks and vectorizes them for retrieval, so the generated cases draw on product documentation, API documentation and business rules rather than only the requirement text.
What is the tech stack behind WHartTest?
Django 5.2 with DRF for the back end and Vue for the front end, with LangChain and LangGraph orchestrating multi-round reasoning and MCP used for tool calling. It is a front and back end separated monorepo rather than a single application.
Can WHartTest import existing API documentation?
Yes. Release v2.7.0 added multi-format import for OpenAPI 3.0, Swagger 2.0, Postman, HAR, curl, Markdown, Insomnia, Apidoc, Apifox, Apipost and YApi, with the source format detected automatically. Export goes to OpenAPI in JSON or YAML plus Apifox, Apipost and YApi.
Is it safe to deploy WHartTest on a public server?
The project's own security warning says to deploy only on an internal or trusted private network, because the Skills module holds high system execution permissions. It states explicitly that you should not expose the service directly to the public internet or grant unauthenticated access.
How do I run WHartTest locally with Docker?
The compose file documents its own usage: run ./run_compose.sh remote to pull prebuilt images from GitHub Container Registry. For local builds, comment out the image line and enable the build section. The stack runs postgres:16-alpine and redis:7-alpine, with Redis acting as the Celery broker.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mgdaaslab-wharttest)