ClawWork: An Economic Benchmark That Makes AI Agents Complete Real Professional Tasks for Pay
"ClawWork: OpenClaw as Your AI Coworker - đź’° $15K earned in 11 Hours"
At a glance
- What is it?
- ClawWork is a Python-based benchmarking framework that places AI agents in an economic simulation where they earn income by completing tasks from the GDPVal dataset, pay for every API token used, and must stay solvent to keep working.
- Who is it for?
- ClawWork is a useful tool for researchers and developers who want to evaluate AI models against real professional tasks rather than academic benchmarks. The economic framing (agents start with $10, pay for tokens, earn from completed work) produces signals that standard accuracy scores do not: cost efficiency, survival under budget pressure, and work quality across diverse professional domains.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 7 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ClawWork Measures and Why It Differs from Standard Benchmarks
Most AI benchmarks measure accuracy on a fixed test set: the model answers a question, the answer is compared to a reference, and a percentage is reported. These metrics are informative but do not capture whether a model can complete real work within cost constraints.
ClawWork takes a different approach. It places an AI agent in an economic simulation with a starting balance of $10. The agent receives tasks from the GDPVal dataset, completing tasks earns income, and every LLM API token the agent uses costs money deducted from that balance. If the balance reaches zero, the agent can no longer operate. The README describes the three evaluation dimensions as work quality, cost efficiency, and economic sustainability.
The GDPVal dataset, referenced in the README as coming from OpenAI's GDPVal index, provides 220 professional tasks spanning 44 economic sectors: manufacturing, finance, healthcare, legal operations, and others. Tasks are priced according to Bureau of Labor Statistics wage data, which ties the income model to real-world economic value rather than arbitrary point scores.
The leaderboard in the README shows that top agents can reach the equivalent of over $2,000 per hour in task earnings before API costs. The differences between models are visible not just in total income but in the ratio of income to cost, which the leaderboard tracks as a pay rate.
The Economic Pressure Model in Detail
The $10 starting balance is not symbolic. It creates real pressure on the agent's decision-making within the simulation. Every search query, every LLM completion, every tool call deducts from the balance. The README states that one bad task or careless search can eliminate the balance.
Release 2026-02-20 improved cost tracking to read token costs directly from API responses, including thinking tokens, rather than estimating them. OpenRouter's reported cost is used directly when available. This change made the economic accounting more accurate and removed a source of optimistic estimation.
The agent faces a daily decision loop in the simulation: work on an income-generating task, or invest time in learning activities that may improve future performance. This mirrors a real economic trade-off between immediate productivity and skill development. The README frames this as a feature that makes the benchmark more representative of sustained work rather than a single-task evaluation.
Agents that fail to manage their token usage effectively or that produce low-quality work (which the scoring system penalizes through reduced payment) will exhaust their balance and stop. Only agents that balance quality, speed, and cost survive the full simulation run.
Running ClawWork Locally
The quickstart in the README requires two terminals. In the first terminal, start the dashboard:
./start_dashboard.shThis starts the backend API and the React frontend. In the second terminal, run the agent:
./run_test_agent.shThen open a browser at http://localhost:3000 to watch the agent work in real time.
The .env.example file documents the required environment variables. The agent and the evaluator can use different API providers. For the agent:
OPENAI_API_KEY=your-api-key-hereThe evaluator uses GPT-5.2 with category-specific scoring rubrics. The README recommends using a real OpenAI API key for evaluation because evaluation requires a specific model that may not be available on all providers.
The setup.py defines the package as livebench with Python 3.10 or later as a requirement. The requirements.txt specifies fastapi, uvicorn, langchain, langgraph, and related packages for the server and agent runtime. Sandbox backends include boxlite (default) and e2b-code-interpreter as a fallback.
The 44-Profession Task Distribution and Quality Scoring
The 220 tasks in the GDPVal dataset cover four broad domains as described in the README: Technology and Engineering, Business and Finance, Healthcare and Social Services, and Legal, Media, and Operations. The 44 specific professions within those domains include roles where AI capability has real economic consequences.
Quality scoring uses GPT-5.2 with category-specific rubrics for each of the 44 GDPVal sectors. The README describes this as rigorous LLM evaluation, meaning the scoring is not a keyword match or a fixed test suite but a model-evaluated judgment on whether the work output meets professional standards for that sector.
The average quality percentage visible in the leaderboard ranges from 37.9% for Qwen3-Max to 66.8% for the ATIC-DEEPSEEK combination. These numbers reflect the fraction of task submissions that meet the scoring threshold for payment, not raw output quality. A submission that scores below the threshold generates no income for that task, making quality directly economic.
The Nanobot integration adds a /clawwork command for on-demand paid tasks within the nanobot gateway framework. The ClawMode wrapper transforms any live Nanobot gateway into an economically-tracked agent.
Architecture: Nanobot, LangChain, and the TrackedProvider
ClawWork's architecture has two main components. The first is the benchmark engine (in the livebench/ directory), which manages task assignment, scoring, and balance tracking. The second is the ClawMode integration layer (clawmode_integration/), which wraps an existing Nanobot gateway with economic tracking.
The flow described in the README comments is: the agent receives a GDPVal task assignment, decides whether to work or learn, executes the decision through Nanobot tools (file, shell, web, message, spawn, cron), submits work for evaluation, and receives payment or deduction. The TrackedProvider component intercepts every LLM call and deducts its cost from the agent's running balance.
The frontend is a React dashboard that visualizes balance changes, task completions, learning progress, and survival metrics. It reads from local files for real-time updates when run locally, and the README notes that the public leaderboard data is periodically synced from the live system.
The requirements.txt lists both boxlite and e2b-code-interpreter as sandbox backends. Boxlite is the default. For web search, tavily-python is listed as the recommended search API.
Limitations and Current Maintenance Status
ClawWork's last push to the repository was on 2026-03-03. That is more than six months before the date of this review. No recent development activity is documented in the repository.
The GDPVal pricing model depends on BLS wage data and the API cost structure of current model providers. Both change over time: BLS wage rates update annually, and API pricing changes with model releases. If the simulation was calibrated to pricing from early 2026, the balance dynamics may not accurately reflect the economics of running newer models.
The evaluator is GPT-5.2 with category-specific rubrics. If that model is updated, deprecated, or changes in behavior, the scoring rubrics may need recalibration to maintain consistency across evaluation runs.
The repository has no GitHub releases. The setup.py version is 1.0.0. There is no changelog or release notes file listed in the repository.
For teams comparing ClawWork to other agent evaluation frameworks: the LiveBench project (separate from ClawWork's internal livebench/ package) evaluates models on questions with answers that update over time to avoid contamination. ClawWork's distinction is the economic simulation layer and the professional task distribution from GDPVal rather than knowledge questions. The two approaches measure different things.
Editorial conclusion
ClawWork is a useful tool for researchers and developers who want to evaluate AI models against real professional tasks rather than academic benchmarks. The economic framing (agents start with $10, pay for tokens, earn from completed work) produces signals that standard accuracy scores do not: cost efficiency, survival under budget pressure, and work quality across diverse professional domains. The framework is built on Nanobot and LangChain, so it integrates with any OpenAI-compatible API provider. The main limitation is the last push date of 2026-03-03, which is more than six months before this review, meaning no recent development activity is documented. Before building on ClawWork for ongoing evaluation, confirm that the GDPVal dataset and the scoring model remain compatible with current API pricing and response formats.
Frequently asked questions
Can AI agents make money with ClawWork?
In the ClawWork simulation, yes. Agents earn income by completing tasks from the GDPVal dataset and the README's leaderboard shows top agents reaching the equivalent of over $2,000 per hour in simulated earnings before API costs. This is a controlled benchmark environment, not a real income source.
Is ClawWork free to run?
The ClawWork repository is MIT licensed and free to run locally. Running the benchmark incurs real API costs for the agent and evaluator model calls; these costs come from your own API provider accounts and are not covered by the project.
What is ClawWork?
ClawWork is a Python-based benchmarking framework that runs AI agents through 220 professional tasks from the GDPVal dataset across 44 economic sectors. Agents start with a $10 balance, pay for every API token used, and earn income by completing quality work. The system measures work quality, cost efficiency, and economic survival.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hkuds-clawwork)