MCPMark: a stress-test benchmark for MCP tool use
MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.
At a glance
- What is it?
- MCPMark runs agentic models against Notion, GitHub, Filesystem, Postgres and Playwright MCP servers and scores the results with automated verifiers. It is aimed at researchers and engineers who need reproducible numbers rather than anecdotes.
- Who is it for?
- Adopt MCPMark if you need comparable, verifier-backed numbers for a model or an MCP server, and start from the filesystem task because it needs no external accounts. Skip it if you want a general chat or coding benchmark, or if you cannot supply the service credentials and a disposable eval organisation that the GitHub and Notion tasks assume.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 111 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap MCPMark fills: tool-use claims without a verifier
Most statements about how well a model drives MCP tools come from demos or from internal task sets that nobody else can rerun. MCPMark is an evaluation suite for agentic models in real MCP tool environments, and it names five: Notion, GitHub, Filesystem, Postgres and Playwright. The unit of work is a task with strict automated verification, not a transcript a human grades.
The audience is narrow on purpose. The README addresses researchers and engineers and lists what the suite promises: one-command tasks, isolated sandboxes, auto-resume for failures, unified metrics, and aggregated reports. If you are choosing a model for an agent that touches a real Postgres instance or a real repository, that framing matters more than a chat leaderboard does. The tasks are built around practical workflows, and the repository ships them under `tasks/<mcp>/<task_suite>/<category>/<task>/`, so the workload is inspectable rather than hidden behind an API.
One design decision shapes everything else. The README states that MCPMark Verified is now the default task set, that every environment is version-pinned and every verifier stabilized, and that results from earlier task versions are deprecated and not directly comparable. Any number you publish has to carry that label, otherwise you are comparing across task revisions without saying so.
How the pipeline, task suites and verifiers fit together
The entry point is `pipeline.py`, invoked as `python -m pipeline` with flags for the MCP service, the number of runs, the model and the task. MCPMark defaults to a built-in orchestration agent called `MCPMarkAgent`; passing `--agent react` switches to a ReAct-style scaffold, and the README says other settings stay the same. That makes the agent scaffold a variable you can hold constant or change while the tasks stay fixed.
Tasks are organized by service and by suite. `standard` is the default and the README describes it as the full benchmark, 127 tasks at the time of writing. `easy` holds 10 lightweight tasks per MCP and is positioned for smoke tests and CI. Results land under `./results/{exp_name}/{model}__{mcp}/run-*/...` for the standard suite, and under `./results/{exp_name}/{model}__{mcp}-easy/run-*/...` when you pass `--task-suite easy`. Multi-run aggregation is where pass@k and avg@k come from, and the suite also produces aggregated reports.
Two operational details are worth knowing before you plan a long run. Failed tasks auto-retry and resume, which matters when a benchmark takes hours and one flaky browser step would otherwise void the run. And `--compaction-token` summarizes long conversations to avoid context overflow, added because tool-heavy episodes grow past the window. Both are accommodations to the reality of multi-turn tool use rather than features of the scoring itself.
Installing MCPMark and running a filesystem task
Clone the repository and install it in editable mode. If you intend to run browser-based tasks, install the Playwright browsers as well; the README lists both steps under the local install path.
git clone https://github.com/eval-sys/mcpmark.git
cd mcpmark
pip install -e .
playwright installConfiguration lives in a `.mcp_env` file at the repository root. Only the credentials for the service you are testing are needed, so a filesystem run needs a model key and nothing else. The README gives this shape for an OpenAI-compatible endpoint, with the service blocks optional.
OPENAI_BASE_URL="https://api.openai.com/v1"
OPENAI_API_KEY="sk-..."
POSTGRES_HOST="localhost"
POSTGRES_PORT="5432"
POSTGRES_USERNAME="postgres"
POSTGRES_PASSWORD="password"
GITHUB_TOKENS="token1,token2"
GITHUB_EVAL_ORG="your-eval-org"The quickest real run is a filesystem task, which the README notes requires no external accounts. This command runs one model once against a single task.
python -m pipeline \
--mcp filesystem \
--k 1 \
--models gpt-5 \
--tasks file_property/size_classificationAfter it finishes, look under `./results/test-run/gpt-5__filesystem/run-1/` if you used `test-run` as the experiment name. Add `--task-suite easy` to run the lightweight dataset where one is available, which changes the output directory to a `-easy` suffix. Docker is the other supported path: the repository ships a `Dockerfile` and `build-docker.sh`, and the image installs PostgreSQL client tools, Git and the Playwright system libraries in separate layers.
Where MCPMark gets expensive or breaks down
The cost is mostly in credentials and isolation, not in compute. Notion tasks want a source API key, an eval API key and a parent page title; GitHub tasks want a token pool and a dedicated eval organisation; Postgres tasks want a reachable instance. The README frames the sandboxes as isolated environments that do not pollute your accounts or data, but isolation is something you have to construct by supplying those separate credentials. Point the eval keys at a workspace you care about and the guarantee no longer holds.
Version pinning cuts the other way too. The project pinned the GitHub MCP Server at `v0.15.0` and the Notion MCP Server at `@1.9.1`, with the note that Notion released 2.0 but it has many bugs and is not recommended. That is a defensible choice for reproducibility, and it also means MCPMark will not tell you how a newer server behaves. If your question is whether to upgrade an MCP server, this benchmark answers it only after the pinned version moves.
Finally, MCPMark is the wrong tool for general capability questions. It measures whether an agent completes verifiable tasks in these five environments. A model that scores well here may still be poor at long-horizon coding or at open-ended research, and nothing in the suite speaks to that. The README also notes a practical hazard the maintainers had to work around: GitHub @mentions are obfuscated during evaluation to prevent notification spam, which is a reminder that these tasks act on live services.
MCPMark compared with MCP-Universe and ad hoc eval scripts
MCP-Universe is the natural comparison, and it appears in the searches people run around this project. Both evaluate agents against MCP servers rather than against static question sets. The difference that the MCPMark README makes explicit is the Verified task set: environments are version-pinned and verifiers stabilized, and older task versions are declared deprecated and not directly comparable. That is a stance about score comparability over time. If you need numbers that survive a re-run six months later, that stance is the reason to pick MCPMark; if you want breadth across many servers, the five services here are the ceiling.
Against a homegrown script, the trade is control versus maintenance. A script you write can target exactly your production MCP server and your own workflows. MCPMark gives you 127 standard tasks, an easy suite for CI, auto-resume, and aggregation into pass@k and avg@k, at the price of adopting someone else's task definitions and their pinned versions. The repository also accepts agent scaffolds through PRs, so a custom harness can live inside the suite instead of beside it, but that means tracking upstream changes to the task format.
One more comparison is worth noting because the README raises it: a community PR from insforge reports that better MCP servers achieve better results with fewer tokens. If token efficiency under a fixed task set is your question, MCPMark can be used that way, since the same tasks run against different servers.
Maintenance, licence and what an upgrade costs you
The repository is not archived and the last push was on 2026-06-12, so it is a project with recent activity rather than an abandoned one. Releases are tagged: v1.2.0 on 2025-09-20, v1.1.0 on 2025-09-01 and v1.0.1 on 2025-08-29. The README's news entries show a steady stream of changes, including the Verified switch, pinned server versions, `--compaction-token`, the easy task set, ReAct agent support and the @mention obfuscation. A `CHANGELOG.md` sits at the repository root, which is where you should look before pinning a version for a published result.
Upgrading is not free. Because Verified tasks are version-pinned and older task versions are deprecated, moving to a new release can invalidate the numbers you already have. Budget for re-running the suite rather than diffing scores across releases. The Python requirement is `>= 3.11`, and the dependency list includes `litellm==1.80.0` pinned exactly, `openai-agents>=0.2.3,<0.3`, `playwright>=1.43.0` and `notion-client==2.4.0`, so a resolution conflict is plausible in a shared environment. The project also declares pixi platforms for osx-arm64, linux-aarch64, linux-64, win-64 and osx-64, while the README says the suite is fully validated on macOS and Linux.
The licence is Apache-2.0, which permits commercial and modified use with the usual attribution and notice conditions. That is a statement about the licence text, not advice about your situation; if you plan to redistribute a modified harness or publish derived task sets, read the `LICENSE` file and the `NOTICE` handling yourself.
Editorial conclusion
Adopt MCPMark if you need comparable, verifier-backed numbers for a model or an MCP server, and start from the filesystem task because it needs no external accounts. Skip it if you want a general chat or coding benchmark, or if you cannot supply the service credentials and a disposable eval organisation that the GitHub and Notion tasks assume. Before reporting anything, confirm which task version you ran, because the README states that results from earlier task versions are deprecated and not directly comparable, and report the run as MCPMark Verified.
Frequently asked questions
What does MCP stand for?
The README does not spell out the expansion. It describes MCP in terms of servers and tool environments, naming Notion, GitHub, Filesystem, Postgres and Playwright as the services MCPMark evaluates against.
What is a MCP certificate?
MCPMark does not issue certificates. It is an evaluation suite that runs agentic models against MCP tool environments and scores them with automated verifiers, and its README asks that new numbers be reported as MCPMark Verified.
What is a MCP in finance?
That is outside what the project covers. MCPMark's environments are Notion, GitHub, Filesystem, Postgres and Playwright, and none of them is a finance workload.
What is an MCP in construction?
MCPMark has nothing to do with construction. It is a Python benchmark repository for evaluating models and agents on MCP tool use, licensed Apache-2.0 and requiring Python 3.11 or later.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/eval-sys-mcpmark)