VibeSearchBench scores the graph an agent builds, and the best reported triplet F1 is 30.3
🔍 The hardest search benchmark in the wild — vague, multi-turn, proactive. 200 long-horizon tasks with persona-driven progressive disclosure, scored by verifiable schema-free knowledge-graph evaluation. No vibes, just triplet F1.
At a glance
- What is it?
- A 200-task search benchmark where a persona-driven simulator reveals constraints one turn at a time and the agent's output is judged as a knowledge graph rather than compared as text. Runs against an OpenAI-compatible endpoint or a wrapped OpenClaw CLI, with triplet F1 as the headline number.
- Who is it for?
- VibeSearchBench earns a slot in an evaluation stack for one reason: it scores a knowledge graph rather than an answer string, so an agent cannot look right by phrasing. Use it when you are choosing between search agents and care whether they find the right entities and attach the right relations, and be ready for a low number, since the best reported triplet F1 is 30.3.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 136 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Triplet F1 on a graph, so phrasing an answer well earns nothing
Scoring happens on a graph, not on prose. Every task ships a ground-truth knowledge graph, and the agent's answer is expected to produce one. Predicted graphs are matched against ground truth by an LLM-as-judge in two phases: node alignment first, then relation equivalence on the pairs that aligned. The headline number is triplet F1, and the best reported score is 30.3, from Claude Opus 4.6 driving OpenClaw.
That 30.3 is worth pausing on. Under a string-match metric a partial answer scores well. Under triplet F1 it does not, because a graph is only right when the entities are right and the relations between them are right. An agent that identifies the right company and attaches the wrong relationship to it has not moved the score.
Node matching is alias-aware and translation-aware, which is what makes a cross-lingual research query survivable at all. For matched entity pairs, the judge then decides whether two relation strings mean the same thing, so acquired by and was bought by count as one edge while a genuinely different relation does not.
Every match is made by the judge rather than by a string comparison. That choice buys tolerance for phrasing and aliases, and it puts a second model in the measurement path, which is worth keeping in mind when two agents score within a point of each other.
The simulator discloses one constraint per turn, and MODE=direct removes it
VibeSearch is the part that makes the benchmark hard, and it is not the retrieval. A persona-driven user simulator plays the requester and reveals constraints gradually rather than stating them in the first message. Agents are expected to interleave partial results with follow-up questions, and the claim being tested is that this converges in both directions.
Two modes make the split explicit on the command line. MODE=simulated is the default and runs the simulator; MODE=direct skips user simulation entirely. For the GeneralAgent path run.py takes the simulator as its own model, with --user-model and --user-model-url, and the shipped example pins --user-model doubao-seed-2-0-pro against an OpenAI-compatible endpoint. The user model is therefore a separate bill from the agent model, and it is not the model under test.
Four OpenClaw environment variables exist for the same machinery: GATEWAY_PORT, default 18789, plus SOURCE_DIR, IDLE_THRESHOLD and MAX_NUDGE. The last two are listed with no description, and their names point at the two ways a scripted user misbehaves, nudging forever or going quiet. No default is printed for either, and nothing in the repository documents what happens when one trips.
The cost of the design is that a run is not one prompt. --num-samples 4 in the documented example means four independent trajectories per task, each a multi-turn conversation, and the evaluator reports both avg@N across samples and best@N, which is the luckiest one.
Two splits of 100, and four fields per task on Hugging Face
200 tasks, two splits of 100, twenty domains. The pro half covers professional research: literature reviews, market analysis and technical due diligence. The daily half covers shopping, travel and lifestyle search with preferences that evolve mid-conversation, and that second property is the more interesting one, because a shifting preference is precisely what a single static query cannot express.
Each task pairs a vague initial query with a ground-truth knowledge graph, and the vagueness is deliberate rather than sloppy. Real users rarely state full intent upfront, so the opening question is written to be startable but not final: enough to begin searching, not enough to know when the work is done.
Four fields ship per task on Hugging Face. qid is the unique identifier and it also names the trajectory file on disk. question is the full research query with its constraints. user_persona drives the progressive-disclosure simulator, which means a task without a persona cannot be run in simulated mode at all. nodes and triples hold the ground-truth graph the judge compares against.
The leaderboard does not break the 30.3 down by split or by domain, so the daily and pro halves have to be read as equally weighted or not read at all. Anyone using this to rank agents is relying on that assumption.
run_all.sh and run_inference.sh differ only in whether the grader runs
Everything starts from run.py, and the shell scripts in scripts/ are thin wrappers that set environment variables for it.
MODEL_NAME=glm-5.1 VLLM_URL=http://host/v1 bash scripts/run_all.sh
MODEL_NAME=kimi-k2.5 VLLM_URL=http://host/v1 bash scripts/run_inference.sh
MODEL_CONFIG=model_config.yaml MODEL_PROFILE=seed2_0_pro bash scripts/run_all.shThe gap between the first two is whether the grader runs. run_inference.sh is labelled inference only and stops at trajectories; run_all.sh is the full pipeline and continues into evaluation. Both reach the model through VLLM_URL, which the README describes as an OpenAI-compatible chat API, so anything speaking that protocol works and MODEL_NAME is only the route string you want that server to take.
MODEL_CONFIG and MODEL_PROFILE reach for the model_config.yaml file at the repository root, which is how a named configuration gets reused across runs instead of retyped. The top-level entries also hold prompts/, test/, assets/ and a directory named viberesearch_query_synthesis/, which the quick start never mentions and the evaluator does not reference.
Re-scoring a finished run is one variable:
TRAJS_DIR=results/trajs/glm-5.1_custom_serper bash scripts/run_eval.shPointing TRAJS_DIR at an existing trajectory directory runs the judge over trajectories you already have, which is how two graders get compared on identical agent output instead of on two different agent runs.
Two entry points: an OpenAI-compatible URL, or a gateway with tools already built
There are two ways in, and they are not interchangeable.
The GeneralAgent path drives an OpenAI-compatible LLM through multi-step web research using a ToolKit with three verbs: search through Serper, visit as a Serper scrape followed by an LLM summarization, and python in an HTTP sandbox. MULTI_AGENT=1 lets a main agent spawn sub-agents for parallel research, while the default 0 keeps one agent on the whole query.
The OpenClaw path wraps an existing CLI instead of reimplementing those tools, and it needs a gateway already running before you start.
bash scripts/run_openclaw.sh
MODE=direct bash scripts/run_openclaw.sh
DATA_PATH=tasks/my_tasks MODE=simulated OPENCLAW_MODEL=my-model bash scripts/run_openclaw.shDATA_PATH is how you point either path at your own task files instead of the shipped 200. In the Python entry point the same choice is --agent-type general or --agent-type openclaw.
Code layout matches. agent/ holds general_agent.py, openclaw_agent.py, llm.py, prompts.py and toolkit.py, and eval/ holds grader.py, whose GraderClient has OpenAI and Gemini backends, plus evaluator.py. The best leaderboard entry, 30.3, comes from the OpenClaw path, which is a data point worth holding onto: the winning configuration arrives with its own tooling already built rather than using the bundled ToolKit, so a comparison against a GeneralAgent run is not purely a model comparison.
Six of the ten environment variables are preset to something never printed
Ten environment variables carry the configuration, and their defaults tell you what the project treats as settled.
MODEL_NAME defaults to glm-5.1 and TOOL_SET defaults to custom. API_KEY is empty and VLLM_URL has no default at all, which means an inference run fails early rather than silently talking to nothing. MULTI_AGENT is 0 unless you set it to 1.
The remaining six are marked preset instead of being written out: SERPER_API_KEY for web search, SUMMARIZE_URL and SUMMARIZE_MODEL for page summarization, CODE_SANDBOX_URL for the Python tool, and GEMINI_API_KEY and GEMINI_API_URL for the grader. SUMMARIZE_MODEL is the single value spelled out, qwen3-30b-a3b-instruct.
Preset is not the same as free. The repository never prints what those presets resolve to, and it does not document how to inspect or override them one at a time, so a run that works on the author's machine may be reaching a summarizer, a sandbox and a judge that you are not running. That is the first thing to check when a score arrives and you cannot account for it.
SUMMARIZE_MODEL being separate from MODEL_NAME is the detail to notice. Page summarization runs on a different model from the agent by default, so the visit tool carries its own quality ceiling independent of the model you are trying to rank, and a weak summarizer can depress an otherwise strong agent's recall.
builtin needs gpt_oss, and the two dependency lists disagree
TOOL_SET has two values and they differ in what you have to install.
custom is the default and needs nothing exotic: search goes through Serper, visit is a Serper scrape plus an LLM summary, and python runs in an HTTP sandbox reached over CODE_SANDBOX_URL. It is also the only set that gives an agent the ability to run code.
builtin offers search, open and find instead, and it requires the gpt_oss package. That package appears in neither requirements.txt nor the README's own dependency line, so choosing builtin means finding the requirement somewhere the quick start does not point to.
python run.py \
--agent-type general \
--model glm-5.1 \
--vllm-server-url http://host/v1 \
--tool-set custom \
--num-samples 4 \
--grader-type gemini \
--grader-api-url https://... \
--grader-api-key YOUR_KEYThe two dependency lists also disagree with each other. requirements.txt pins openai>=1.0.0, httpx>=0.24.0, aiohttp>=3.8.0, pyyaml>=6.0 and tqdm>=4.60.0. The README's Dependencies block names openai aiohttp httpx tqdm transformers json_repair, which drops pyyaml and adds transformers and json_repair. Nothing states which list is authoritative. json_repair looks like something the evaluator needs in order to parse imperfect model output, so an environment built from the README line alone is a candidate for failing at grading time rather than at startup.
One JSONL line per sample, three JSON files per experiment, no releases
Results land in two trees. Trajectories go to results/trajs/{experiment}/ with one JSONL file per task named {task_id}.jsonl, and every line is one sample:
{"qid": "task_042_...", "sample_idx": 0, "question": "...", "messages": [...], "response": "...", "termination": "answer", ...}Evaluation goes to results/eval/{experiment}/, where {task_id}_sample{N}.json holds one trajectory's node and triplet metrics, item_ratings.json collects every per-item result, and summary.json carries the aggregates. The termination field on each line is what lets a run stopped by a turn budget be scored on the same footing as one that ran to an answer.
Two housekeeping details. No GitHub releases exist for this repository, so there is no version to pin and you track main. And the README ends mid-sentence in its License section, so whatever terms it spells out in prose are not visible in the copy a reader gets; the project is MIT licensed and a LICENSE file sits at the top level next to .gitignore and README.md.
The last push was on 2026-05-28 and the repository is not archived.
Editorial conclusion
VibeSearchBench earns a slot in an evaluation stack for one reason: it scores a knowledge graph rather than an answer string, so an agent cannot look right by phrasing. Use it when you are choosing between search agents and care whether they find the right entities and attach the right relations, and be ready for a low number, since the best reported triplet F1 is 30.3. Do not use it as a smoke test. 200 tasks times a multi-turn simulator times four samples is a long run, and the presets behind SERPER_API_KEY, SUMMARIZE_URL, CODE_SANDBOX_URL and the Gemini grader are not printed anywhere. Before trusting any number from it, resolve those presets and confirm the judge model is not the model you are benchmarking.
Frequently asked questions
What score does VibeSearchBench report for a search agent?
Triplet F1 is the primary metric, produced by an LLM-as-judge in two phases: node alignment against the ground-truth entities, then relation equivalence on the pairs that matched. Precision, recall and F1 are reported at both node and triplet level and aggregated as avg@N and best@N across samples.
How many tasks are in VibeSearchBench and how are they split?
200 tasks across 20 domains, divided into a pro split of 100 for professional research such as literature reviews, market analysis and technical due diligence, and a daily split of 100 for shopping, travel and lifestyle search with evolving preferences.
What does the persona-driven user simulator do in VibeSearchBench?
It plays the requester and reveals constraints progressively instead of stating them upfront, so the agent has to interleave partial results with follow-up questions. MODE=simulated is the default and MODE=direct turns the simulation off entirely.
What does the custom tool set in VibeSearchBench require?
The default custom set covers search through Serper, visit as a Serper scrape plus an LLM summarization, and python in an HTTP sandbox. The builtin set offers search, open and find but requires the gpt_oss package, which is not listed in requirements.txt.
How do I run only the evaluation stage of VibeSearchBench?
Set TRAJS_DIR to a directory of existing trajectories and run scripts/run_eval.sh, for example TRAJS_DIR=results/trajs/glm-5.1_custom_serper bash scripts/run_eval.sh. Per-trajectory metrics, item_ratings.json and summary.json are written under results/eval/{experiment}/.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vibebench-vibesearchbench)