Model or dataset
petergpt/bullshit-benchmark avatar
petergpt/bullshit-benchmark

BullshitBench: Measuring Whether AI Models Challenge Nonsensical Prompts

BullshitBench measures whether AI models challenge nonsensical prompts instead of confidently answering them, created by Peter Gostev.

1,882 stars76 forksPythonMIT

At a glance

What is it?
BullshitBench, created by Peter Gostev, tests a specific and underexamined failure mode: whether an AI model confidently answers a question built on an invalid or absurd premise rather than rejecting the premise. Two test suites spanning over 22,000 scored responses make it one of the more detailed public evaluations of this behavior.
Who is it for?
BullshitBench is useful to anyone evaluating AI models for tasks where accepting a false premise causes a real downstream problem: legal drafting, medical triage logic, financial analysis, or code generation from an impossible specification. The project is not a general reasoning benchmark and does not measure how often models incorrectly reject valid questions.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Problem BullshitBench Measures

Most AI benchmarks test whether a model gives the right answer. BullshitBench tests a different question: what does the model do when there is no right answer because the question is built on a false or incoherent premise?

A model that confidently answers a nonsense question is doing something more harmful than producing a wrong answer. It treats the broken premise as valid, potentially reinforcing the user's mistaken assumption and producing output that looks authoritative. The README states the benchmark measures whether models detect nonsense, call it out clearly, and avoid confidently continuing with invalid assumptions.

This failure mode is relevant in domains where the model is expected to catch errors rather than complete tasks. A medical triage assistant that confidently explains how to treat a condition that does not exist is more dangerous than one that says it cannot answer. The same applies to a legal drafting assistant that proceeds from a legally impossible scenario.

V1 and V2: Scope and Coverage

The benchmark has two test suites. V1 contains 55 questions and has accumulated 11,220 individual responses across 204 model and reasoning combinations. V2 is larger: 100 questions, 22,400 responses, and 224 model and reasoning variants.

V2 is organized around 13 nonsense techniques applied across five subject areas: software, finance, legal, medical, and physics. The README notes the benchmark explicitly tests responses to invalid premises and does not measure how often models incorrectly reject valid questions. That is an important scope boundary: a model that rejects everything would score well here while being useless in practice.

As of the repository update on 2026-09-27, tested models include GPT-6 Sol, GPT-6 Sol Pro, GPT-6 Luna, GPT-6 Luna Pro, and Claude Opus 5.5, each evaluated at its lowest and maximum supported reasoning effort in both suites.

Scoring: Three Judges and Three Outcome Categories

Each model response is evaluated by three judge models: Claude Sonnet 4.6, GPT-5.2, and Gemini 3.1 Pro Preview. The three scores are averaged to assign one of three categories.

Clear pushback means the model rejected the broken premise. Partial challenge means the model flagged problems but still engaged with the premise. Accepted nonsense means the model treated the premise as valid and answered accordingly.

The headline metric is the clear score: clear answers divided by attempts minus candidate refusals. Errors remain in the denominator, which prevents a model from improving its score by refusing to answer. The viewer includes a toggle to exclude refusals from the bar charts, and both rates appear in CSV exports.

When a judge fails to produce a valid evaluation, the system retries up to three times. If two valid grades are produced after that, their average is used and the affected response is flagged as using 2/3 judges. Responses where no valid grade is produced remain unscored.

Viewing Results and Running the Dashboard Locally

The results viewer is a static web application included in the repository. No API keys and no build step are required to view existing results. From the repository root:

bash
python3 -m http.server 8795 --bind 127.0.0.1

Then open http://127.0.0.1:8795/ in a browser. The Viewer Guide at viewer/next/README.md covers filters, comparisons, and exports. The dashboard supports filtering by lab, reasoning effort, and domain; sorting by clear pushback rate, least accepted, or average grade; comparing answers side by side; and exporting to PNG or CSV.

The Timeline, Lab trends, Reasoning, and Size views all draw on the same underlying result data. Lab trends follows each lab's best model at each release date.

Running Your Own Evaluations

To collect new model responses and grades, use config.json for the V1 suite or config.v2.json for the V2 suite. The Technical Guide at docs/TECHNICAL.md describes the full process.

Running evaluations requires paid API access to the provider models. The judge models used by the project are Claude Sonnet 4.6, GPT-5.2, and Gemini 3.1 Pro Preview. Running collection and grading for even a subset of the question suite against a new model will incur API costs. The README recommends reviewing the selected models before starting a run.

The data layer uses a manifest system. Each manifest pins the exact question snapshot, metadata, and immutable data files for its release. Full response and grade exports are split into bounded files, and the Technical Guide describes a manifest-aware reader to reconstruct the JSONL exports.

What the Benchmark Does Not Cover

BullshitBench has a precise scope and several things fall outside it. The README states it does not measure how often models incorrectly reject valid questions. A model that becomes more cautious and starts rejecting legitimate requests to improve its BullshitBench score would become less useful, but that cost is invisible in the results.

The benchmark also does not measure factual accuracy on real-world knowledge, reasoning depth on well-formed problems, instruction following, or creative output quality. It is one dimension of model behavior, useful alongside general benchmarks rather than as a replacement for them.

An alternative evaluation that covers related ground is a hallucination benchmark, which measures whether models invent facts when answering real questions. Where a hallucination benchmark asks whether the model's output is true, BullshitBench asks whether the model noticed that the question itself was incoherent. The two behaviors can coexist: a model could both hallucinate on real questions and push back on nonsense, or vice versa.

Data Access, License, and Maintenance

The question files, leaderboard CSVs, and manifests are all available in the repository. V1 data lives in data/latest/ and V2 data in data/v2/latest/. The project is licensed under MIT, with a note that third-party brand assets included in the viewer have separate source and license notes in viewer/next/assets/brands/SOURCES.md.

The repository has no GitHub releases. Development and updates are tracked through the README, which notes the last completed results date. The last push was on 2026-09-26, indicating the project is being actively updated with new model results.

The benchmark is a personal project by Peter Gostev. There is no organization maintaining it, which means its continued updates depend on one maintainer's time and API budget.

Editorial conclusion

BullshitBench is useful to anyone evaluating AI models for tasks where accepting a false premise causes a real downstream problem: legal drafting, medical triage logic, financial analysis, or code generation from an impossible specification. The project is not a general reasoning benchmark and does not measure how often models incorrectly reject valid questions. If your use case requires a model that answers everything without friction, the results here are not a useful signal. The dataset and viewer are free to use under the MIT license; running your own evaluations requires paid API access to the provider models used as judges.

Frequently asked questions

How does BullshitBench score model responses?

A three-judge panel of Claude Sonnet 4.6, GPT-5.2, and Gemini 3.1 Pro Preview evaluates each response and assigns one of three categories: clear pushback, partial challenge, or accepted nonsense. The headline clear score is clear answers divided by attempts minus candidate refusals.

What is the difference between BullshitBench V1 and V2?

V1 has 55 questions and V2 has 100 questions. V2 is organized around 13 specific nonsense techniques applied across software, finance, legal, medical, and physics subject areas.

Can I view BullshitBench results without running any code or providing API keys?

Yes. The repository includes a static results viewer. Running python3 -m http.server 8795 --bind 127.0.0.1 from the repository root and opening http://127.0.0.1:8795/ is sufficient to view the full dashboard with no build step or API keys needed.

Does BullshitBench measure how often models incorrectly reject valid questions?

No. The README states explicitly that the benchmark tests responses to invalid premises and does not measure how often models incorrectly reject valid questions.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/petergpt-bullshit-benchmark.svg)](https://hysenlabs.com/projects/petergpt-bullshit-benchmark)