# Veridrop: a quarter of the Claude weight rests on one signature, and Gemini gets the least checking

> Veridrop probes AI API relay services to see whether they forward to the model they claim, and scores them against official baselines. It is bilingual, commercially operated, and its evidence is much stronger for one of the three protocols it supports than for the other two.

**canarybyte/veridrop** — AI API 中转站检测工具：Claude 中转站检测、OpenAI 中转站检测、Gemini 中转站检测，中转站真伪检测、长上下文验证、思维签名验证、中转站红黑榜，自托管开源。

- Repository: https://github.com/canarybyte/veridrop
- Website: https://veridrop.org
- Stars: 568 · Forks: 50
- Language: Python
- License: AGPL-3.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/canarybyte-veridrop

## Four names for one detector, and a section explaining which of its two domains is live

The project presents itself under at least four identifiers.

The repository is `canarybyte/veridrop`. The distribution name in the manifest is `relay-detector`, with the description crediting Veridrop by name. Two console scripts are installed, `relay-detector` and `veridrop`, and both point at the same CLI entry. And there are two web domains.

The domain question is handled by a dedicated section headed with currently available URLs, dated 2026-09-20. It names `veridrop.app` as the recommended entry, reachable domestically without a detour, and `veridrop.org` as the main domain or fallback for when the first is restricted. It says domain changes will be announced in that block first, and suggests bookmarking the GitHub project page and treating that block as authoritative.

A tool that needs to publish which of its own two domains currently works is telling you something about the hosting situation, and the mitigation it offers is to read a dated block in a readme.

There is also a mismatch in the metadata. The manifest's project URLs list the homepage as `veridrop.org`, which is the domain the page describes as the fallback, while the recommended entry is the other one.

The page is bilingual. The English section gives the framing and the headline numbers; the Chinese sections carry the per-detector tables and the weighting detail.

## A quarter of the Claude weight rests on a signature the project calls unforgeable

The distinguishing claim is about cryptography, and it applies to one protocol.

When thinking is enabled on the Anthropic Messages API, the response carries a `signature` field of roughly five hundred to two thousand characters. The project describes this as a cryptographic artefact signed by Anthropic with a server-side key, and states that a relay theoretically cannot forge it. It bills this as the only authenticity indicator in the category that is cryptographically verifiable and cannot be bypassed, and gives it twenty-five percent of the Claude protocol's weight as the core detector.

The honesty of the framing is worth noting. The same passage says that OpenAI and Gemini have no equivalent server-side signing mechanism, so verification strength for those two reaches only protocol level and behavioural level. The project does not pretend the three protocols are equally provable.

What remains for the other two is fingerprinting, described as catching a relay that has swapped the underlying chip by looking at residual `usage` fields such as the `claude_cache_creation_*` family.

So the architecture is: one cryptographic check that is genuinely hard to defeat, and a family of statistical and structural checks that are defeatable in principle. The weighting is the design decision that follows from that, and it is why a Claude verdict and an OpenAI verdict from the same tool are not equally strong claims even when they carry the same label.

## OpenAI authenticity is read off leftover field names, and the depth is twelve, eight, and seven

The two weaker protocols are detected in noticeably different ways, and neither is the way you would guess.

For OpenAI Chat Completions the decisive signal is the `usage` object. If a response still carries `claude_cache_creation_*` entries, or a `usage_source` of anthropic, or Anthropic-style field naming such as `input_tokens` and `output_tokens`, the relay is judged critical and the verdict is capped at marginal. In other words, the primary question for an OpenAI relay is whether it is secretly an Anthropic relay, and the evidence is field names the upstream forgot to rewrite.

For Gemini the signal is packaging. The killer detector is described as adaptation for Gemini 3 thinking-by-default behaviour: the genuine service is supposed to carry thinking metadata, and imitation wrappers frequently omit the fields or return a structure that does not match.

The detector counts underline how uneven the coverage is. Claude gets twelve, OpenAI gets eight, Gemini gets seven. The cross-protocol long-context tier, a three-layer needle-in-a-haystack probe that checks whether a relay really honours its advertised context window, is listed as supported for Claude and OpenAI and explicitly not yet supported for Gemini.

So the protocol with no cryptographic signature is also the one with the fewest detectors and the only missing capability probe. Gemini verdicts from this tool are the weakest of the three, and the project's own tables say so.

## The hosted service says the key never touches disk, and the self-hosted path writes it to a file

The key handling claim and the self-hosted instructions are in direct tension.

The hosted service is described as free, requiring no signup, and never persisting API keys. The Chinese version is blunter: the key is not written to disk. That is a property of the hosted deployment, and it is a good one to offer given that what users paste in is a working credential for a paid relay.

The self-hosted path cannot work that way. Installation is four commands, and the configuration step copies the example environment file and edits it, filling in three values.

```bash
# 配置中转站凭据
cp .env.example .env
nano .env  # 填 ANTHROPIC_BASE_URL / ANTHROPIC_API_KEY / ANTHROPIC_MODEL
```

So on your own machine the key is written to `.env` in the working tree, which means it is covered only by whatever `.gitignore` says and by whatever filesystem permissions you happen to have.

Nothing visible claims otherwise, and the distinction is fair. But a reader who takes the no-persistence line from the service description and then follows the self-host instructions is storing a live third-party credential in a plaintext file without being told that this is the trade.

The self-host install also clones over SSH rather than HTTPS, so it assumes you have GitHub SSH keys configured before any of this begins.

## The cost ladder is documented in dollars, from near zero to eight

Probing someone else's relay costs money, because every probe is a request against a metered endpoint. This project publishes what each level costs.

The cheapest command is a connectivity test described as taking seconds at near-zero cost. A full detection run is given as about one minute and roughly twelve cents, using a Haiku-class model. The long-context standard tier is quoted at about five to fifty cents. The extreme tier, which probes proportionally to a model's advertised limit, ranges from about five cents to eight dollars.

That last number is the one to notice. A single authenticity check can cost eight dollars depending on the model being tested, which is a real sum to pay for information rather than output.

The mitigation is a stopping rule. The three probe layers are ordered by depth, and the run stops immediately when any layer fails, explicitly to avoid burning tokens on a relay already shown to be short of its claims. Which is the right design, and it also means the advertised cost range is an upper bound you only reach on a relay that passes everything.

The command surface matches those tiers, and the full run writes JSON that a separate compare step later diffs against baselines found automatically under the data directory.

## Two console scripts, one distribution name, and a default model no documented command uses

Three small inconsistencies sit in the packaging, and each would cost a new user a minute.

The manifest installs two scripts, `relay-detector` and `veridrop`, both bound to the same CLI application. Every documented command uses the first one. The second exists and is never used in the documentation, which is a reasonable compatibility alias and an undocumented one.

The distribution is named `relay-detector` while the project, the domain, the hosted service, the service unit file, and the second console script are all `veridrop`. The manifest description credits Veridrop in its name field, so the name you search for and the name you install under differ.

The third is in the example environment file. It ships a base URL pointing at the official Anthropic API, a key placeholder, and a model of `claude-opus-4-7`. Every documented detection command specifies `claude-haiku-4-5` instead, because the probes are run against whichever model you are testing and the cheap class is the sensible default for that.

So copying the example file as instructed leaves you testing Anthropic itself, at the most expensive model in the list, unless you also edit the model line. The example and the commands disagree, and neither notes the other.

## A paid certification channel exists for the project that scores those same relays

The project is open source under AGPL-3.0-or-later, and it also sells something to the operators of the relays it grades.

The support section offers two routes: a GitHub star, and business cooperation or certification listing, with the commercial route linked from the hosted site. Alongside it is a written assurance that cooperation does not change detection scores, verdicts, critical findings, or the leaderboard algorithm, and a line saying the project's value rests on public evidence rather than human endorsement.

The assurance is more than marketing, because the mechanism behind it is checkable. Weight allocation and sub-check details are said to be in `DESIGN.md`, the scoring logic and report evidence are stated to be public in the repository, and the leaderboard is described as Bayesian-weighted over community-submitted detections. So the claim is that you can recompute a score from the published weights rather than take it on trust.

What the repository cannot show you is whether the hosted service applies those same weights. The open source and the live service are two artefacts, and only one of them is in the tree.

That is the reason to read `DESIGN.md` before trusting a verdict rather than after: the conflict of interest is structural, the mitigation is openness, and openness is only a mitigation if you actually check the published weights against the number you were given.

## Every detection report gets a permanent public link, and the tests never touch a real relay

Two things about reports and tests, one risky and one reassuring.

The hosted output includes a permanent share link at `/r/{id}` and a JPG card at `/r/{id}.jpg`, described as directly shareable to chat groups and forums. Permanent means it does not expire, and the report is derived from the base URL and model you supplied, so a shared report identifies which relay you use and what it scored. The key is not in the report, but the endpoint is.

If that matters to you, read one report before sharing one.

The tests are the other side. The manifest configures pytest with `asyncio_mode` set to auto and the test paths pointed at the tests directory, and the dev extra brings in `respx`, which is a mocking library for the HTTP client the detectors use. So every probe in the suite runs against mocked responses.

That is the correct way to write a test suite, and it has one honest limit: no automated test can confirm that any of these fingerprint detectors actually fires against a real relay in the wild. The detectors are validated against fixtures, including, in the Claude protocol package, bundled JSON and PDF data files shipped inside the wheel for the document capability probe.

One more detail. `reportlab` sits in the dev extra rather than the runtime, with a comment saying it is only for a single test PDF build script. The image path is separate again, with Pillow in the web extra for the share cards. Three optional dependency groups for one small tool is tidy.

## Conclusion

Use Veridrop before you commit spend to a relay you did not build, and treat a clean report as evidence about capability rather than proof of identity, since only the Claude protocol has a cryptographic check and the other two rest on fingerprints that a determined relay can rewrite. Read DESIGN.md rather than the score, because the weights are published and a score is only meaningful next to them. Two things to settle first: the key is passed to whatever you point the tool at, so only test relays you trust with it, and a self-hosted run writes that key to a .env file in the working tree even though the hosted service claims keys are never persisted. Finally, if you run a relay, expect to be scored publicly on a permanent share link by a project that also sells certification listings.

## FAQ

### What does Veridrop check?

Given a base URL, API key and model, it runs probe requests against an AI API relay and diffs the responses against official baselines at field, protocol and cryptographic level. It answers whether the relay is genuine, whether capabilities such as PDF, tool use, thinking and long context are silently stripped, and whether response fields, streaming events and usage accounting match the official specs.

### Which API protocols does Veridrop support?

Three: the Anthropic Messages API, OpenAI Chat Completions, and the Gemini OpenAI-compatible API. The detector counts differ by protocol, with twelve for Claude, eight for OpenAI and seven for Gemini, and the long-context tier is not yet supported for Gemini.

### Why is the Claude thinking signature treated as stronger evidence?

Because when thinking is enabled, the Anthropic Messages API returns a signature field signed server-side, which the project describes as something a relay cannot forge. It carries twenty-five percent of the Claude protocol's weight, while OpenAI and Gemini have no equivalent signing mechanism and are verified only at protocol and behavioural level.

### How much does running a Veridrop detection cost?

A connectivity test is described as taking seconds at near-zero cost, a full detection about a minute at roughly twelve cents, the standard long-context tier about five to fifty cents, and the extreme tier from about five cents up to eight dollars depending on the model's advertised limit. A run stops at the first failed layer to avoid spending more.

### How do I self-host Veridrop?

Clone the repository over SSH, create a virtual environment, and install with the dev and web extras. Then copy the example environment file to .env and fill in the relay base URL, API key and model, using the ping and detect commands before comparing the result against baselines.

### Is Veridrop free and open source?

The code is published under the AGPL-3.0-or-later licence with the detection logic, scoring weights and field evidence public in the repository. The hosted service is free and needs no signup, and the project also offers business cooperation and certification listing while stating that such arrangements do not change scores or verdicts.

## Sources

- [canarybyte/veridrop on GitHub](https://github.com/canarybyte/veridrop)
- [License: AGPL-3.0](https://github.com/canarybyte/veridrop/blob/main/LICENSE)
- [Project website](https://veridrop.org)
- [README](https://github.com/canarybyte/veridrop/blob/main/README.md)
- [Releases](https://github.com/canarybyte/veridrop/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/canarybyte-veridrop
