hlwy-ai-checker: fingerprinting third-party AI APIs to see if the model you paid for is the model you got
检查第三方AI API是否掺假以及渠道一致|基于llm指纹的AI模型识别
At a glance
- What is it?
- hlwy-ai-checker compares the statistical behaviour of an official API endpoint against a reseller's endpoint, using repeated sampling rather than prompt tricks. It is a local, Python-launched tool for people who suspect a cheap relay channel is substituting a smaller model.
- Who is it for?
- Adopt hlwy-ai-checker if you resell, resubscribe through, or buy from relay channels and want a local, LGPL-2.1 statistics-based signal before you renew a contract. Do not adopt it as evidence in a refund dispute: the README states plainly that results cannot serve as an absolute legal or factual basis, and it is a statistical comparison, not proof of which model answered.
- Can I use it commercially?
- Yes, with conditions. LGPL-2.1 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: relay channels that quietly swap the model behind an API key
A large share of cheap AI API access does not come from the model vendor. It comes from a relay or reseller that holds an official key and re-exposes it under its own Base URL, often at a discount. The buyer sends an OpenAI-compatible request and gets a plausible answer back. Nothing in the response tells the buyer which model produced it. The README frames the project around exactly this gap: checking whether a third-party AI API is adulterated, and whether a channel behaves consistently with the official one. The audience is narrow and practical. It is people who already pay for API access through an intermediary, or who are deciding whether to, and who want a local check before committing money. The README also states that the tool can be deployed entirely locally, which keeps the API key and the test data off anyone else's server. That matters here, because the test itself requires sending your official key to your own machine rather than to a hosted checker. A hosted verification service would need both keys, which is precisely the thing a suspicious buyer is unwilling to hand over. The project's answer is to ship the whole thing as a local front end and launcher instead of an API.
Why a language model can be fingerprinted at all
The method rests on a claim the README makes directly: a large language model is not a true random number generator. When you ask a model to pick a number at random, the output is shaped by training data, architecture, RLHF alignment, tokenization strategy and sampling parameters. Those influences leave statistical bias, and repeated sampling turns that bias into a distribution. Compare the distribution from an official endpoint with the distribution from a third-party endpoint on the same prompt set, and you get a similarity measure. The README is explicit that this is statistical detection. It lists the things that move the result: model version, sampling parameters, server-side configuration, sample size and request success rate. It also states that the method cannot on its own prove that a third-party channel did or did not use a particular model. That sentence is the most important one in the repository, and it is the reason the tool should be read as a screening instrument rather than a verdict machine. The README's own feature section claims high discrimination, low random variation and low token consumption, and those claims sit inside the repository without published methodology, so they are the maintainer's assertions rather than independently reproduced measurements.
Calibrate first, then compare: the two-stage workflow
The design is two-stage, and the order is not optional. Stage one is calibration against the official endpoint. You supply the official API key and a Base URL that must include /v1, and the tool records the fingerprint for a chosen model. Stage two is verification. You supply the third-party channel's API key and Base URL, select the same model you calibrated, and run the test. The result screen then puts the official fingerprint and the third-party fingerprint side by side with similarity and related statistics. The README calls this adaptive: because calibration happens before testing, the reference is built from the live behaviour of the official endpoint at that moment rather than from a fixed table. That choice has a cost. Every comparison requires a fresh calibration run against the official endpoint, which doubles the requests and the wall-clock time compared with a tool that ships fixed reference tables. The repository keeps a baselines/ directory at the top level, which suggests stored reference fingerprints ship with the project, but the README does not document what is inside it or how a stored baseline is selected. Treat that directory as something to inspect before you trust a comparison you did not calibrate yourself.
Installing it and running a first calibration
There is no package on a registry. The README says to download the latest ZIP from the Releases page and extract it completely, then run the launcher. The repository also ships a text file named 使用方法_打开start.py(用python).txt, which restates the same instruction in the repository's own language. Extraction matters: the tool is launched from the extracted directory, and the HTML front end sits next to the Python entry point.
python start.pyAfter that command the README shows an interface with an automatic mode and a manual mode. Automatic mode is the one-click test. Manual mode is the three-step path: calibrate with the official key, verify with the third-party key, then read the comparison. For the manual path, the README's only hard constraint on input is that the Base URL must contain /v1, as shown in the calibration step. The third-party step takes the same shape, with the channel's own key and Base URL and the same model selected. What you should see at the end is a similarity score and accompanying statistics for the two fingerprints. The README does not document a pass/fail threshold, so the number is something you interpret, not something the tool adjudicates. It also does not document a timeout, a retry policy or what happens when a request fails mid-run, which means a flaky channel can quietly shrink your sample.
Where the method breaks down
The failure modes are stated in the README and they are not minor. Sampling parameters on the two sides must match, because temperature and related settings change the output distribution independently of the model. Server-side configuration at the relay can alter behaviour without any model substitution taking place. A low request success rate shrinks the effective sample and widens the noise. A model version bump on the official side invalidates a calibration you ran last week. The README's own conclusion is that the test cannot prove use or non-use of a model on its own. There is a second limitation that follows from the architecture: the tool talks to whatever Base URL you give it, so a channel that detects probing traffic and returns canned or proxied responses can distort the sample. And the licence is LGPL-2.1, which is a copyleft licence. If you plan to embed this in a product rather than run it as a tool, the terms of that licence apply to the combined work, and that is a question for your own counsel rather than something the README answers. One more gap worth naming: the README does not document rollback, an upgrade procedure, or how a stored baseline is invalidated when a vendor ships a new model version.
How it differs from prompt-based model-identification tools
The obvious alternative is asking the model who it is. Plenty of people test a suspicious endpoint by sending a prompt like "which model are you" and reading the answer. That approach fails immediately against any relay that wants to hide the substitution, because the system prompt or a wrapper can simply assert a different identity, and the answer costs one request to fake. hlwy-ai-checker takes the opposite route: it never asks the model to describe itself, it measures how the model behaves across repeated samples and compares that distribution against a calibrated reference. The trade-off is real. The prompt approach is instant and readable but trivially spoofable. The fingerprint approach is harder to fake but slower, consumes tokens across many requests, and returns a similarity number that requires interpretation instead of a yes or no. The README's claim of low token consumption is relative to that sampling requirement, not a claim that the test is free. A third option, reading the vendor's own usage dashboard, only works when the vendor exposes per-request model attribution to the reseller's account, which is exactly what a reseller arrangement tends to obscure.
Maintenance, releases and what the repository tells you
The repository is not archived, and the last push was on 2026-09-06. Three releases are listed: 2.4.0 on 2026-08-28, 2.3.0 on 2026-08-10, and 2.5-pre1 on 2026-09-01, which is marked as a pre-release by its version string. That cadence suggests the project is still being worked on, but a pre-release sitting alongside a stable line means you should pin to 2.4.0 unless you specifically want to test the newer build. The primary language is HTML, which matches the repository layout: a single hlwy-ai-checker.html front end, a start.py launcher, and a baselines/ directory. There is no documented upgrade mechanism, no migration notes, and the README does not describe rollback. Upgrading means replacing the extracted directory, which also means re-running calibration, because a new build may change how fingerprints are computed. The licence is LGPL-2.1, and the README carries a disclaimer that the maintainer does not participate in commercial disputes between users and API providers. The project states it was published on Linux Do, and the README links that community as a friend link, which is where release discussion is most likely to appear.
Who this is for, and what to check before you trust a result
This is a tool for a specific decision: whether to keep buying API access from a particular channel. It is not a general AI detector, and the search traffic around AI content checkers has nothing to do with what it does. The people who get value from it are teams routing production traffic through a relay, and individuals who noticed that a cheap endpoint answers suspiciously fast or suspiciously badly. The people who should not use it are anyone hoping for a certificate. The README says the results are for reference only and cannot serve as an absolute legal or factual basis in a commercial dispute. A sensible first run is a control: calibrate against the official endpoint and then verify against that same official endpoint. If the tool does not score a channel against itself as highly similar, your sampling parameters or sample size are wrong, and any third-party number you produce afterwards is noise. Run that control before you accuse anyone of anything.
Editorial conclusion
Adopt hlwy-ai-checker if you resell, resubscribe through, or buy from relay channels and want a local, LGPL-2.1 statistics-based signal before you renew a contract. Do not adopt it as evidence in a refund dispute: the README states plainly that results cannot serve as an absolute legal or factual basis, and it is a statistical comparison, not proof of which model answered. Before relying on it, run the calibration step against the official endpoint yourself, keep the same model and sampling parameters on both sides, and check the baselines/ directory to see whether a reference fingerprint for your model already exists.
Frequently asked questions
Does hlwy-ai-checker prove that a third-party API substituted a different model?
No. The README states that the method is statistical detection and cannot on its own prove that a third-party channel did or did not use a particular model.
How do I install and start hlwy-ai-checker?
Download the latest ZIP from the Releases page, extract it completely, and run python start.py from the extracted directory.
What has to match between the calibration run and the third-party test?
The model must be the same in both stages, and the Base URL must include /v1. The README also lists sampling parameters, model version, server-side configuration and request success rate as factors that affect the result.
What licence does hlwy-ai-checker use?
The repository lists LGPL-2.1. The README does not discuss the implications of that licence for products that embed the tool.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hanlinwenyuan-hlwy-ai-checker)