Model or dataset
tianchong-zerotemp/dianxing avatar
tianchong-zerotemp/dianxing

DianXing publishes 8,451 findings, a hash table, and no code

DianXing - AI-Driven End-to-End Code Security Auditing

871 stars67 forksUnknownLicense varies

At a glance

What is it?
An AI code auditing system whose implementation is deliberately sealed, offering bounty payments for whatever it missed and a SHA-256 table committed before each challenge window. The headline count is exploitation chains rather than distinct defects, the severity grades are assigned by the model itself, and 704 of the 8,451 findings appear in no severity column at all.
Who is it for?
This is a claim with a verification wrapper, not a tool, and the right response is to treat it that way. Nothing here can be run, installed or audited, because the repository holds no code and no licence, so every number is a self-report and the hash table only fixes the order in which the reports were written.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 108 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

No code, no licence, and two QR codes at the root

The repository contains no source code. Its top level is a readme, a gitignore, a file naming a public challenge hash table for one specific project version, a markdown document describing the severity rating standard, and two images whose filenames translate as a group QR code. The repository's language and its licence both come back unclassified. So there is nothing to install, nothing to run and nothing to fork, and the page is candid about the first of those: it explains at length why the implementation is withheld and says the alternative is to publish verifiable audit data and a bounty challenge instead. It also anchors that decision to two named third-party programmes, one model and one collaborative effort, invoked as precedent for a company choosing not to publish capability. Neither is linked anywhere in the visible page, so the precedent is asserted rather than offered for checking.

Eight thousand findings counts attack chains, not defects

The most useful paragraph in the repository is the callout that tells you not to trust the headline. It states the counting rule: the same vulnerability site, when several independent exploitation chains exist because of different trigger entry points or different exploitation methods, is counted as separate findings. So the cumulative figure reflects the total independently exploitable attack surface rather than a deduplicated count of code defects, and the page says that choice was made for transparency rather than to make the number look good. It is an unusual and creditable disclosure. It also means the number cannot be compared to a scanner's finding count, a bug bounty total or another tool's report without first converting both to the same unit, and it means a single missing chain would have inflated the figure rather than deflated it.

Seven hundred and four of the findings have no severity grade

The per-project table has a total column and three severity columns, and in twelve of fourteen rows the three severities do not add up to the total. LiteLLM reports 2,074 total against 1,999 graded, a shortfall of 75. Jenkins reports 1,469 against 1,170, a shortfall of 299, the largest single gap. Grav CMS is short by 35, AcePanel by 132, phpMyAdmin by 26. Only the Aave V4 row and the firmware row reconcile. Adding every row, the totals column sums to exactly 8,451, matching the headline and the language table below it, while high, medium and low sum to 7,747. So 704 findings, just over eight percent, are counted in the total and absent from every severity bucket. One number does reconcile: the later claim of 1,796 high-or-above findings is exactly 1,785 high plus 11 critical.

Every severity grade was assigned by the model

A note under the project table states that all severity gradings were produced by AI rather than by a human. The grading scheme itself is a separate markdown file in the repository, so the standard is inspectable even though the grading is not. Critical is defined narrowly: unauthenticated remote code execution, reproduced on a target machine with retained evidence of full command execution. Remote code execution found by static reading alone goes into the high bucket instead, and firmware findings with preconditions also go to high. That is a defensible definition and it is stricter than most vendor severity schemes. What the page does not say is how many of the 8,451 were reviewed by anyone, or whether the target machine reproductions were witnessed. The eleven critical findings are the only ones with a verification step described at all.

The zero percent miss rate is measured against its own output

The head-to-head table compares the system with three commercial auditing products on one project version, at one time, under what the page calls the same evaluation criteria. Its row reads 1,215 findings, a 98.6% true positive rate, a 0% miss rate, and 14 remote code executions reproduced on a target machine. The rivals report 7, 2 and 3 findings with true positive rates of 100%, 50% and 33%, and every miss-rate cell for them is empty. A zero percent miss rate needs a known denominator, and the only denominator on the page is the named system's own finding count. So the one column that makes the comparison decisive is measured against a ground truth produced by the system being measured, and the columns that would let a reader check it are blank for everyone else. The remote execution column is the one that does compare cleanly: 14 against a rival high of 3.

Nine benchmark languages, seven audited ones

The recall benchmark sets out its constraints carefully: nine languages covering C, Java, JavaScript, Python, Go, PHP, .NET, Rust and Ruby, at least five distinct vulnerability types per language, every vulnerability disclosed after the model's training cutoff, no network access during the run, no human intervention, and a single run with no parameter retries. The result line says all nine languages, zero misses. Two things are missing. No counts appear anywhere, so the size of the test set per language is unstated and a zero-miss result on five findings and a zero-miss result on five hundred read identically. And the language sets do not line up with the audited projects, which cover Java, Go, JavaScript, PHP, Python, Solidity and binary firmware. Solidity and firmware appear in the audit list and not in the benchmark, and C, .NET, Rust and Ruby appear in the benchmark and not in the audit list.

The hash table fixes the order of the reports, not their truth

The verification design is more thought through than most. For each audited project the team publishes a CSV of SHA-256 hashes ahead of opening the challenge window, using a git commit timestamp as the tamper-evident record. Each finding gets two hashes, one over the canonical record and one over the description text, and either matching counts as a hit. A submitter who finds something the system missed receives the record text by email, hashes it locally and compares. What this proves is narrow: that a particular piece of text existed in the vendor's inventory before the window opened. It does not prove the vulnerability is real, it does not prove the record describes the right project or version, and since the text is the vendor's own prose, the commitment is over self-authored description. Its actual job is de-duplication, deciding who was first. The flow is one line:

code
漏洞描述原文 → SHA-256 → 0x7a3f...b2c1 → Git commit 存证(含时间戳)

Editorial conclusion

This is a claim with a verification wrapper, not a tool, and the right response is to treat it that way. Nothing here can be run, installed or audited, because the repository holds no code and no licence, so every number is a self-report and the hash table only fixes the order in which the reports were written. What the page does unusually well is disclose its counting method, which is the one thing that makes the headline number readable at all. Read the two count columns as different quantities, notice that the severity breakdown accounts for 7,747 of 8,451 findings, and treat the zero percent miss rate as a claim about the named system's own output rather than about the others. If you want a code auditor you can inspect, this repository is not one.

Frequently asked questions

What does DianXing do and can I run it?

It is described as an end-to-end AI code auditing system that takes uploaded source code and returns a structured vulnerability list with no human intervention. You cannot run it. The repository holds no source code, no package manifest and no licence, only a readme, a severity standard document, a public hash table for one project version and two group QR code images.

How does DianXing count its vulnerabilities?

In exploitation chains rather than deduplicated defects. The page states that the same vulnerability site counts as several findings when independent exploitation chains exist, with different trigger entry points or different exploitation methods, and that the figure is meant to reflect independently exploitable attack surface.

How many vulnerabilities did DianXing claim to find, and how severe were they?

8,451 across 14 projects, of which 11 are described as unauthenticated remote code execution reproduced on a target machine. The severity columns report 1,785 high, 5,310 medium and 652 low, which sums to 7,747, leaving 704 findings counted in the totals but absent from every severity bucket. All gradings are stated to be produced by AI.

How does the DianXing vulnerability hunter challenge work?

Each round the team publishes one already audited project, then commits a CSV of SHA-256 hashes for every finding before the challenge window opens, using the commit timestamp as proof of order. Submitters post publicly or email privately, receive the matching record text, hash it locally and compare. A match proves the text existed in the inventory first, which settles precedence rather than correctness.

Which projects did DianXing audit, and which are hidden?

Named ones include LiteLLM, Jenkins, Grav CMS, New API, AcePanel, phpMyAdmin, Redash, Damn Vulnerable DeFi, Aave V4 and a Tenda router firmware audit. Four rows have their names and versions blacked out because the system found unauthenticated remote code execution in them; the redaction keeps the language, the totals and the severity counts of those rows.

Official sources

  1. Issues
  2. README
  3. tianchong-zerotemp/dianxing on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tianchong-zerotemp-dianxing.svg)](https://hysenlabs.com/projects/tianchong-zerotemp-dianxing)