ai-text-detector: the 84 in the sample output and the 18.4 in the field data
A cautious, explainable AI-like text risk analyzer for local workflows and coding agents.
At a glance
- What is it?
- ai-text-detector is a standard library Python project that scores text for AI-like writing risk with weighted signals, shipped as a CLI, a Python API, and a Codex skill. Its own evaluation puts the mean AI score at 18.4 and covered accuracy at 0.317, so the 84 in the sample output is a property of a bundled example file rather than of the detector in the field.
- Who is it for?
- Use this project as a queue builder and a source of explanations, not as a verdict. It ships without runtime dependencies, its strongest output is the name and note on each signal, and its schema cannot express high confidence at all.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 57 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The 84 in the sample output is a property of the bundled example file
The most quoted number in this repository is the one in the install section, and it comes from a file that ships with it. The install step is two lines:
pip install -e .
ai-detect examples/sample_ai_like.txt --jsonPointed at examples/sample_ai_like.txt, the detector answers:
Conclusion: AI-like signals are present, but this medium-confidence score is a risk estimate rather than proof.
Score: 84/100
Confidence: medium
Verdict: high_ai_likelihood
Words analyzed: 256The field data points the other way. On the public slice the evaluation describes, the mean score for AI answers is 18.4 and the mean for human answers is 5.4, a separation of 13.0 points on a 0 to 100 scale, and 45 is the only score in the repository attached to a number. Both means sit far below it. So 84 belongs to a file named for what it is, a sample written to look AI-like, while the averages come from real answers. The mapping from score to verdict is not written down anywhere: score-bands.svg exists in the assets folder, but the bands it draws are never stated as text.
Coverage runs 0.427 on human answers and 0.920 on AI answers
Coverage, not accuracy, is the figure that says where this detector declines to answer. On the same slice, human coverage is 0.427 and AI coverage is 0.920. Read literally, a little under 43 percent of human answers receive a score at all, against 92 percent of AI answers, so the guardrail withholds a judgement on more than half the human side while speaking on nearly all of the AI side. The evaluation credits the short-text guardrail for that, especially on the shorter human answers. The accuracy figure behind it is thinner: covered accuracy at score >= 45 is 0.317, meaning fewer than a third of covered cases land on the correct side of the only threshold that carries a number. The project says plainly that the detector separates human and AI answers on average but only weakly on HC3, that the thresholds are conservative, that conservatism keeps false confidence down while also lowering recall, and that the result is better read as triage plus explanation than as a stand-alone classifier.
Confidence has no high setting and the tool refuses text under about 80 words
The result schema carries a deliberate hole. Confidence is currently low or medium, so there is no high value the tool can ever emit, which makes the medium in the example output a ceiling rather than a description. Four verdicts exist, and one of them, insufficient_text, is not a judgement about authorship at all but a refusal to make one. The floor for that refusal is given as about 80 words in the exclusion list, while the sample that does produce a score analyses 256 words. The intent is stated in both directions: the tool is built so an agent can say that AI-like signals are present, that a sample is too short for a meaningful estimate, or that a passage should be reviewed against known writing samples, and it is built not to say that text was definitely written by AI or that the detector proves misconduct. Evidence arrives through strongest_signals() as name and note pairs, and caveats stay attached to the result instead of being dropped.
requirements.txt points at a dependency list pyproject.toml never declares
Two dependency files disagree about where dependencies live. requirements.txt states that runtime dependencies are defined in pyproject.toml and that the project currently uses only the Python standard library at runtime. The pyproject.toml in this repository carries no dependencies list at all, only a build requirement of setuptools 68 or newer, so the file requirements.txt points at is not holding the pointer. Nothing breaks, because a standard library only runtime needs nothing installed, but the second sentence carries the meaning and the first one is now a dangling reference. The packaging around it is compact. The distribution is named ai-detector-skill at version 0.1.0, it requires Python 3.9 or newer, packages are found under src, tests run with src on the python path, and a single console script is registered: ai-detect, bound to aidetect.cli:main. The project classifier calls it alpha.
The skill install copies ai-detector-skill out of a checkout named ai-text-detector
The local skill install is one line, and it depends on a directory name the repository does not use. The documented command copies ai-detector-skill into a skills folder under CODEX_HOME, while the repository is named ai-text-detector and the printed project structure labels its own root ai-detector-skill, so the copy only finds its source when the checkout directory carries the name the command expects. The surrounding text stays quiet about CODEX_HOME as well: no location is given for it and no instruction covers the case where it is unset, in which case the destination path collapses to a root level skills folder. The portability story is otherwise explicit. The root SKILL.md is the portable skill definition, and AGENTS.md stays at the repository root for repo-aware agents. A .claude directory also exists in the tree, and nothing in the README explains what it holds or how it differs from the SKILL.md install.
The printed project structure shows eight of the seventeen root entries
The structure section prints eight entries and the repository root holds seventeen. In the printed tree there are SKILL.md, a scripts folder, a references folder holding api-reference.md, an assets folder with hero.svg, score-bands.svg, workflow.svg and templates/report.md, src/aidetect, tests, AGENTS.md, and README.md. Absent from it are the files that actually run the project: the Makefile, pyproject.toml, and requirements.txt. Also absent are SECURITY.md, CONTRIBUTING.md, the Chinese README, the docs folder that holds both evaluation reports, the examples folder with sample_ai_like.txt and sample_human_like.txt, and the .github and .claude directories. The examples folder is what the demo target feeds to the module, and the header of the README links two sibling projects, lynote-ai/humanize-text and lynote-ai/ai-image-detector, alongside this one.
CI covers Python 3.9, 3.11, and 3.13 and never runs the demo target
The automation runs three commands on three interpreters, and the shape of that matrix is unusual. make test executes on Python 3.9, 3.11, and 3.13, so the even-numbered releases in between are never exercised. make benchmark regenerates a synthetic benchmark report and make eval-hc3 regenerates the HC3 evaluation report, and both files are then uploaded as workflow artifacts. The demo target is the one entry the Makefile offers that CI does not run, even though it is the command that produces the sample output shown earlier. Both report targets end in a markdown flag pointing at a file under docs, which means those reports are written by the runs themselves rather than committed by hand. Every figure in the README evaluation list is therefore a snapshot of whatever the last run produced. The repository's only release, v0.1.0, carries the same 2026-08-05 timestamp as the last push.
Four use cases, four exclusions, and one tool that says it is not a classifier
Every documented use case ends with a person reading the output. A teacher receives a polished 400 word reflection and runs ai-detect submission.txt --json before manual review. An editor wants to catch formulaic product reviews and guest posts, where the notes claim medium-length prose behaves better than short snippets. A moderation team sorts suspicious long-form posts into a manual queue instead of removing them. A content QA team compares human and AI-assisted drafts across versions and treats the score as a relative signal rather than an absolute one. The exclusion list is just as specific: no disciplinary decisions about a named student or employee, no single score treated as proof of cheating or fraud, nothing under about 80 words, and no high-stakes authorship dispute without a known-sample comparison. That is a narrow remit, and the repository states it rather than hiding it, which is the main reason to read the evaluation numbers before trusting the score line.
Editorial conclusion
Use this project as a queue builder and a source of explanations, not as a verdict. It ships without runtime dependencies, its strongest output is the name and note on each signal, and its schema cannot express high confidence at all. Before trusting a number, check which threshold produced it, because the one cutoff in the repository is written as 45 while the example output scores 84 on a file written to be AI-like, and the public slice averages 18.4. The 43 percent human coverage figure is the one to read first, because it decides how often the tool simply declines to speak.
Frequently asked questions
Where does the 84/100 score in the ai-text-detector example output come from?
It comes from examples/sample_ai_like.txt, a bundled sample analysed at 256 words, and it returns medium confidence with the verdict high_ai_likelihood. On the public evaluation slice the mean score for AI answers is 18.4.
Can ai-text-detector report high confidence?
No. Confidence is currently low or medium, so there is no high value in the schema, and the medium shown in the sample output is the ceiling the tool can reach.
What are the runtime dependencies of ai-text-detector?
There are none at runtime, according to requirements.txt, which says the project uses only the Python standard library and points dependency definitions at pyproject.toml. That pyproject.toml file declares no dependencies and requires setuptools 68 or newer only for the build.
How does ai-text-detector score on the HC3 dataset?
On the English finance, medicine, and open_qa subsets at the first 100 rows each, the human mean is 5.4 and the AI mean is 18.4, a separation of 13.0 points. Human coverage is 0.427 against AI coverage of 0.920, and covered accuracy at score >= 45 is 0.317.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lynote-ai-ai-text-detector)