The demo report gives Performance 10.0 for finding nothing, which the skill's own rules forbid
An evidence-based AI code audit skill. Professional output. Zero emotional bullshit.
At a glance
- What is it?
- A prompt-only audit skill for Codex, Claude Code, Copilot and Gemini that asks a coding agent to review a codebase along 26 named dimensions and return a scored HTML or JSON report with file-and-line evidence. The evidence discipline is the interesting part. The sample output it ships to demonstrate the skill breaks the two rules the feature list states most firmly.
- Who is it for?
- Use it as a prompt scaffold if you want a structured second opinion that refuses to return vague praise and insists on file-and-line evidence, since that is what the skill genuinely contributes. Do not treat its score as a measurement, because the overall figure is an unweighted mean of seven sub-scores and the shipped example demonstrates the not-found-becomes-a-perfect-score behaviour the documentation disclaims.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The sample output awards a perfect score for finding nothing
The feature list states two rules about missing findings. Areas not deeply covered are explicitly marked `Not assessed`. And it says the skill will never dress up not having found anything as a full mark.
Then there is the sample report, which the project ships as a demo. The Performance row reads 10.0, grade S, with the reason given as no obvious I/O blocking or hot-path overhead found.
That is the excluded behaviour, executed in the project's own example. A dimension where the agent searched and found nothing receives the maximum score and the top grade, rather than an Not assessed marker or a partial. The letter grades in the same table run S at 10, A at 8 and 7, B at 6 and 5, and C at 4, so the perfect score is not an artefact of rounding.
The scale itself is only ever shown by example. Seven rows give you seven data points, and there is no specification of where A ends and S begins, or what 9.0 would score. So a reader cannot tell whether S means ten out of ten or eight and a half upward, which is a gap for a feature whose central claim is a quantitative score.
The rest of the sample is more careful in its wording. The 4.0 on Testing is attributed to mostly ineffective mocks and missing assertions on key branches rather than to an absence of tests, and the 8.0 on Security names a specific missing guard on a key route plus a sensitive default in a config value. That is the behaviour the evidence rule is meant to produce, and it shows up in the reasons even where it does not show up in the grading.
Overall is the plain average of seven scores, and nothing says how 26 dimensions become 7
The sample panel has seven scored dimensions: Security, Stability, Performance, Testing, Maintainability, Design and Release. The dimension catalogue the project publishes has 26 entries. So 26 things get examined and 7 numbers come out.
The Overall figure is 6.6. Add the seven scores and you get 46.0. Divide by seven and you get 6.57, which rounds to 6.6. So the overall is an unweighted arithmetic mean.
That arithmetic has a consequence worth stating. Performance at 10.0 and Testing at 4.0 pull the average by exactly the same amount in opposite directions. A codebase with a real concurrency defect and a fast hot path scores the same overall as a codebase with excellent concurrency and no tests, provided the other five dimensions match.
Unweighted means also let a dimension with nothing to measure inflate the result. Performance is the clearest case, since a codebase with no meaningful I/O has almost nothing to find, so the dimension that is easiest to score highly is also the one where a high score carries the least information.
The 26-to-7 reduction is the other undocumented step. The catalogue names data integrity, type safety, fallback behaviour, accessibility, dependency weight, comment coverage, documentation, frontend state and backend API as dimensions in their own right, and the README never says which of the seven panels each of them feeds. Whether a data integrity finding moves Security or Stability, and whether a type assertion problem counts as Maintainability or Design, changes what a score means and is not recoverable from the documentation.
The headline says 26+ and the list contains 26
The claim is repeated in three places. The feature list refers to a matrix of more than 26 specialist dimensions covering seven core scored dimensions. The full mode is described as covering all 26+ dimensions automatically. The dimension list is behind a disclosure element labelled click to expand the full list of 26+ dimensions.
The list itself contains exactly 26 entries: architecture, security, stability, concurrency, performance, testing, testing-authenticity, maintainability, design, release, configuration, observability, data-integrity, privacy, accessibility, supply-chain, cost, ai-safety, fallback, type-safety, frontend-state, backend-api, dependency-weight, code-consistency, comment-coverage and documentation.
So the plus sign in 26+ is doing work the enumeration does not support. Either there is a twenty-seventh dimension in the skill that is not published, or the number was written as a bound and the list caught up exactly to it.
That distinction is not pedantic for anyone evaluating coverage. Seven of the 26 produce a scored panel. The other nineteen produce findings with evidence but no number. So a report saying Security 8.0 and Stability 6.0 is reporting on two dimensions out of twenty-six, and the score panel gives no hint that nineteen others were involved.
The most interesting entry is testing-authenticity, described as over-mocking, testing implementation details rather than behaviour, false-green tests, and missing real end-to-end paths. Having a whole dimension for whether the tests lie is unusual and is arguably the most valuable idea in the catalogue, since a passing suite is the most common false signal in a codebase. It also has no score panel of its own.
Ten of the 26 dimensions are reachable only through full mode or by naming the ID
There is a mapping table from human intent to internal mode identifiers, and it is the practical entry point, because the project asks you not to memorise the parameters and instead to describe what you want in ordinary engineering language.
Eight rows are given. A full sweep maps to `full`. Pre-release compliance maps to release, stability, observability and configuration. Permissions and supply chain maps to security, privacy and supply-chain. Concurrency and deadlock maps to concurrency and stability. A pull request review maps to `incremental`. AI and LLM work maps to ai-safety, privacy, cost and observability. Suspected false-passing tests maps to testing and testing-authenticity. Preparing for a refactor maps to maintainability, architecture, design and code-consistency.
Cross-referencing that against the catalogue, ten dimensions never appear in any row: performance, data-integrity, accessibility, fallback, type-safety, frontend-state, backend-api, dependency-weight, comment-coverage and documentation.
So performance is only reachable through a full audit or by asking for it by name. Accessibility likewise. So is anything about fallback behaviour, which the catalogue describes as dangerous silent degradation, empty catch blocks swallowing errors and type guessing. Those are three of the failure modes that most damage a codebase quietly, and they sit outside every suggested phrasing.
This is not a criticism of the mapping, which is presented as examples rather than as an exhaustive table. It is a limitation for anyone who asks in natural language and expects a specific dimension to be picked up: the model is choosing the scope, and it will choose from what the prompt implies, so a dimension nobody thinks to mention is simply not examined.
Which connects to the one Python file the project names. The context FAQ says the system builds a project map first using project_inventory.py, then prioritises high-risk surfaces such as authentication, gateways, payment and core data flows.
Installation is a URL pasted into a chat, and all four hosts activate differently
There is no installer. The documented first step is to send the repository link to an agentic IDE and ask it to install the skill directory from the repository, which in practice means pasting a URL and a sentence into a chat window.
The manual route is a directory copy into the host's skills folder, and the four hosts differ in how the skill becomes active:
| Platform | Path | Activation | | :--- | :--- | :--- | | Codex | `~/.codex/skills/fuck-my-shit-mountain/` | A new conversation after copying | | Claude Code | `~/.claude/skills/fuck-my-shit-mountain/` or `.claude/skills/...` | Type the slash command, or invoke in natural language | | GitHub Copilot | `~/.copilot/skills/fuck-my-shit-mountain/` or `.github/skills/...` | Run `/skills reload` in the chat window | | Gemini CLI | `~/.gemini/skills/fuck-my-shit-mountain/` or `.gemini/skills/...` | Run `/skills reload`, and trust the workspace first |
Three different activation mechanisms for the same copied directory: restart the conversation, type a command, or reload the skill list. And Gemini adds a workspace-trust step that the others do not have, so an install that works everywhere else can silently do nothing there.
The terminal shortcut covers only one of the four:
git clone https://github.com/XiNian-dada/Fuck_My_Shit_Mountain.git /tmp/shit-mountain
mkdir -p ~/.codex/skills
cp -R /tmp/shit-mountain/fuck-my-shit-mountain ~/.codex/skills/
rm -rf /tmp/shit-mountainIt clones the whole repository to a temporary directory, copies one subdirectory into the Codex path, and deletes the clone. For Claude Code, Copilot or Gemini you have to rewrite the destination yourself. The temporary path also means the project name ends up in your shell history, which is a small cost for a project whose name is the main thing it is known for.
The report language is a request, not a default
The example prompt sets three things explicitly, and each of them is a choice the reader has to know exists:
请使用 fuck-my-shit-mountain 对当前项目进行全量审计。
报告语言:中文
输出格式:htmlReport language: Chinese. Output format: html. The language line is not decoration. The documentation, the skill and the example are all Chinese, and nothing in the visible text indicates what language a report comes back in if you do not ask.
The three output formats are described as an interactive single-page HTML report, JSON for automation pipelines, and a minimal Markdown document. Since the project is aimed at pipeline use, with incremental mode described as embedding into daily review and continuous delivery, the JSON form is the one that matters for automation, and the HTML form is the one demonstrated.
The HTML report is also where the project publishes its demo, hosted as a GitHub Pages site. The demo is described as supporting light and dark adaptation, sidebar scroll tracking, risk-level filtering and detailed remediation cards. It is a static page, so what a reader sees is a rendering of a fixed report rather than a live run.
One more piece of framing worth knowing. The read-only guarantee in the FAQ says the audit will never modify business logic, configuration or lock files, and that it only writes audit reports and remediation suggestions into the project root or a specified path. The guarantee is about your code, not about the filesystem. Reports still land in the repository directory, so an audit on a git working tree produces untracked files unless they are ignored.
The AI picks the scope, and the FAQ admits coverage is prioritised rather than complete
Two design decisions combine into the project's main limitation, and both are stated rather than hidden.
The first is scope selection. The profiling feature is described as automatically identifying the tech stack, dependency manifest and key risk surfaces, then recommending audit dimensions in natural language. Combined with the instruction that you express your need in ordinary engineering language and the model maps it to mode identifiers, the selection of which of the 26 dimensions to apply is made by the model, not by a fixed rule. Two runs on the same repository with the same request can examine different dimensions.
The second is what happens on a large repository. The context FAQ asks what to do when a large codebase does not fit in context, and answers that there is a built-in progressive profiling mechanism: build a project map first, prioritise high-risk surfaces, and focus on incremental changes when appropriate.
Prioritise is the operative word. The answer to a question about not fitting is that less gets examined. So the honest description of the skill is that it reads a whole codebase shallowly and reads the high-risk parts deeply, and the Not assessed marker exists to distinguish the first from the second.
Given that, the reliability characteristics are the ones to judge it on, and the project is unusually explicit about them. It says an AI audit is not magic and cannot fully replace deep human review and production testing, and that its value is doing exhaustive risk sweeps at high speed so that architecture and security problems surface early. It also names the strictness explicitly: even a clean overall architecture will lose points on a dimension if a core call chain lacks timeouts and circuit breaking, or if there is a serious deadlock risk.
The reporting discipline it does commit to is worth keeping. Findings must carry line numbers, real trigger conditions, an impact derivation and an effort estimate. A finding without that evidence is supposed to be rejected rather than reported.
The repository homepage points at a forum, and there is no issue tracker or CI
The top-level listing is six entries: .gitignore, LICENSE, README.en.md, README.md, assets/, docs/, the skill directory itself, and a single file called html.html.
That html.html at the root is the demo report, and it carries a doubled extension. It is also the reason the repository's detected primary language is HTML rather than Markdown, since the skill is prompt and instruction text.
The README that most visitors land on is in Chinese. An English version, README.en.md, is present in the repository but was not the copy retrieved, so both exist and the English one is discoverable only by going looking for it.
The metadata homepage is a community forum, not the project. The project's own demo lives on GitHub Pages, and the community section points at two third-party directories that have listed the project alongside the forum. There is no issue tracker described, no contribution guide, and no continuous integration directory in the tree.
For a project whose premise is that an agent should find defects a human reviewer missed, the absence of a way to report a bad audit is a reasonable thing to notice. There is nowhere in the visible documentation to say that the skill produced a false finding or missed a real one, which is the feedback loop the rest of the design implies.
The licence is MIT and the last push was 2026-09-21, so the branch is current. There are no GitHub releases, so a copy of the skill directory is the only way to pin a version.
Editorial conclusion
Use it as a prompt scaffold if you want a structured second opinion that refuses to return vague praise and insists on file-and-line evidence, since that is what the skill genuinely contributes. Do not treat its score as a measurement, because the overall figure is an unweighted mean of seven sub-scores and the shipped example demonstrates the not-found-becomes-a-perfect-score behaviour the documentation disclaims. Verify first which dimensions a run actually applied, since the agent chooses the scope in natural language and unreferenced dimensions are marked Not assessed rather than scored zero.
Frequently asked questions
What does the Fuck My Shit Mountain skill actually do?
It is a prompt-only audit skill for coding agents that reviews a codebase along 26 named dimensions and produces a report with code evidence, file-and-line references, trigger conditions and prioritised fix suggestions. Outputs come as an interactive HTML report, JSON for automation, or Markdown. Incremental mode limits a run to files changed in a git diff.
How do I install the Fuck My Shit Mountain skill?
There is no installer. Either send the repository link to an agentic IDE and ask it to install the skill directory, or copy the fuck-my-shit-mountain/ directory into the host's skills folder. Activation then differs by host: a new conversation on Codex, a slash command on Claude Code, and /skills reload on Copilot and Gemini, with Gemini also requiring workspace trust first.
How is the overall audit score calculated?
The sample report shows seven scored dimensions, Security, Stability, Performance, Testing, Maintainability, Design and Release. The seven scores in that example sum to 46.0 and the displayed overall is 6.6, which is 46 divided by seven, so the overall is an unweighted mean. No mapping from the 26 analysed dimensions to those seven panels is published.
Will the audit change my code or write to my repository?
The documentation states an absolute no on modifying source or committing git records: the audit follows a read-only contract and will not change business logic, configuration or lock files. It does write audit reports and remediation suggestions into the project root or a specified path, so a run on a git working tree creates untracked files.
How does the skill handle a codebase too large to fit in context?
With a progressive profiling mechanism. The agent first builds a project map using project_inventory.py, then concentrates on high-risk surfaces such as authentication, gateways, payment and core data flows, and can narrow further with incremental mode focused on changed files. Areas that are not deeply covered are marked Not assessed rather than scored.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/xinian-dada-fuck-my-shit-mountain)