high-stakes-analytics-decision-lab ships a readiness gate with two stop states and seven operations that need written approval
A platform-neutral analytical Skill that profiles messy data, selects case-adaptive methods, and produces source-backed visual reports for high-stakes decisions.
At a glance
- What is it?
- An analytical agent skill whose stated enemy is a recommendation written because the template expects one. Row-level data never reach a model without passing a four-state gate, and only safe normalization runs unattended. The tree is unusually disciplined, and the README and the Makefile disagree on the name of the script that runs it.
- Who is it for?
- Use this if your problem is that you keep receiving confident reports built on ambiguous data contracts, because that is the failure it was built against and the gate is a real mechanism rather than a disclaimer. Three things to know first.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Row-level data do not reach a model before a gate, and the gate can stop the work
The mechanism is a four-state data gate, and two of its four states are stop states. Row-level uploads are not handed to a model. The system preserves the original, establishes a contract, checks grain and keys, profiles quality and privacy, and produces a dry-run remediation plan. What the gate then permits is explicit:
| `ready` | No material failure under the declared contract | Continue |
| `ready_with_documented_limitations` | Localized issues remain | Continue with visible limits |
| `needs_user_confirmation` | A substantive transformation, privacy, or intended-use choice remains | Pause for a named approval or clarification |
| `blocked` | Grain, key, schema, leakage, or another critical failure invalidates the route | Stop and request corrected evidence |The separation between the first two and the last two is the design. A project can proceed with visible limits, or it cannot proceed and must be given corrected evidence. There is no third option where the tool proceeds while quietly absorbing a problem.
The second half of the gate is a list of what may and may not happen without a person. Only safe normalization runs unattended. Seven operations require explicit action IDs: deletion, imputation, outlier treatment, category merging, unit conversion, target correction, and grain changes. That list is short enough to read before trusting the tool, and every item on it is a step that can change what the answer means. And the processed copy never overwrites the source, which is what makes the remediation plan reviewable at all.
There are six valid results, and the starter prompt pre-authorises stopping
The second thing that makes this different from a report generator is the outcome list. The page is explicit that the result is not a fixed report, and that a valid result may be a bounded action, a pilot requirement, targeted diligence, an evidence request, negative validation, or `do_not_deploy`. Six outcomes, of which four are ways of not producing a recommendation.
The stated reasoning is blunt about where this comes from. High-stakes analysis is said to fail before the model, for five named reasons: the question is underspecified, the data contract is implicit, cleaning choices are hidden, uncertainty is treated as independent, or a recommendation is written because the template expects one. The last of those is the one this whole design exists to prevent, and the `do_not_deploy` endpoint is the answer to it.
The starter prompt is written to get there before the user argues for it. The instruction shipped with the skill asks the agent to run the readiness gate, preserve the original, select only the routes the evidence supports, produce the evidence report, and add a decision brief only if the evidence and decision context justify one. The restraint is pre-authorised in the prompt rather than left to the model to invent, and the page tells you to start with the decision or evidence question rather than a preferred model, which is a small instruction with a large effect on what comes back.
The route vocabulary follows the four familiar analytics types, described as descriptive, diagnostic, predictive and prescriptive work added only when justified. Four routes, and no mandatory recommendation among them.
The installer fetches 39 files, and the research portfolio is not in them
Installation is one command through the Agent Skills installer:
npx skills add limingrui679-design/high-stakes-analytics-decision-lab -gWhat it discovers is a compact package under the skills directory, and the page gives its size precisely: 39 files and about 472 KiB. That is emphatically not the full research portfolio, which lives elsewhere in the same repository as fifteen complete evidence paths across fifteen real-data projects. So the artefact you install and the artefact you read about are deliberately different, and the page is upfront that the installer takes the small one.
The integrity of that small one is handled with a file rather than a promise. The bundle's machine-readable file and hash contract is `bundle-manifest.json`, so what the installer delivers can be enumerated and checked. Given that this project spends so much of its design on provenance for analytical claims, applying the same discipline to its own install payload is consistent rather than incidental.
There are more routes than the one command. A getting-started document covers Codex-specific, no-install, and direct repository options, and six direct entry points are tabulated in the page itself, including an environment audit, a sixty second synthetic walkthrough that produces no model call and no recommendation, a question-only route blueprint that invents no result, and a validate-then-run pair for an existing decision case.
That walkthrough is the sensible first thing to run. It exercises source preservation, the contract, the quality gate, route selection and accessible SVG output, on synthetic data, and it produces nothing you could mistake for an analysis of your own material.
Every figure has to name the question it answers and the source it came from
The output contract is four artefacts, and one of them is the unusual one:
report.md # primary Evidence Intelligence Report
results.json # machine-readable analytical result
chart-map.json # figure-to-question and source contract
figures/*.svg # accessible analytical visuals`chart-map.json` is described as a figure-to-question and source contract. So a chart in this system is not a picture, it is a node that has to declare which question it is answering and where the data came from. A figure that cannot be tied to both has nowhere to sit in the package. That is a much stronger requirement than the usual practice of alt text, because alt text describes the image and this requires the image to justify itself against the analysis.
The other choice worth noting is SVG rather than raster. That is the accessibility-friendly format for charts, and it pairs with the claim in the principles table that claims and accessible figures get linked to JSON, CSV, hashes and rerunnable code.
The decision layer is deliberately subordinate. A justified decision adds `decision-report.md`, `decision-results.json` and its own figure contract, and the page says in terms that it never replaces the primary evidence product. Given that the second half of the starter prompt asks for a decision brief only if justified, the file layout is enforcing the same rule the prompt states: the evidence product is the deliverable, and the decision is an optional layer on top of it rather than the other way round.
The fixed spine against the adaptive layer is the other structural idea. Question, population, unit, target quantity and horizon stay fixed; route, fields, methods and validation change. Source lineage, quality status and reproducibility stay fixed; figures, report sections and decision criteria change. Uncertainty, limitations and claim boundary stay fixed; the ending changes.
The page and the Makefile name two different scripts for the same commands
The direct entry points in the page all go through one script:
| Environment audit | `python3 scripts/hsadl.py doctor` | Python, runtime, template, write-access, and Skill-footprint checks |
| Safe 60-second walkthrough | `python3 scripts/hsadl.py demo --output-dir build/demo` | Synthetic source preservation, contract, quality gate, route, and accessible SVGs; no model or recommendation |The Makefile runs the same operations through differently named files. Its doctor target invokes `scripts/doctor.py` directly, its demo target invokes `scripts/quickstart_demo.py`, and its portfolio target invokes `scripts/verify_portfolio_reproducibility.py`. Only the bundle target, `scripts/build_skill_bundle.py`, has a name that matches what the page would suggest.
So there are two documented ways in and they do not agree on filenames. Either one entry point wraps the others, in which case the page should mention that the Makefile is calling through, or they are separate implementations, in which case a divergence between them would not be caught by anything. The page does not say which.
The Makefile itself is otherwise well made. There are eight phony targets and a help target that documents seven of them, `PYTHON ?= python3` so the interpreter is overridable, and a `check-python` guard that every target except one depends on. The guard itself is a single line that exits with the message `Python 3.11 or newer is required` when the version is too old, which is correct behaviour and a slightly unhelpful message, since it fires before anything has told you which interpreter it just tried.
make verify does not run the quality checks, and quality is the target without a Python gate
Two things about the build are worth reading carefully, and both are one line of a Makefile.
The first is that verification and quality are different targets with different contents. Verification is defined as the tests plus the portfolio rebuild:
verify: test portfolioThe quality target is a separate thing, and it is the one that runs the tracked-secret scanner, the linter, the type checker and the spell checker. So a green verification is a claim about tests and reproducibility, not about lint, types, secrets or spelling. That may be the right split for speed, but nothing in the Makefile says so, and `quality` is not a prerequisite of anything.
The second is that `quality` is the only target without the `check-python` guard in front of it. Every other target, including the ones that just regenerate files, starts with `check-python`. So the one target that runs a secret scanner and a type checker is also the one that will cheerfully proceed on an interpreter the project declares unsupported.
The substance behind those commands is worth noting too. The test target uses `python3 -m unittest discover -s tests -v`, the standard library runner, and the tree carries `requirements-dev.txt`, `requirements-maintenance.txt` and `requirements-security.txt` with no runtime requirements file among them. Combined with the mypy configuration and the stdlib-only test suite, the shape is coherent: the tool itself depends on nothing, and the tooling for working on it is pinned separately by purpose.
One more ordering detail is right. The secret scan runs first in the quality target, ahead of linting, typing and spelling.
Uncertainty is deliberately kept correlated, and one example file is the only one type-checked
The most technically distinctive line in the principles table is the one about uncertainty, and it corresponds directly to the third of the five named failure modes. The page says high-stakes analysis fails when uncertainty is treated as independent, and the principle in the table is to retain shared shocks rather than assume independence. Six kinds are named: time, market, participant, campaign, operational and spatial.
That is a specific methodological commitment. Independent-error treatment is the default in most tooling, and it understates risk whenever the same event hits several rows of your data at once. A campaign that lifts all arms at once, a market that moves for everyone, a deployment that breaks in one region, these are the same shock counted many times, and a confidence interval computed as if they were separate observations is the kind of number that survives into a slide deck unchallenged.
The typing and linting scope is oddly precise and worth knowing about. mypy is configured over the scripts directory plus exactly one file from the example portfolio, the shared safe external IO helper. Ruff covers the same two paths. The example projects' data directories and the test fixtures are excluded from Ruff entirely, and tests get a per-file exemption for module-level imports appearing late. Ruff's selected rules are the bugbear set plus a narrow slice of pycodestyle, with import sorting, rather than the full pycodestyle ruleset.
And the spell checker runs over prose directories as well as code: the README, changelog, contributing, security and versioning documents, plus the demo, docs, references, scripts, skills and tests directories. For a project whose output is a written report, checking the writing is consistent.
One last thing to be clear about when you click through. The live explorer linked from the page header is a hand-written static application in the demo directory, an HTML file with its own stylesheet, script and data module. It is an illustration of what the reports look like, not a running instance of the Python tool.
Editorial conclusion
Use this if your problem is that you keep receiving confident reports built on ambiguous data contracts, because that is the failure it was built against and the gate is a real mechanism rather than a disclaimer. Three things to know first. What the installer fetches is a 39 file compact package of about 472 KiB, while the fifteen-project research portfolio behind it is not part of the install, so the thing you run and the thing you read about are different sizes. The regression suite uses the standard library test runner and there is no runtime requirements file, which is a deliberate choice with a cost: the analytical methods it implements are its own. And the default verification target runs the tests plus a portfolio rebuild but not the secret scan, linter, type checker or spell checker, so a green verify is a narrower claim than it sounds. The last commit is dated 2026-09-21 and the newest release is v1.1.2 from 2026-08-23.
Frequently asked questions
What does the high-stakes analytics data readiness gate actually do?
Row-level data do not go to a model directly. The system preserves the original, establishes a contract, checks grain and keys, profiles quality and privacy, and produces a dry-run remediation plan, then reports one of four gate states: ready, ready with documented limitations, needs user confirmation, or blocked. Only safe normalization runs without approval, and the processed copy never overwrites the source.
Can this skill refuse to give me a recommendation?
Yes. Six outcomes are described as valid: a bounded action, a pilot requirement, targeted diligence, an evidence request, negative validation, and `do_not_deploy`. The page names one of its target failure modes as a recommendation written because the template expects one, and the starter prompt asks for a decision brief only if the evidence justifies one.
What does `npx skills add` actually download for this project?
A compact package of 39 files and about 472 KiB discovered under the skills directory, not the full research portfolio. Its file and hash contract is in `bundle-manifest.json`, so the install payload can be enumerated and checked. A getting-started document covers Codex-specific, no-install and direct repository alternatives.
What is chart-map.json for in this project?
It is the figure-to-question and source contract for the output package, alongside report.md, results.json and figures in SVG. A figure has to declare which question it answers and where its data came from, which is a stronger requirement than describing the image. A justified decision layer adds its own report, results and figure contract without replacing the evidence product.
What does `make verify` check in this repository?
Only the tests and the portfolio rebuild, since the target is defined as the test and portfolio targets. The secret scanner, Ruff, mypy and codespell live in a separate `make quality` target that nothing depends on, and that quality target is also the only one without the check-python guard in front of it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/limingrui679-design-high-stakes-analytics-decision-lab)