Model or dataset
Prysai/Prysai-LLM-Playbook avatar
Prysai/Prysai-LLM-Playbook

Prysai LLM Playbook: the curriculum that disqualified its own eval data

An evidence-led, eight-locale LLM playbook: a transferable core, the Codex flagship track, and adapters for ChatGPT, Claude Code, Gemini, DeepSeek, and Grok.

437 stars7 forksPythonNOASSERTION

At a glance

What is it?
An eight-locale LLM curriculum built around one five-step loop and three learner-written artefacts, where the unusual part is the evidence ledger: a collected model-output packet the project marked analysis-ineligible after its own integrity review found the prompt hashes did not bind the bytes.
Who is it for?
Adopt this if you want a written method for inspecting model output and you are the kind of reader who reads an evidence table before a curriculum, because the five-step loop and the three artefacts are concrete and the ledger tells you precisely what has and has not been measured.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 17 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One loop, five units, no platform attached

The whole method fits on one line, and the project puts it near the top:

text
define the task → choose a bounded action → inspect the result → keep evidence → state the limit

The default route turns that into five units: explain what a language model is and is not, write a small request carrying goal, context, limits and output shape, identify omission, invention, forced ambiguity and overconfidence, check and minimally repair an answer while stating one limit, then repeat the method on an unseen task with no complete prompt template supplied. Each unit has to leave a learner-authored artefact, and the project is explicit about what does not count: a copied prompt, a polished model answer, or a green structural check. The three artefacts the route targets are a bounded task card, a checked result record, and a transfer attempt. What makes the positioning unusual is the refusal to extend any of this upward: the advanced material on Codex, Skills, Agents and named-platform adapters is described as useful reference whose current structure is not learner evidence and does not establish cross-platform equivalence.

The project marked its own collected outputs unusable

The evidence ledger has three rows and the middle one is the interesting one. A Shift Handoff output packet was collected, containing 18 de-identified fictional model outputs. Its input-integrity review then found that the historical prompt hashes do not bind the prepared Windows prompt bytes. That single finding reclassifies the whole packet as captured, unscored and analysis-ineligible, with a list of conclusions a reader may not draw from it that runs from time and percentage through efficiency, productivity, learning, safety and accuracy to model quality. In other words the authors collected data, checked whether the data could be trusted to mean anything, concluded that it could not, and published that conclusion in the same table as their results. Very few projects put a disqualification of their own material in the results section rather than in a changelog, and if you are trying to evaluate this repository on rigour rather than on the size of its claims, that row is the most informative sentence in it.

Observed means seven checks on one Windows worktree

The row above it is the only positive claim, and its scope is deliberately narrow. The record is seven local checks run five times sequentially, with raw milliseconds and a chart. The permitted conclusion is that these named engineering checks were stable in one current local Windows worktree, and the ledger states outright what that is not: not a speed result, not a Skill result, not a learner result, not a safety result and not a model result. So the strongest empirical statement this project can make about itself is that its own test suite did not vary across five runs on one machine. That is a real and useful thing to know, and it is worth holding onto while reading the rest of the table, because the third row is titled Unknown and covers learner completion, transfer, real-work productivity and IQ, with the note that no conclusion is available and that the Playbook does not measure or claim IQ improvement.

Status is a controlled vocabulary, not an adjective

The way this project writes about itself is the most transferable thing in it. Status values are set in backticks and behave like enum members: the project is `candidate`, learner completion, transfer and long-term retention are `not_run`, and the eval packet is `Captured, unscored, analysis-ineligible`. Warnings are folded into the prose the same way, as in the instruction not to stop at a plausible output, followed by three questions to ask instead: what changed, what was checked, and what remains unproven. The measurement research record behind the table defines task-scoped completion, rework, time and fixed-rubric measures, and sets conditions on any future result: it must keep its commit, its conditions, its raw de-identified records and any scorer disagreements, and even then it remains a small descriptive observation rather than a universal efficiency claim. Governance sits alongside, with a core course contract, a scope freeze and a content inventory kept in YAML.

Eight locales and a script that audits their language

The root carries eight locale entry points as separate files: English, Simplified Chinese, Spanish, Japanese, Korean, German, Traditional Chinese and French, with `README.md` acting as GitHub's compact English entry and `README-EN.md` as the detailed source. Each one carries machine-readable metadata in an HTML comment, naming content id, locale, language, default locale, a compatibility entrypoint and a canonical source, which is how a static site can route readers without parsing prose. English is the default, and the switcher is annotated with the honest caveat that eight entry points are registered while translation review and learner evidence remain in progress. That caveat has tooling behind it: the repository's test surface includes `audit_translation_language_quality.py`, invoked as `python -X utf8 scripts/audit_translation_language_quality.py --verbose`, where the UTF-8 mode flag is a direct consequence of checking eight files that include Simplified and Traditional Chinese, Japanese and Korean.

The npm package exists only to drive Playwright

The repository's package manifest is named `prysai-llm-playbook-browser-checks`, marked private, and requires Node 20 or newer. Its single entry point is `npm run test`, which runs `python scripts/run_tests.py`, so the package whose purpose is browser checks launches a Python test runner. The one dependency declared is Playwright, pinned to 1.62.1, and the remaining scripts are a family of visual checks: `browser_smoke.mjs`, `reader_visual_smoke.mjs`, `visual_asset_geometry_smoke.mjs` and `visual_guide_smoke.mjs`, one per visual surface. Read together, that is a curriculum repository that treats rendered output as something to be tested, which is why there are four separate visual smoke scripts and a geometry check for assets. The examples directory is sparse in a way that suggests deliberate selection rather than incompleteness: labs numbered 001, 008 and 013, plus `skill-sandbox/` and `universal-seam-v1/`.

Two licences, a DCO, and a repository that calls itself a book

Licensing is split by artefact and documented rather than implied: curriculum text and teaching assets are CC BY 4.0, scripts and tooling are Apache-2.0, and any individual file may state otherwise. Two licence files at the root, `LICENSE` and `LICENSE-CODE`, plus a licensing boundary document under `docs/sources/`, exist to make that boundary checkable. Alongside them sit the usual governance files, and one that is worth naming: `DCO.md`, which signals sign-off commits rather than a click-through contributor licence agreement. The repository describes itself as a textbook and reference library rather than a required menu, and instructs you not to enter the Codex, Skills or professional tracks until the core route says to continue. Its structure backs that up, with `book/` holding the routes and chapters, `site/` for the guided reading experience, `skills/`, `tasks/`, `evals/`, `tests/` and `docs/`. The last push to main was on 2026-09-23, and the single release is v0.1.0-alpha from 2026-08-16.

Editorial conclusion

Adopt this if you want a written method for inspecting model output and you are the kind of reader who reads an evidence table before a curriculum, because the five-step loop and the three artefacts are concrete and the ledger tells you precisely what has and has not been measured. Do not adopt it expecting research backing, named-platform equivalence or a finished course, since the project labels itself candidate, marks learner completion and transfer as not_run, and states that its platform adapters do not establish equivalence. Verify first which locale you are reading, because eight README entry points are registered while translation review is still in progress.

Frequently asked questions

Is the Prysai LLM Playbook finished?

No. It describes itself as a candidate: its structure and static checks exist, but learner runs, transfer runs, repeated evaluations and independent review are still pending. Learner completion, transfer and long-term retention are all marked `not_run`.

What method does the Prysai LLM Playbook teach?

A single loop: define the task, choose a bounded action, inspect the result, keep evidence, state the limit. The five-unit route turns it into explain, initiate, identify, repair and transfer, with the last unit repeating the method on an unseen task without a complete prompt template.

Why did Prysai mark its Shift Handoff outputs analysis-ineligible?

The packet contains 18 de-identified fictional model outputs, and its input-integrity review found that the historical prompt hashes do not bind the prepared Windows prompt bytes. Without that binding the collection cannot be compared, scored or aggregated, or used to infer any time, percentage, benefit, productivity, learning, safety, accuracy or model-quality result.

Does the Prysai LLM Playbook measure intelligence improvement?

No, and it says so directly. Learner completion, transfer, real-work productivity and IQ are all listed as unknown with no conclusion available, and the project states that it does not measure or claim IQ improvement.

How is the Prysai LLM Playbook licensed?

Split by artefact. Curriculum text and teaching assets are CC BY 4.0, scripts and tooling are Apache-2.0, and a file may state otherwise. The root carries both a `LICENSE` and a `LICENSE-CODE` plus a licensing boundary document under `docs/sources/`.

Official sources

  1. Issues
  2. Project website
  3. Prysai/Prysai-LLM-Playbook on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/prysai-prysai-llm-playbook.svg)](https://hysenlabs.com/projects/prysai-prysai-llm-playbook)