Model or dataset
tjboudreaux/cc-thinking-skills avatar
tjboudreaux/cc-thinking-skills

cc-thinking-skills' only positive result is below its own usefulness margin

28 eval-informed mental models and critical-thinking skills for Claude Code, GitHub Copilot, Codex, Cursor, and other Agent Skills-compatible tools

1,408 stars166 forksJavaScriptMIT

At a glance

What is it?
A catalog of 28 reasoning skills for coding agents, installed as a package, as a plugin, or copied into a directory, with a meta-skill that routes to one or a few of them. What makes it unusual is the evidence section: it reports no retention verdict, a portfolio with zero model calls and an unmeasured result, and one positive row that falls below the project's own margin for usefulness. The page then tells you to read the audit before repeating any of its claims.
Who is it for?
cc-thinking-skills is worth installing for its content rather than its evidence, because the catalog itself is competently assembled and the evaluation apparatus around it is more careful than the headline suggests. Two things to hold onto.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 57 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The one positive row sits below the project's own usefulness margin

The evidence section is four bullet points and three of them are negative.

- 28 active skills, all manual-only - no automatic-retain verdict - `portfolio-v1` with zero model calls and an unmeasured result - a provisional `thinking-scientific-method` row at +4.0 percentage points

Then the sentence that governs how to read all four: the positive row is below the five-point utility margin and has evidence gaps, so treat it as directional evidence rather than an accuracy claim. So the entire empirical case for this catalog is one skill, moving four points, in a sample the project itself calls provisional, against a threshold the project set at five. Most projects publish a number and hope nobody asks for the denominator. This one publishes the denominator, the threshold, and the verdict that it failed. That is a stronger position than a passing number would be, because you can see exactly what is and is not known.

The page tells you to read the audit before repeating its numbers

This instruction appears early, in the section arguing for the catalog rather than the evidence.

> Read the [catalog audit](analysis/AUDIT.md) before making performance claims.

It is an unusual thing for a README to say, because it points the reader at the file that qualifies the page. The same pattern holds in the evidence section, which links not one but two artefacts: a decision-ready audit document and an evidence registry in a machine-readable format. So the claims can be checked against the raw record rather than against prose, and the layout section confirms that the analysis directory is a first-class part of the repository rather than a scratch file. The repository summary agrees, listing an analysis directory for the evidence registry and the audit as one of its five top-level directories. It is the difference between a project asserting a number and a project publishing the basis for it.

A portfolio with zero model calls measured nothing, and the page says so

The second bullet deserves reading on its own. A portfolio run was executed, and it is recorded as having made zero model calls and having an unmeasured result. In other words, an evaluation was defined, run, and produced nothing, and that nothing is printed in the feature list rather than quietly dropped. The likely cause is visible elsewhere on the page: the routing evaluation requires an authenticated external command-line tool, which a portfolio run without credentials would not have. So the honest chain is that the harness exists, the local structural gate runs without credentials, the routing gate does not, and a portfolio run attempted without them yielded an empty measurement. Most projects would either not run it or would report it as a failure. Reporting it as unmeasured is the more useful answer, because it tells you which piece is missing.

The router is allowed to answer NONE

The meta-skill at the top of the catalog is a router, and its stated contract is that it can return one of three things: no framework at all, one frame, or up to three complementary ones. Being permitted to return nothing is the design detail that matters. A catalog of twenty-eight reasoning frameworks invites an agent to reach for one every time, which turns a library into a tic, and a router that can decline is the mechanism that prevents it. The catalog supports the same restraint from two other directions. One skill exists to classify what kind of cause-and-effect situation you are in before you pick a method, and another exists to set a good-enough threshold and stop searching. A set of skills about rigorous reasoning containing a skill whose explicit job is to make the agent stop escalating rigour is a coherent position rather than a contradiction.

The router's invocation string changes depending on how you installed it

The page gives four invocation examples with bare skill names, then adds that the plugin uses a namespaced identifier as the exact router name. So a prompt that works after a package install does not work after a plugin install without editing, and there is no single string that is correct in both cases. That is a small piece of friction in a project whose entire value is that you can ask for a frame by name, and it is the kind of detail that only shows up if you install both ways. The four install routes compound it: a package manager command with an all-skills flag, a plugin marketplace add followed by a plugin install, a clone-and-copy into a dot-directory, and the two locations the host will read from. Four routes, four conventions, and a naming difference in the middle.

The routing evaluation needs an authenticated tool nothing else depends on

Three commands are given for checking the project, and they are not equivalent.

bash
node scripts/validate-skills.js
bash
EVAL_RUN=local node evals/run-structural.js
bash
EVAL_RUN=local node evals/run-routing.js

The first validates that each skill has the right structure and needs nothing. The second is the local structural gate and runs on plain Node. The third is the routing evaluation, and the page specifies that it runs with an authenticated external command-line tool. That is the one piece of the verification story that cannot be reproduced from a clone, which is consistent with the portfolio run having produced zero calls. It also means a contributor cannot check the claim the project cares about most, namely that the router picks sensibly, without credentials for a third-party tool. The project is explicit about this, and it points contributors at a harness guide and the contribution guide before changing the catalog.

Most of the repository is evaluation machinery rather than skills

The top level holds a skills directory, a scripts directory for validation, an evals directory described as structural, routing and outcome evals, an analysis directory for the evidence registry, and plugin metadata. Plus an experiments directory that the summary mentions but the layout section does not. So one of five or six top-level directories is the actual product and the rest is the apparatus built to measure it, including a studies subdirectory for study artefacts and a note that external datasets are used within their licence limits. That ratio is unusual for a catalog of prompts and is the clearest signal of what this project actually is: the skills are the payload, and the interest is in whether they can be shown to work. The catalog section is arranged into five groups whose sizes are uneven, with the largest holding eight skills against a smallest of two, and the stated count of active identifiers matches the directories on disk, which is checkable.

Editorial conclusion

cc-thinking-skills is worth installing for its content rather than its evidence, because the catalog itself is competently assembled and the evaluation apparatus around it is more careful than the headline suggests. Two things to hold onto. The description calls these eval-informed skills, and the audit behind that description reports no verdict at all, so read the audit rather than the tagline. And the routing evaluation requires an authenticated external command-line tool that the rest of the project does not depend on, which means the strongest claim the project could make about itself is the one you are least likely to reproduce. None of this makes the skills bad. It makes them unevaluated.

Frequently asked questions

What does the cc-thinking-skills catalog contain?

Twenty-eight agent skills grouped into five areas: routing and composition, diagnosis, decision and evaluation, creation and improvement, and risk and execution. Each skill is a directory containing a markdown file with front matter, and the page states that the active identifiers match the directories on disk.

Has the cc-thinking-skills catalog been evaluated?

Not to a positive result. Its own audit reports no automatic-retention verdict, a portfolio run with zero model calls and an unmeasured result, and one provisional row at four percentage points, which the page says is below its own five-point utility margin and should be treated as directional only.

How do I install the cc-thinking-skills catalog?

Four ways: a package-manager command that can install all skills without prompts, a plugin marketplace add followed by a plugin install, a clone-and-copy into a dot-directory under your project, or loading from either of two directories the host reads from. The router's exact invocation string differs between the plugin and the other routes.

What can the thinking-model-router return?

Three outcomes: no framework at all, a single frame, or up to three complementary frames. The page's examples show it being used to pick a frame for a problem, to localise a bug, to classify an architecture decision, to stress-test a plan, and to find a bottleneck.

Can I run the cc-thinking-skills checks myself?

Structure and the local structural gate run on plain Node from a clone. The routing evaluation additionally needs an authenticated external command-line tool, which is why the portfolio run recorded on the page produced zero model calls and an unmeasured result.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. tjboudreaux/cc-thinking-skills on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tjboudreaux-cc-thinking-skills.svg)](https://hysenlabs.com/projects/tjboudreaux-cc-thinking-skills)