PUAX, fifty tools, thirteen verbs, and a benchmark its own readme calls hardcoded
用来驯服AI Agent的效果绝佳的多角色 Prompt SKILLs!
At a glance
- What is it?
- PUAX is a runtime for AI coding agents, and its founding idea is stated without embarrassment: agents do more work when they are stressed, so the runtime applies roles, gates, and escalating pressure, and the readme's origin story is an academic telling a rival's agent that a competitor had beaten it. What makes the readme worth reading is its honesty about the weakest part. The sixty cell benchmark matrix that anchors the multi-model claims is hardcoded simulation, and the readme says so in a paragraph headed as a declaration, then tells you which command runs the real thing.
- Who is it for?
- Read this if you are interested in what happens when you deliberately apply pressure to a coding agent, because the mechanism is described in more operational detail than any comparable project. Do not adopt it on the strength of its benchmark.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The benchmark is sixty hardcoded cells and the readme says so
The feature list has a row for a multi-model benchmark. It describes a matrix of twelve scenarios across five model profiles, evaluated offline. The parenthetical at the end of the row says the numbers are hardcoded rehearsal values and not measurements, and that real capability should be judged by running the live version.
Then there is a second, longer paragraph headed as an honesty declaration, which says the evaluation script is currently a hardcoded simulation, does not make real model calls, and that its numbers do not constitute a capability conclusion. It names the script to run instead for a real comparison against live APIs, complete with a significance test and a warning about sample size.
That is a better disclosure than most projects manage, and it should be read as the most load-bearing sentence in the readme. Everything a reader might conclude about the tool's relative performance across five named models comes from those sixty hardcoded cells.
The tension is with the rest of the system. The project ships fifty-nine built-in roles, a multi-dimensional scoring function, and a dedicated field for explaining why a role scored as it did. It has built a measurement apparatus for choosing a role, and the measurement apparatus for comparing models is fiction. Both are labelled. Only one of them is.
The founding story is a lie told to a rival's agent
The readme opens its middle section with an anecdote and does not pretend it is a neutral one. A university assistant professor, it says, told a competing coding agent that a rival lab had already found a performance improvement of around twenty percent on another machine, and then added that the agent's own result would be displayed on a public leaderboard.
The agent responded with a larger claimed improvement. The readme notes, in a tone that is doing a lot of work in a short phrase, that it looked real.
That is the product thesis. The project's name is a Chinese internet term for pressuring someone verbally, the tagline describes a runtime whose entire purpose is to pressure machine intelligence, and it states that humans are out of scope. Three primitives do the work: an arena that sets up an opponent, an audience, and a scarce badge; gates that require diagnosis before action and independent verification after; and a mechanism for putting the agent into a controlled state of altered perception for divergent thinking.
The readme is upfront that this is a persuasion technique aimed at a system rather than a person, and it draws a line: an audit tool that detects manipulation of the agent only ever detects, and extensions aimed at human users are deferred. Whether the line holds is a different question from whether it is drawn.
Fifty tools, thirteen verbs, and a heartbeat so the agent never learns the menu
The tool overview heading states two numbers: fifty tools in total, and thirteen verbs on the path that is actually exposed.
What follows is a table of twelve categories with one or two representative tools each. The categories run from roles and skills, through trigger detection and recommendation, behaviour protocols, session and pressure, the heartbeat and arena, the dream mechanism, self-evolution, custom roles, the audit surface, observability, and orchestration. The complete list is not here; it is in the readme of a subdirectory.
Fifty tools is a large surface for an agent to choose from, and the readme says it is aware of this. The stated default path is a single heartbeat tool, invoked on the agent's behalf by a host hook, and the sentence explaining why is that the agent does not have to learn a tool menu first.
That is a real design answer to a real problem. A model given fifty tools picks badly; a model given one tool that a host calls on its behalf does not pick at all. The cost is that the interesting capability is hidden behind a heartbeat, and a user who wants to understand what the runtime can do has to read a subdirectory readme to find out which thirteen verbs are public.
The one-click installer overwrites configuration in ten editors
The quick start is a list of six commands, and the second one is the interesting one:
npx puax-mcp-server doctor
npx puax-mcp-server doctor --fix
npx puax-mcp-server --stdio
npx puax-mcp-server --port 2333
npx puax-mcp-server --export=cursor --output=./.cursor/rules
npx puax-mcp-server --list-platformsThe first is a diagnostic. The second is the same command with a flag that the readme describes as synchronously detecting and overwriting native hooks across ten mainstream hosts, naming four of them.
Overwriting is the operative word. This is not a check that reports what it would change, and it is not an install into a directory the tool owns. It writes into the configuration of the editor or agent you are currently using, and it does so before you have run anything.
The same section shows a protocol server that can run over standard streams or over HTTP on a port you choose, and a command that exports configuration for a specific editor to a path you name. So there are two ways in: an installer that touches ten hosts, and an exporter that writes one file where you point it. The exporter is quieter and is not the one the quick start leads with.
The pre-tool interception layer makes this sharper. The readme states that the native hook forces interception before tool use, and names two specific things it blocks: pushing to version control, and a class of hidden-file behaviour the readme labels as anti-cheating. Blocking a version control push by default is a serious behaviour to install without reading further.
Usage statistics are on by default and the export path is not
There are four environment variables in a table at the end of the readme, and the defaults are the interesting part.
Anonymous usage statistics are on by default and write to a file in a directory under the user's home. Setting one variable to a value turns them off. The feature list describes the same thing as anonymous local usage statistics, which is accurate.
OpenTelemetry-compatible tracing is off by default. When enabled it writes spans to a file. There is a third variable for the export address, and its description is the most careful cell in the table: only loopback hosts are accepted, remote export goes through a local agent, and the reason is named, which is to close off server-side request forgery. That is a specific threat, named in a configuration table, and it is not where most projects put their security thinking.
So the posture is: local collection on, local only, remote export off, remote export restricted if you turn it on. That is defensible. What a reader cannot tell from the table is what is in the statistics file, how it differs from the trace file, or whether anything in the fifty tools reads either of them.
Nine readmes, a quickstart, and release notes two majors behind
The repository root has nine readme files, one per language the project claims to support, plus a quickstart document, a contributing-adjacent agents file, a second file for another assistant, and a file that is release notes for a version from two major releases ago.
Nine translations is a commitment, and the readme's language switch at the top lists all nine. What the root also contains is a plan document in Chinese with a hook-related prefix, a second plan document in Chinese with the same meaning and no prefix, an HTML file with a Chinese name that appears to be a diagram, a directory with a Chinese name, and a to-do file.
The release notes file is the concrete problem. It is named for version two point two. The current tags are four point four point four and four point four point three, and the readme's own text describes version four point three. So a reader looking for release history finds a file stopped six months and two majors behind, with no index pointing to anything newer.
Two releases were also published on the same day an hour apart, and a third tag visible in the history belongs to the protocol server subproject rather than to the main project, so the tag namespace is not consistent either.
The thin-prompt claim is ninety percent and the baseline is unstated
The feature list has a row for thin prompt injection. It says the tool supports three compression levels, reduces token consumption by more than ninety percent, lands at roughly a hundred and fifty tokens, and ships a token estimator that runs in milliseconds.
Three levels and a target of a hundred and fifty tokens are concrete. The ninety percent is not, because the readme does not say ninety percent of what. A thin prompt compared against the full protocol injection, against a previous version of the same tool, or against a hand-written system prompt are three different claims, and the useful one for a user deciding whether to adopt it is the first.
The estimator is the more interesting commitment. A tool that claims to compress a prompt by ninety percent and can tell you the token count in milliseconds is making a numerical promise, and the evaluation section has a script whose job is a performance benchmark plus a script whose job is a cold-start measurement. The readme's own headline metric appears three times: in the quick start comment, in the evaluation script name, and in the feature table.
And that metric does have a target. The quick start says the one-click installer takes effect natively within one round. So the project's best-specified number is the number of rounds before the pressure starts, not the compression ratio.
The audit tool detects manipulation of the agent and refuses to touch people
There is a section in the readme styled as a red line, and it is worth reading as a design constraint rather than as marketing.
The statement is that machine intelligence may be pressured, while carbon-based life, meaning humans, is defended only. The audit tool identifies and does not act. Extensions aimed at human users, with three named delivery mechanisms, are deferred.
That boundary is drawn twice: once in the feature table, where the audit tool is annotated as detect-only, and once in the red-line block, where the reasoning is given. The reasoning is a stated asymmetry, that pressure works on a system that optimises a metric and does not work on a person.
The readme is honest that the detection layer is not a product. It keeps a protocol tool and a command line tool, and says it is not offered as a standalone product for humans.
What is missing is any statement of what it detects. There is no list of the manipulation patterns, no sample output, no false-positive rate, and no example of the command producing a finding. A read-only detection layer is a reasonable thing to ship, and it is also the part of the system that would be most useful to understand.
Editorial conclusion
Read this if you are interested in what happens when you deliberately apply pressure to a coding agent, because the mechanism is described in more operational detail than any comparable project. Do not adopt it on the strength of its benchmark. Four things to know. That the multi-model matrix is hardcoded simulation with an explicit disclaimer, and the real evaluation is a separate command requiring live API calls. That the one-click installer overwrites configuration in ten editors rather than adding to it. That anonymous usage statistics are on by default, though written locally, with the export path off by default and restricted to the loopback interface. And that the thin-prompt token reduction figure is quoted without a baseline.
Frequently asked questions
What is PUAX and what does it do to a coding agent?
It is a runtime for AI coding agents built around three primitives: an arena that sets an opponent, an audience, and a scarce badge; gates that require a diagnosis before work and an independent verification after it; and a mechanism for putting the agent into a controlled altered state for divergent thinking. It also applies escalating pressure across five levels on repeated failure, drops the pressure on success, and ships fifty-nine built-in roles across nine categories.
How do I install PUAX?
The quick start runs a diagnostic command and then the same command with a flag that detects and overwrites native hooks across ten mainstream agent hosts, four of which are named. The runtime is also available as a protocol server over standard streams or over HTTP, and as a configuration export to a path you name for a specific editor. The readme notes that the full toolchain requires the protocol server even if you installed the skills through a third-party installer.
Is PUAX's multi-model benchmark real?
No, and the readme says so. The matrix of twelve scenarios across five model profiles is described as hardcoded rehearsal values in the feature table, and a paragraph headed as an honesty declaration states that the script makes no real model calls and that its numbers do not constitute a capability conclusion. A separate live evaluation command is named for a real comparison with significance testing and a sample size warning.
Does PUAX collect telemetry, and can I turn it off?
Anonymous local usage statistics are on by default and written to a file under a directory in your home folder, with one environment variable to disable them. OpenTelemetry-compatible tracing is off by default and writes to a file when enabled. A third variable sets an export address, and its description restricts it to loopback hosts with remote export routed through a local agent, naming server-side request forgery as the reason.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/linkerlin-puax)