Model or dataset
paralleldrive/riteway avatar
paralleldrive/riteway

Riteway: a test framework whose real feature is the shape of the assertion, plus a CLI that scores agent prompts

Simple, readable, helpful unit tests. Optimized for AI Driven Development.

1,195 stars40 forksJavaScriptMIT

At a glance

What is it?
A Node testing library organized around four letters and five questions, shipping a riteway ai subcommand that runs AI prompt evals against a pass-rate threshold using OAuth-authenticated agent CLIs.
Who is it for?
Riteway is two products that happen to share a repository. The first is a small assertion library whose argument is that the shape of a test determines whether it is worth reading, and that argument holds regardless of who writes the test.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four letters that constrain how you write a test

Riteway's README opens with an acronym rather than an API: Readable, Isolated or Integrated, Thorough, Explicit. The claim that follows is unusually blunt for a testing library. Riteway forces you to write Readable, Isolated and Explicit tests because that is the only way you can use the API. Thorough is the odd one out: it is not enforced, and the README argues it becomes easier once assertions are simple enough that you want more of them.

The enforcement argument is the interesting part. Most assertion libraries will accept a one-liner that checks a boolean and moves on, because refusing it means fighting the tool. Riteway's structure makes the minimal form require the parts a good test needs, which is why the framework describes itself as producing a good bug report when a test fails.

Underneath that is a five-question framework the README says every unit test must answer: what is the unit under test, what should it do in prose, what was the actual output, what was the expected output, and how do you reproduce the failure. The given and should pairs map onto the first two. Questions three through five are the diagnostic ones, and they are the reason the library bills itself on failure output rather than on assertion ergonomics.

For AI-driven development the same structure is the argument. The README claims a minimal API surface reduces agent confusion and hallucinations, and that concise syntax saves context window. Both are reasonable: an agent writing a test against a two-function API produces better output than one navigating a large assertion vocabulary, and fewer tokens per test means more of the window holds actual source code.

Install is one dev dependency and one npm script

The install path is short. Riteway requires Node.js 16 or newer and native ES modules, so `"type": "module"` has to be present in your `package.json`. Then:

shell
npm install --save-dev riteway

The test script is where the design shows. The basic form is a glob over your test files:

json
"test": "riteway test/**/*-test.js",

For a project that mixes plain module tests with JSX component tests, the README shows a dual runner, chaining the Riteway entry point to Vitest:

json
"test": "node source/test.js && vitest run",

And because Riteway supports the full TAPE-compatible usage syntax, the advanced form pipes through `nyc` for coverage and `tap-nirvana` for colorized output with source line identification and diffs:

json
"test": "nyc riteway test/**/*-rt.js | tap-nirvana",

Three takeaways from those three lines. Riteway does not insist on owning your test process; it can hand off to Vitest for the parts it does not want. It inherits TAPE compatibility, which means existing TAPE-style tests can run under it. And the filename conventions in the globs, `-test.js` versus `-rt.js`, suggest the project expects teams to sort their tests into the plain assertion suite and the fuller TAPE suite deliberately, rather than converting everything at once.

JSX support does require a build tool that transpiles JSX, which is the main setup cost for a component-heavy React project.

The riteway ai subcommand turns prompts into assertions

The second half of the project applies the same given/should shape to AI behavior. You write a `.sudo` file in SudoLang syntax, and each `- Given ..., should ...` line becomes an independently judged assertion:

shell
riteway ai path/to/my-feature-test.sudo

A file looks like a small script. It imports a spec, assigns a `userPrompt` string, then lists what the response should contain:

code
import 'path/to/spec.mdc'

userPrompt = """
Implement the sum function as described.
"""

- Given the spec, should name the function sum
- Given the spec, should accept two parameters named a and b
- Given the spec, should return the correct sum of the two parameters

The mechanism is that the agent is asked to respond to the prompt, with the imported spec as context, and a separate judge agent scores each assertion. Because scoring is per assertion rather than per test, a run reports which specific expectations held across passes, which is the granularity you want when you are tuning a prompt.

Agent support covers Claude, Cursor and OpenCode, selected with `--agent`. Authentication is OAuth through each vendor's own CLI, `claude setup-token` for Claude and `agent login` for Cursor, with no API keys passed through Riteway. That is a sensible design choice for a tool whose whole purpose is to run against whichever agent you use, though it does mean Riteway inherits whatever each CLI does with your session.

For agents outside the built-in list, `--agent-config` takes a flat JSON file with `command`, `args` and `outputFormat`, which is how you would point the harness at anything else that can be invoked as a subprocess.

Defaults decide whether an eval run is cheap or noisy

The default parameters are worth reading carefully because they define the cost of a single run. By default Riteway executes 4 passes per assertion, requires a 75% pass rate to count as passing, uses the `claude` agent, runs up to 4 tests concurrently, and allows 300 seconds per agent call.

Four passes is the smallest number that can distinguish a reliably correct behavior from a coin flip, and 75% is deliberately below certainty, which fits a system with genuine run-to-run variance. But the arithmetic matters. Four passes times the number of assertions in your file is four agent invocations per file before you factor in the judge, and the judge call is itself a model call that scales with assertion count. A twenty-assertion eval file is on the order of eighty calls before concurrency.

The flags let you tune both axes. `--runs 10 --threshold 80` is the shape the README shows for a tighter check, and `--timeout` raises the per-call ceiling for agents that think slowly. `--concurrency` caps parallel work, which matters because files run sequentially by design to respect rate limits, while tests within a file can overlap.

Output goes to a TAP markdown file under `ai-evals/` in the project root. Adding `--save-responses` writes a companion `.responses.md` beside it with the raw agent output and per-run judge detail, including passed, actual, expected and score for every assertion. That is the difference between an eval you rerun blindly and one you can diagnose, and it is the flag to reach for first when a result looks wrong rather than merely failing.

The tree describes a project being rebuilt around agents

The repository layout says more than the README does. Alongside the conventional `source/`, `bin/`, `docs/` and `tasks/` directories there are `ai/`, `aidd-custom/`, `.cursor/`, `AGENTS.md`, `vision.md`, `plan.md` and `activity-log.md`. A `tea.yaml` sits next to them. The package name in `package.json` is plain `riteway`, typed as an ES module, and the exports map splits into `.`, `./vitest`, `./bun`, `./match`, `./render` and `./render-component`, each with ESM variants. That export surface tells you Riteway is not only an assertion runner: it ships a Vitest adapter, a Bun adapter, a `match` helper, and two rendering paths.

Release tooling is equally visible, with `.release-it.json`, a `release.js` at the root and a dedicated `RELEASING.md`.

Then there is the discrepancy. The repository was last pushed on September 19, 2026 and is not archived, with roughly 1,200 stars and 21 open issues, so the work is current. But the only published release is v6.1.1 from November 2019, described as minor fixes to TypeScript definitions, documentation and dependencies, while `package.json` declares version 9.3.0. A project on its ninth major version with an active AI tooling roadmap has not cut a GitHub release since 2019.

The reasonable read is that releases moved to npm automation and the GitHub releases list stopped being maintained, which is common and mostly harmless. It does mean the tag history is not a usable changelog. The repository root tells you where the real history lives: `activity-log.md`, `vision.md` and `plan.md` sit alongside the code as first-class files, which is a choice about how a project stays legible when it is being rebuilt rather than incrementally extended.

Where the AI-native framing holds, and where it is a bet

The claims split cleanly into two groups, and it helps to treat them differently.

The assertion-style claims are checkable. Readable, isolated and explicit tests are better tests, and the framework's structure pushes you that way regardless of who writes the test. A human writing tests by hand will find the given/should form natural. Token efficiency is a real property of a small API surface. Nothing here depends on a model behaving a particular way.

The prompt-eval claims depend on a moving target. Whether 75% across 4 runs is the right threshold for a given prompt is a question about agent behavior that changes with model versions, with provider-side system prompts, and with the agent CLIs Riteway shells out to. A threshold tuned in one quarter can be meaningless in the next. The framework does not pretend otherwise, which is why the per-assertion scoring and the saved-responses file matter: they are the observability you need precisely because the thing being measured is non-deterministic and its implementation is not yours.

There is also a practical ordering to respect. Prompt evals are only meaningful if the underlying code has unit tests, because an assertion like should return the correct sum is checking an agent against a specification you have not yet verified in code. The project describes itself as the standard testing framework for AI driven development, and the practical reading is that it wants to be the first layer you write, before the agents get involved.

Editorial conclusion

Riteway is two products that happen to share a repository. The first is a small assertion library whose argument is that the shape of a test determines whether it is worth reading, and that argument holds regardless of who writes the test. The second is `riteway ai`, which applies the same given/should logic to prompts and is the part that depends on agent CLIs behaving predictably across versions. Start with the assertion style on one module, since adopting a test framework is cheap and easy to reverse, and treat the prompt evals as a separate decision with its own cost model: defaults of 4 runs at a 75% threshold per assertion add up fast once a suite has more than a handful of assertions. One thing to settle before trusting either is the release history. `package.json` reads version 9.3.0 while the only published release is v6.1.1 from November 2019, so the gap between what npm installs and what the release page shows is wide. Pin your version deliberately, and read `source/` for the assertion implementation and `RELEASING.md` for how releases are cut before assuming the tags describe the current state.

Frequently asked questions

How do I install Riteway in an existing Node project?

Riteway requires Node.js 16 or newer and native ES modules, so add `"type": "module"` to your `package.json`, run `npm install --save-dev riteway`, and set the test script to a glob such as `"test": "riteway test/**/*-test.js"`. For a project with JSX component tests, the README shows a dual runner that chains your Riteway entry point to `vitest run`.

What does riteway ai do exactly?

It runs prompt evaluations. You write a `.sudo` file in SudoLang syntax with a `userPrompt` and a list of `- Given ..., should ...` lines, then run `riteway ai path/to/file.sudo`. Each line becomes an independently judged assertion, a judge agent scores it across 4 passes by default, and results are written as TAP markdown under `ai-evals/`. Add `--save-responses` for a companion file with the raw agent output and per-run judge detail.

Do I need API keys to use the riteway ai evals?

No. The README states that all agents use OAuth authentication with no API keys needed. You authenticate once through each vendor's own CLI, `claude setup-token` for Claude and `agent login` for Cursor, with OpenCode following its own documentation. Riteway shells out to those CLIs rather than calling model endpoints itself, which also means it inherits whatever each CLI does with your session.

Official sources

  1. Issues
  2. License: MIT
  3. paralleldrive/riteway on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/paralleldrive-riteway.svg)](https://hysenlabs.com/projects/paralleldrive-riteway)