# Research Harness separates a green check from a good answer

> This project treats a research run as a recoverable file tree rather than a conversation, pinning the pipeline contract, hashing every artifact, and refusing to let an approved outline drift after sign-off. Its most valuable section is the one that says what a passing check does not establish: it does not establish that the answer is good, that the retrieval was exhaustive, or that the result holds beyond the cases that were evaluated.

**WILLOSCAR/research-units-pipeline-skills** — Research pipelines as semantic execution units: each skill declares inputs/outputs, acceptance criteria, and guardrails. Evidence-first methodology prevents hollow writing through structured intermediate artifacts.

- Repository: https://github.com/WILLOSCAR/research-units-pipeline-skills
- Stars: 512 · Forks: 39
- Language: Python
- License: not declared
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/willoscar-research-units-pipeline-skills

## A PASS establishes consistency, not quality

The file publishes a three-row table separating three claims that it says are easy to blur, and in each row it names both what a passing check establishes and what it does not. Execution integrity means attempts, state, manifests, hashes and provenance agree, and explicitly does not mean the answer is good. Contract acceptance means the required artifacts satisfy observable workflow checks, and explicitly does not mean scientific truth or exhaustive retrieval. Research quality means usefulness and correctness on realistic inputs, and explicitly does not mean validity beyond the evaluated cases. Then the sentence that matters most for anyone deciding whether to trust the output: the repository implements the first two layers, while the third needs repeated runs, held-out evaluation and expert judgment. Reports are described as using qualified evidence rather than turning every green check into a quality claim. The framing around it is equally careful: the file says up front that this is not a claim about autonomous science, it is infrastructure for making agent-assisted work inspectable, resumable and honest about what has and has not been proven. The diagram behind that shows the whole chain in eight hops, from a goal through a workflow, a pinned contract, recoverable units, artifacts and completion checks, to run evidence and a bounded diagnosis that repairs and reruns rather than starting over.

## The gate came from a measured failure, not from taste

There is a section explaining why one particular check exists, and it is the strongest argument in the file because it comes with a number. The survey writer can bootstrap provisional prose from structured evidence packs and versioned templates, which is a reasonable design until the scaffold survives into the finished document. An early version completed the delivery path and still matched template fragments in 96 of 140 sentences, or 68.6 percent, in the historical course-paper sample. The file states that this failure is now a contract rather than a warning, and names the unit that enforces it: a front-matter writer that checks five sections of the finished document, the abstract, the introduction, the related work, the discussion and the conclusion. That is a narrow gate on a narrow failure, which is the shape of a check that can actually pass. A repository that measures its own output quality failure at sentence resolution, then encodes the fix as a pipeline unit, is doing something rarer than adding a review step.

## The pipeline contract is pinned and a run fails closed

The first of three mechanisms that make the trail useful is a lock file. A versioned lock snapshots the selected pipeline and hashes three things: its inheritance bundle, the skill implementations, and the harness kernel. Then the rule that makes it worth having: an active run fails closed if the pipeline or the kernel drifts, so it cannot quietly continue under different rules than the ones it started with. For a system that may run unattended for hours, that is the difference between a reproducible result and a plausible one. The same discipline is applied to the human checkpoints, where approval is bound to the hashes of the artifacts that were reviewed, which means editing an approved outline, scope or protocol revokes the authorisation rather than leaving a stale signature attached to different content.

## Failure is given an address, and improvement does not edit in place

The second and third mechanisms are about what happens when something is wrong. Completion is evidence-backed, which the file states as a rule rather than an aspiration: a done marker on its own is not success, and the attempt, the required outputs, the artifact hashes, the workflow checks, the manifest and the completion event all have to agree. Failure then has an address. A doctor pass, an audit pass, scorecards and a failure ledger are used to distinguish an observable defect from the surface that owns repairing it, and the improvement command diagnoses rather than rewriting the harness in place. That last constraint is the one worth internalising: a repair tool that edits its own harness can make a failure disappear without making the underlying problem go away, and the file forbids that.

## The install is a source checkout with two runtime dependencies

Getting it running means a clone, a sync against the lock file, and a command, with Python 3.10 or newer and a Python package manager as the stated prerequisites:

```bash
uv run rh goal create \
  --goal "Understand test-time adaptation for robotics and decide what to read" \
  --workflow research-brief \
  --workspace workspaces/robot-adaptation
```

The declared runtime dependencies are two: a PDF reader library and a YAML parser. Everything else in the description, the hashing, the manifests, the event log, the contract checks, the LaTeX and PDF delivery, is standard library. Two details in the packaging are worth noticing. The command a user types is defined in a package named for tooling rather than in the package that holds the research logic, so the entry point and the engine live in different subtrees. And the lint configuration selects five rules, which is a syntax-and-errors gate rather than a style gate, while the test extra pulls in one test runner and that same linter. The distribution version is 0.1.0 and there are no tagged releases, so nothing pins the build for you beyond the dependency lock, which covers dependencies rather than the package itself.

## The workspace is the deliverable

After a run, the value is the tree on disk rather than anything the assistant said. Six named locations carry the whole state:

```text
GOAL.md                  requested outcome and constraints
UNITS.csv                explicit plan and current Unit state
DECISIONS.md             human checkpoints and choices
papers/ + outline/       research evidence and intermediate structure
output/                  deliverable, scorecards, audits, repair reports
.harness/                Run identity, Attempts, Events, hashes, provenance
```

Two of those names do the heavy lifting. The first says what was asked for and under what constraints, which is the only place the original request survives once the run has been going for hours, and the third records the checkpoints a person actually approved. The plan is a spreadsheet, which means the unit list and its states are inspectable and diffable rather than buried in a transcript, and the research evidence and intermediate structure sit in their own directories rather than inside the final document. The hidden directory holds provenance, so the question of what ran and what it hashed is answerable after the fact. The framing the file opens with is the four-step line goal, run, evidence, artifact, and the opening argument is that a long research task can still produce a polished document while leaving the supporting questions unanswered, including which sources back a given paragraph and whether the work can resume tomorrow without reconstructing a conversation. Driving it is a small verb set: create a goal against a workspace, start the run, then ask for status, approve a named checkpoint, resume, or inspect the evidence with excerpts rather than in full. A run advances until it either finishes or stops on a prerequisite it cannot meet, which is the point at which a person is expected to look.

## Seven workflows, and one that is deliberately not executable

The seven deliverables are named files in most cases, not formats: a snapshot, a review, a synthesis, a draft, a report, and for the tutorial a tutorial with an article PDF and slides. In the two coding hosts the file supports, activation is one sentence, asking the assistant to use a named workflow for a stated topic and a stated audience, which keeps the prompt surface deliberately small. A user picks a workflow by the outcome they want, and the skills and units underneath stay hidden until something needs inspecting or repairing. Seven are offered, each with a required starting point and a named deliverable: a brief producing a snapshot file, a single manuscript review producing a review file, an evidence synthesis producing a synthesis file, a survey producing a draft, the same survey delivered as LaTeX and a PDF, an idea brainstorm producing a report, and a tutorial built from a fixed source pack. An eighth name appears and is explicitly excluded from that list, being described as a research-stage Chinese thesis path rather than one of the executable contracts. The input boundaries are stated as refusals: the review unit will not invent a manuscript, the tutorial unit will not invent a source pack, and the evidence unit writes its protocol and then pauses for approval before it retrieves anything.

## Twenty-five top-level entries and no licence file

The repository root holds twenty-five entries for a project whose user-facing surface is one command. Alongside the source, tests, templates, pipelines and a locked dependency file sit two agent configuration directories, one for each of two hosts, an agent instructions file, a context file, a contribution guide, two readmes in different languages, and a skills standard paired with a skills index. One entry is worth flagging on its own: a directory named readme in lowercase, sitting next to the uppercase readme file, and it is where the English usage guides live. The examples directory carries six named cases, three of which end in the word proof, including one named for the residue failure above and one named for a real source rather than a typed input. What is missing from all of it is a licence: the repository records none, and no licence document appears anywhere in the listing, in a project that does publish a contribution guide. Naming runs one layer deeper than the repository does: the readme titles the project Research Harness, while the repository and the distribution are both named after pipeline skills. There is a single badge link, pointing at the workflow file that runs verification, and the repository has 512 stars, 39 forks and one open issue, with the default branch last moving on 2026-10-01.

## Conclusion

The distinction this project insists on is the reason to read it. Most agent frameworks equate a completed run with a good result, and this one publishes a table saying that a PASS establishes internal consistency and contract acceptance and nothing more, then admits it implements only the first two of the three layers it describes. Use it if you run multi-hour research where losing the thread or shipping template filler is a real failure, and read the layer table before you trust any of its green marks. Before you depend on it, check the three things the file itself leaves open: version 0.1.0 with no tagged release means nothing is pinned for you, there is no licence file anywhere in a twenty-five-entry listing that does invite contributions, and the third quality layer, the one that would tell you whether a deliverable is any good, is explicitly left to you.

## FAQ

### What is WILLOSCAR/research-units-pipeline-skills?

It is a Python project that turns a research goal into a file-first, recoverable run rather than a conversation, organising focused skills into explicit workflows, keeping intermediate artifacts, and checking observable contracts. Its declared runtime dependencies are a PDF reader and a YAML parser, and it requires Python 3.10 or newer.

### What does a PASS mean in Research Harness?

The file separates three layers. Execution integrity means attempts, state, manifests, hashes and provenance agree. Contract acceptance means required artifacts satisfy observable workflow checks. Neither establishes that the answer is good, that retrieval was exhaustive, or that the result holds beyond the evaluated cases, and the repository states that it implements only those first two layers.

### How does Research Harness keep a run from drifting?

A versioned lock file snapshots the selected pipeline and hashes its inheritance bundle, the skill implementations and the harness kernel. An active run fails closed if the pipeline or the kernel changes, so it cannot continue under different rules, and a human approval is bound to the hashes of the artifacts that were reviewed.

### Which research workflows does Research Harness offer?

Seven executable ones: a brief, a paper review, an evidence synthesis, a survey, the survey delivered as LaTeX and a PDF, an idea brainstorm, and a tutorial from a fixed source pack. An eighth name is described as a research-stage Chinese thesis path rather than one of the executable contracts.

### Is research-units-pipeline-skills licensed?

No licence is declared. The repository records no licence identifier, and no licence document appears among the twenty-five top-level entries, even though the project publishes a contribution guide and an English and Chinese readme.

### How did Research Harness decide what to check?

From a measured failure of its own. An early survey writer completed the delivery path but left template fragments in 96 of 140 sentences, 68.6 percent, of a historical course-paper sample, and that is now enforced as a contract by a unit that checks the abstract, introduction, related work, discussion and conclusion.

## Sources

- [Issues](https://github.com/WILLOSCAR/research-units-pipeline-skills/issues)
- [README](https://github.com/WILLOSCAR/research-units-pipeline-skills/blob/main/README.md)
- [WILLOSCAR/research-units-pipeline-skills on GitHub](https://github.com/WILLOSCAR/research-units-pipeline-skills)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/willoscar-research-units-pipeline-skills
