Harness-1: the state lives in the harness, and the policy only decides
🚀 Ultra Recipe for Training Long-Horizon Search Agents - matching frontier AI's search capability with a 20B model + stateful harness
At a glance
- What is it?
- Harness-1 is a 20B search agent whose contribution is a recoverable retrieval harness: candidate documents, curated evidence, evidence links, verification records and budget-aware context are held outside the policy, which only makes semantic decisions. The repository also has a dependency comment that says it is pinning a version while its specifier allows any newer one.
- Who is it for?
- This repository suits a researcher who wants the harness design rather than the checkpoint, since the state, the tool environment, the training scripts and the evaluation runners are all here and the state machine is legible. It does not suit someone expecting a self-contained reproduction, because the retrieval backend has to be built separately, the most reproducible path runs through another project's pipeline, and a full evaluation reaches an external model.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 110 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
State outside the policy, decisions inside it
The design claim is the first paragraph, and it is a division of labour rather than a feature list. The harness maintains recoverable search state, and the items it holds are named: candidate documents, curated evidence, evidence links, verification records, and budget-aware context. The policy keeps the semantic decisions, also named: what to search, which documents to inspect or curate, what claims to verify, and when the evidence is sufficient. So a long-horizon run is not a single growing transcript. The evidence a step produced, the claim it was supporting, and whether that claim has been checked are all durable objects the model does not have to remember. That is what makes the run resumable, and it is also what makes the cost of a step inspectable after the fact.
A comment says it is pinning, and the specifier does not
One line in the runtime dependency list carries an inline comment: a floor on the harmony-format library, followed by a note attributed to a named maintainer saying higher versions break on their machine so it is pinned for now. The specifier itself is a greater-than-or-equal constraint, which permits every newer release. So the stated intent and the declared constraint disagree, and the resolution depends on the lockfile rather than on the manifest. A second, much tighter pair sits in the same list, a vector database and an ONNX runtime, both held at an exact version with no comment. Neither pattern is wrong, but a reader deciding whether the floor is safe has to check the lock file to know what is actually installed.
The minimal smoke test installs the entire training stack
The quickstart calls its own path a minimal local smoke test, and the requirement list beside it is short: Linux with Python 3.11 or newer, a project runner installed, a CUDA-capable environment, vLLM with the relevant model support, and access to the checkpoint. The install is one command, and the development path below it runs two scripts directly rather than through a test runner:
uv sync --extra vllm
export HARNESS1_HF_MODEL=pat-jj/harness-1What the install actually pulls is much larger, because there is a single runtime dependency list and exactly one optional extra, named for the serving engine. That list includes a parameter-efficient fine-tuning library, a reinforcement learning training library, a hosted training client, a serverless compute client, an experiment tracker, a Wikipedia client, a terminal formatting library and a second ONNX-based embedding library. Serving a checkpoint locally should not need the training and deployment stack, and it gets it anyway.
The corpora are released, the index is not
The dataset repository is published and it is substantial: one training split carrying a stage column, plus chunk text and cleaned metadata for four corpora. The README is explicit about what that does and does not buy you. The files provide the released chunk text and metadata, but the code still expects a compatible retrieval backend for full search evaluation, and a full BrowseComp+ run additionally needs a vector collection built from those chunks whose document identifiers match the query relevance judgments. So the released data is the expensive part and the index is the part you supply. For the most reproducible path, or if you want to rebuild or customise the indexes, the recommendation is to regenerate the corpora and collections with a separate data-generation pipeline maintained by a different organisation.
Fifteen example variables, five documented scopes
The environment template lists fifteen variables. The credentials section documents five scopes, and they cover fewer than half of the names. Documented are a hub token, used only if checkpoint access requires authentication, an OpenAI key for retrieval and evaluation workflows, a pair for the vector database, a pair for the optional reranker, and a hosted-training key. Undocumented in that section are an Anthropic key, a second model provider key, two other retrieval-service keys, and four BrowseComp+ path variables for the gold relevance file, the evidence file, the query file and the answer file. The template is the more complete document. The names in it are a better inventory of what a full evaluation actually reaches for than the prose list is.
A local vLLM run still wants a retrieval credential
This is the detail most likely to surprise someone who picked the project for locality. The minimum serving setup is genuinely local: a project sync, an environment variable pointing at the checkpoint, and a documentation file to follow. But the full evaluation path lists OpenAI credentials as a requirement, described as retrieval support used by the harness, and the manifest makes an Anthropic client a hard runtime dependency as well. So a locally served twenty-billion-parameter model doing BrowseComp+ evaluation still calls out to a proprietary service, and the local cache the README promises applies to the checkpoint weights rather than to the retrieval step. The optional reranker adds a third hosted dependency if you switch it on, and its client is installed either way.
Training queries come from filings, evaluation comes from browsing
The two halves of the training data come from different places and the two halves of the evaluation do not match either. The supervised half is 899 trajectories, raw, generated by a larger model and produced by a script whose name carries a date. The reinforcement half is 3,453 query records drawn from the training split of a financial regulator's filing dataset, selected by two environment variables naming the dataset and the split. The evaluation, meanwhile, is BrowseComp+, a web browsing benchmark, and the corpora published alongside cover browsing, the web, patents and the regulator filings. Nothing in the visible text discusses the gap between the domain the policy was reinforced on and the domain it is scored on, and no score is quoted anywhere in the document.
A chat template override sits in the repository root
Among the top-level entries is a single file with a template extension and a name that gives away its purpose: an override for the harmony chat template. That is consistent with the rest of the inference surface. There are three ways to run the model, a Hugging Face load script, a vLLM path that uses raw completions endpoints with token identifiers rather than chat messages, and hosted inference on the training platform, and all three have to agree on how the conversation is rendered. Overriding the template in-repo rather than in configuration means a reader who serves the checkpoint outside this repository has to notice the file and apply it by hand. The package version is 0.1.0 and describes itself as reproduction code, and there are no tagged releases.
Editorial conclusion
This repository suits a researcher who wants the harness design rather than the checkpoint, since the state, the tool environment, the training scripts and the evaluation runners are all here and the state machine is legible. It does not suit someone expecting a self-contained reproduction, because the retrieval backend has to be built separately, the most reproducible path runs through another project's pipeline, and a full evaluation reaches an external model. Check three things first. Check which retrieval index you will build, since the released chunks are not an index. Check which credentials your path needs, since a local vLLM run still wants a retrieval key. And check the domain of the training queries against the domain of the benchmark, because they are not the same one.
Frequently asked questions
What does Harness-1 actually contribute?
A stateful retrieval harness that holds recoverable search state outside the policy, namely candidate documents, curated evidence, evidence links, verification records and budget-aware context. The policy keeps the semantic decisions about what to search, what to verify and when evidence is sufficient.
What does it take to serve the Harness-1 checkpoint locally?
Linux with Python 3.11 or newer, a project runner, a CUDA-capable environment, vLLM with GPT-OSS support, and access to the checkpoint. You sync the vllm extra, export the model variable, and follow the vLLM and BrowseComp+ guide.
Do I need to build my own retrieval index for BrowseComp+ evaluation?
Yes. The corpus chunks and metadata are published, but the code expects a compatible retrieval backend, and a full run needs a Chroma collection built from those chunks with document IDs matching the qrels. The README recommends an external data-generation pipeline for rebuilding them.
What training data was Harness-1 trained on?
One dataset split with a stage column. The supervised stage holds 899 raw trajectories generated by a larger model, and the reinforcement stage holds 3,453 query records from the training split of the SEC filing dataset, selected with the dataset and split environment variables.
Which credentials does a full Harness-1 evaluation need?
The environment template lists fifteen variables. A full BrowseComp+ run needs a vector database key and database name, OpenAI credentials for the retrieval support the harness uses, the BrowseComp+ query, relevance, evidence and answer files on disk, and optionally a reranker key and model URL if reranking is enabled.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pat-jj-harness-1)