Model or dataset
RyanAlberts/best-of-Agent-Harnesses avatar
RyanAlberts/best-of-Agent-Harnesses

Best of Agent Harnesses: Weekly-Ranked Curated List of AI Agent Runtimes

🏆 Curated, ranked list of AI agent harnesses (100+) — plus an MCP server, llms.txt & JSON so agents can recommend them too. Rescored weekly.

1,011 stars56 forksPythonCC-BY-SA-4.0

At a glance

What is it?
best-of-Agent-Harnesses is a curated and weekly-rescored list of over 100 AI agent harnesses, with comparison guides, setup templates, playbooks, and an MCP server that lets an agent query the list itself to pick the right harness. It is aimed at engineers and teams who need to evaluate and select agent runtimes for production use.
Who is it for?
best-of-Agent-Harnesses is the right starting point for any team that needs to evaluate agent harnesses systematically rather than by word of mouth. The comparison pages, decision guides, and benchmark data cited in the README give a structured basis for the choice.
Can I use it commercially?
Yes, with credit. CC-BY-SA-4.0 allows commercial use as long as you credit the authors and indicate what you changed. It is written for creative content, so check how it applies to any code.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Ranking Agent Harnesses, Not Agent Frameworks

best-of-Agent-Harnesses is organized around a specific definition. The README cites Simon Willison's formulation: an LLM agent runs tools in a loop to achieve a goal. The harness is everything around that loop: which tools exist, what needs approval, what the model sees each turn, and what survives a crash. Andrej Karpathy's 2023 framing is also cited in the README: the model is the kernel process of a new OS, and the harness is the rest of that OS: its scheduler, permissions, and memory.

The SWE-agent paper's concept of the agent-computer interface is the core rationale: how tools and feedback are presented changes what a model can do, independent of the model. The README cites analysis showing that on SWE-bench Pro, swapping the agent harness changed pass@1 more than many model upgrades: 23% to 52% for one model, 15% to 36% for another. Rank correlation between harnesses barely transfers across models at -0.05, which means a harness that performs well with one model may not rank the same with a different model.

This is the problem the list exists to solve: engineers need harness recommendations matched to their specific model and task, not just a global leaderboard.

How Projects Are Scored and Ranked

The README describes a ranking that combines relevance to harness concerns (environment, orchestration, lifecycle, guardrails) with GitHub stars and activity. The list distinguishes two dimensions for evaluation: a simplicity/capability axis labeled as the adoption surface area, and two additional attributes: headless readiness (marked with a star in the tables) and durability (marked with a cross). Headless-ready projects can run without a human in the loop; durable projects survive process restarts.

Rankings are regenerated weekly. The repository includes releases tagged with the week of the update, such as list-2026-09-2 (Mid-September 2026: the harness-matters-more catalog, QM, Prime Agent, OpenJarvis) and list-2026-09 (September 2026 update). The landscape.md file and associated charts are regenerated from the list data on every refresh, plotting projects by adoption surface area against GitHub stars and by autonomy against recovery.

The repository layout separates the machine-readable data (harnesses.json, harnesses.jsonld, llms.txt, projects.yaml), the comparison pages in comparisons/, the templates in templates/, and the playbooks in playbooks/. The attributes/ directory holds structured attribute data per project. The tags taxonomy is in TAGS.md.

Querying the List: MCP Server, llms.txt, and harnesses.json

The MCP server in the mcp/ directory exposes a recommend tool and a pick_harness tool. The README states that pointing an agent at this MCP server lets the agent call recommend or pick_harness to choose a harness matched to its model and task, instead of inheriting whichever harness was benchmarked elsewhere.

Three machine-readable formats are included. harnesses.json is a JSON listing of all projects in the list with their attributes. harnesses.jsonld provides the same data in JSON-LD format for semantic web consumers. llms.txt follows the llms.txt convention for LLM-readable site summaries. The server.json file appears to hold server configuration for the MCP endpoint.

The web interface at ryanalberts.github.io/best-of-Agent-Harnesses/ provides a searchable page per harness with filter capabilities by capability, autonomy, and recovery. This is the human-browsable view of the same data that the MCP server exposes programmatically.

Comparison Pages and Decision Guides

The repository includes ten comparison pages in the comparisons/ directory, each addressing a specific decision in the harness selection process. The README enumerates them: why the harness matters more than the model (citing the ARC-AGI-3 result where Prime Agent's harness took one model from 30% to 95.5%), how to pick a harness (six questions that turn the list into a decision), and how to test-drive a harness (a two-week protocol with seven measurements and a walk-away test).

Head-to-head comparisons include OpenClaw versus Hermes for always-on personal agents, managed versus self-hosted always-on agents (covering Grok Bot, Claude Managed Agents, QM, OpenClaw, Hermes, OpenJarvis), terminal coding agents (opencode, Codex, Gemini CLI, crush, goose), multi-agent orchestration (OpenAI Agents SDK, CrewAI, AutoGen, LangGraph), and agent memory layers (Mem0, Letta, claude-mem).

The comparison pages cite specific measurements: the ARC-AGI-3 benchmark showing 30% to 95.5% on the same model with a different harness, the SWE-bench Pro data showing 23% to 52% pass@1 variance, and the -0.05 rank correlation across models. The comparisons/why-the-harness-matters.md page documents the sources and what the claims do not mean.

Templates, Playbooks, and the CLAUDE.md Agent Integration

The templates/ directory provides copy-paste setup files for harnesses in the list. The playbooks/ directory provides step-by-step guides for specific scenarios. The README describes these as practical resources for engineers who have made a harness decision and need to start using it.

The repository includes a CLAUDE.md file at the top level, which is a project instructions file for AI coding assistants. This indicates the repository is designed to be used directly by agents, not just browsed by humans. The AGENTS.md file (referenced in the Ellama README's section on agent loops, and common in recent repositories) would be for similar purposes.

A CITATION.cff file is included for academic citation of the repository. The contributors.json file lists contributors to the curation effort. Traffic history is tracked in traffic-history.jsonl. These files suggest a maintained project with more than a single author.

Coverage Gaps and Known Limitations

The README itself acknowledges a key finding embedded in the benchmark data: harness rankings barely transfer across models, with a rank correlation of -0.05. This means the list cannot provide a single definitive answer. The recommended response is to use the list as a starting point, then run the two-week trial from comparisons/how-to-test-drive-a-harness.md with the specific model and task the user is evaluating.

The list covers harnesses and orchestration frameworks, not the underlying models. Model quality, cost, and API availability are outside its scope. Teams that have already chosen a model and just need a harness comparison benefit most; teams still evaluating models should make the model decision first.

The list is strongest on general-purpose and coding agents. Domain-specific harnesses for fields like robotic control, scientific computing, or finance are less likely to appear. The curation queue is tracked in curation-queue.json, so it is possible to check what candidates are pending.

License, Maintenance, and a Note on Automated Alternatives

best-of-Agent-Harnesses is licensed under CC-BY-SA-4.0, which permits sharing and adapting the content as long as the same license is applied to derivative works and attribution is given. This covers the curated list content, the comparison pages, and the templates. Software in the mcp/ and scripts/ directories may carry different terms.

The last push was on September 23, 2026, and the repository is updated weekly. The two most recent releases are list-2026-09-2 from September 14, 2026, and list-2026-09 from September 9, 2026.

For comparison, the best-of-ml-python project (part of the best-of ecosystem) generates automated rankings from GitHub metrics on a weekly basis across a wide range of ML libraries. That approach covers more projects and requires no manual curation but applies the same star-based scoring to all categories, which does not distinguish harness-specific attributes like headless readiness, durability, or recovery behavior. best-of-Agent-Harnesses adds that harness-specific layer on top of the star data.

Editorial conclusion

best-of-Agent-Harnesses is the right starting point for any team that needs to evaluate agent harnesses systematically rather than by word of mouth. The comparison pages, decision guides, and benchmark data cited in the README give a structured basis for the choice. Because rankings are regenerated weekly and the underlying benchmark data show that rank correlation between harnesses barely transfers across models, treat the current week's ranking as a data point and run a two-week trial from comparisons/how-to-test-drive-a-harness.md before committing to a harness for production.

Frequently asked questions

What is the difference between an agent and a harness?

The README defines this directly: an agent runs tools in a loop to achieve a goal (the model thinks and acts), while the harness is everything around that loop: which tools exist, what needs human approval, what the model sees each turn, and what state survives a crash. The model provides reasoning; the harness provides the execution environment, permissions, and lifecycle management.

Which is the best agent harness?

The README states that the best harness depends on the specific model and task: rank correlation between harnesses barely transfers across models (measured at -0.05), so a harness that ranks first with one model may rank much lower with another. The list provides comparison pages and a two-week trial protocol in comparisons/how-to-test-drive-a-harness.md for evaluating harnesses against your specific setup.

What are examples of agent harnesses listed in this repository?

The README names specific harnesses in its comparison pages, including OpenClaw, Hermes, Grok Bot, Claude Managed Agents, QM, OpenJarvis, opencode, Codex, Gemini CLI, goose, and Prime Agent. The full list of 100+ projects is in harnesses.json and searchable at ryanalberts.github.io/best-of-Agent-Harnesses/.

Official sources

  1. License: CC-BY-SA-4.0
  2. Project website
  3. README
  4. Releases
  5. RyanAlberts/best-of-Agent-Harnesses on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ryanalberts-best-of-agent-harnesses.svg)](https://hysenlabs.com/projects/ryanalberts-best-of-agent-harnesses)