SWE-AF: Autonomous Engineering Fleet That Ships Code from a Single API Call
Autonomous software engineering fleet of AI agents for production-grade PRs on AgentField: plan, code, test, and ship.
At a glance
- What is it?
- SWE-AF is an open-source autonomous software engineering runtime built on AgentField that spins up a coordinated team of planning, coding, review, and verification agents from one API call, targeting teams who want production-grade pull requests generated without human scaffolding.
- Who is it for?
- SWE-AF is the right tool for teams running on AgentField who want to automate non-trivial software tasks, from single-repository refactors to cross-service changes that span multiple codebases. It is the wrong choice for anyone who needs a standalone coding assistant or who is not willing to operate an AgentField control plane.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SWE-AF is and who it targets
SWE-AF is not a single-agent code assistant. The README describes it as an autonomous engineering factory: a control stack of planning, execution, and governance agents that runs on the AgentField platform, handling the full lifecycle from goal specification to a shipped pull request.
The target user is a team that needs to automate production-grade software tasks, not a developer who wants autocomplete or a code review comment. The README points to PR 179 on the AgentField repository (Go SDK DID/VC Registration) as a concrete example: a real pull request built entirely by SWE-AF from one API call using haiku-class models.
The design explicitly separates itself from single-agent wrappers. The README states that the architecture encodes the engineering strategy, not the prompts. Planning, execution, and governance agents run as a coordinated control stack. Each agent type handles a specific concern: scoping the work, writing code in isolated git worktrees, running tests, reviewing diffs, and tracking compromises when scope is relaxed. The resolved complexity of an issue determines how many of these stages are engaged, which the README calls hardness-aware execution.
Triggering a build: the af CLI and the HTTP API
SWE-AF exposes two entry points. The af CLI, which requires version 0.1.87 or higher, streams live progress and prints the result:
af call swe-planner.build --in '{
"goal": "Refactor and harden auth + billing flows",
"repo_url": "https://github.com/user/my-project",
"config": {
"runtime": "claude_code",
"models": { "default": "sonnet", "coder": "opus", "qa": "opus" },
"enable_learning": true
}
}'The same request goes over raw HTTP:
curl -X POST http://localhost:8080/api/v1/execute/async/swe-planner.build \
-H "Content-Type: application/json" \
-d '{"input": {"goal": "Add JWT auth", "repo_url": "https://github.com/user/my-project"}}'The models object accepts per-role overrides: coder, qa, architect, and similar role keys. This means the team can run a cheaper model for planning and a more capable one for code generation. Supported runtimes include claude_code, open_code (OpenCode CLI v1.2+), and codex (OpenAI Codex CLI). The runtime key defaults to auto, which selects open_code when only an OpenRouter or Infron key is present, and falls back to claude_code otherwise.
Single-repository and multi-repository modes
SWE-AF works on one codebase or several at once. The single-repository mode takes either a remote repo_url or a local repo_path. Multi-repository mode passes a config.repos array where each entry has a role:
curl -X POST http://localhost:8080/api/v1/execute/async/swe-planner.build \
-H "Content-Type: application/json" \
-d '{
"input": {
"goal": "Add JWT auth across API and shared-lib",
"config": {
"repos": [
{"repo_url": "https://github.com/org/main-app", "role": "primary"},
{"repo_url": "https://github.com/org/shared-lib", "role": "dependency"}
],
"runtime": "claude_code",
"models": {"default": "sonnet"}
}
}
}'A primary repository drives the build; failures there block progress. A dependency repository receives supporting changes; failures are captured but do not stop the primary from proceeding. Typical use cases include an API plus a shared SDK, monorepo sub-projects, or a feature spanning multiple microservices.
The isolation mechanism is git worktrees, one per repository, so concurrent agent writes do not collide on the same branches.
Continual learning, compromise tracking, and resume_build
Three architectural features distinguish SWE-AF from simpler agentic loops: continual learning, explicit compromise tracking, and checkpointed execution.
When enable_learning is set to true in the build config, conventions and failure patterns discovered early in a multi-issue build are injected into the context of downstream issues. If the first ten issues reveal that the target repository requires a specific test fixture setup, that knowledge is carried forward without re-discovery.
When scope is relaxed during a build, the compromise is typed, severity-rated, and propagated. The README calls this explicit compromise tracking. This is a governance feature: it ensures that shortcuts taken under time or context pressure are recorded rather than silently merged into the output.
Checkpointed execution means that a long build can be resumed after a crash or interruption using resume_build. For builds that span hundreds of agent invocations, this is a reliability requirement rather than a convenience.
The Rust Python compiler benchmark
The README includes one performance benchmark built autonomously by SWE-AF: a Rust-based Python snippet executor, with results documented in examples/llm-rust-python-compiler-sonnet/README.md.
The comparison measures CPython subprocess spawning (approximately 19 milliseconds per call, approximately 52 operations per second) against a RustPython pre-warmed interpreter pool (in-process). The README records a geometric mean of 253.8 times faster, with a peak throughput of 31,807 operations per second versus roughly 52 for CPython subprocess. The README explicitly explains that these numbers measure different execution models: repeated subprocess spawning versus a persistent in-process pool. The design goal was to replace repeated subprocess invocations with a persistent pool for short-snippet execution, and the benchmark documents the real-world throughput difference for that specific workload.
The artifact trail for that build includes 175 tracked autonomous agents across planning, coding, review, merge, and verification. This is the kind of multi-agent invocation count that motivates the checkpointing and continual-learning features. Without checkpointed execution, a crash at agent 160 of 175 would mean starting over. Without continual learning, failure patterns found in early agents would not carry forward to later ones.
The examples directory also includes an agent-comparison folder and diagrams, suggesting that the benchmark is part of a broader effort to document what SWE-AF produces, but only the Rust Python compiler example has a full write-up in the repository.
Infrastructure requirements and known limitations
Running SWE-AF requires an AgentField control plane. The Docker Compose file starts it as the control-plane service on port 8080, with an ephemeral PostgreSQL database on a separate build-db service used for build-time integration checks. The swe-agent service connects back to the control plane at http://control-plane:8080 via the AGENTFIELD_SERVER environment variable. The agent listens on port 8003 and sets AGENT_CALLBACK_URL to http://swe-agent:8003 so the control plane can push results back.
The pyproject.toml pins claude-agent-sdk==0.1.20 with a comment noting that newer SDK builds have surfaced streaming errors. It also pins agentfield>=0.1.113 because versions below 0.1.96 did not correctly report empty build failures, a gap the CHANGELOG identifies as issue 82 Gap 2.
The Dockerfile installs several system-level tools: git for worktree management, the GitHub CLI (gh) for draft pull request creation, jq for JSON processing in agent bash scripts, Node.js and npm for the Codex CLI, and the OpenCode CLI via a curl-piped installer. Teams with locked-down Docker build environments may find the curl-piped installs inconsistent across runs if the upstream URLs change.
SWE-AF does not support offline operation. Every build requires the AgentField control plane and outbound access to the AI runtime provider. Teams with strict network isolation cannot use it as described in the README.
The benchmark and the PR example both use Claude models. The README documents multi-provider support including OpenRouter, OpenAI, and Google, but no benchmarks for those providers appear in the repository. The docker-compose.yml sets the SWE_DEFAULT_RUNTIME and SWE_DEFAULT_MODEL environment variables to empty by default, with a comment explaining that the empty values trigger automatic runtime selection: open_code when only an OpenRouter or Infron key is present, and claude_code otherwise. This means the default runtime depends on which API keys are in the environment at startup, not on a hardcoded value in the Compose file.
Editorial conclusion
SWE-AF is the right tool for teams running on AgentField who want to automate non-trivial software tasks, from single-repository refactors to cross-service changes that span multiple codebases. It is the wrong choice for anyone who needs a standalone coding assistant or who is not willing to operate an AgentField control plane. Before deploying it in production, verify the af CLI is at version 0.1.87 or higher, confirm which AI runtime (claude_code, open_code, or codex) your environment supports, and read the CHANGELOG.md to understand that the agentfield package pin is at 0.1.113 due to a specific gap in how empty builds are reported.
Frequently asked questions
What version of the af CLI does SWE-AF require?
The README specifies that the af CLI must be version 0.1.87 or higher to use the af call command. The agentfield Python package must be 0.1.113 or higher.
Can SWE-AF resume a build after a crash?
The README documents checkpointed execution and a resume_build parameter for recovering long-running builds after crashes or interruptions.
Does SWE-AF work with AI providers other than Anthropic?
The README lists multi-provider support including Claude, OpenRouter, OpenAI, and Google. Model keys are assigned per role in the config object. The runtime key selects the execution environment: claude_code, open_code, or codex.
What is hardness-aware execution in SWE-AF?
The README describes hardness-aware execution as a mechanism where easy issues pass through quickly and hard issues trigger deeper adaptation and DAG-level replanning instead of blind retries.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/agent-field-swe-af)