ATLAS: a local coding agent that wraps small open models in planning, scoring and sandbox repair
Adaptive Test-time Learning and Autonomous Specialization
At a glance
- What is it?
- ATLAS is a self-hosted coding agent from itigges22 that puts candidate generation, a Geometric Lens scorer and sandboxed verification around a GGUF model you run yourself. The design buys reliability with tokens and hardware, and the install path forks between Docker Compose and a K3s script.
- Who is it for?
- Adopt ATLAS if you already run GGUF weights locally and want verification around a compact model rather than a hosted API, and if you can give it a GPU with enough VRAM for the context you size. Do not adopt it if you need a supported Windows install, a pinned dependency set, or a tool that never executes generated code; the sandbox runs commands and has outbound network access unless ATLAS_SANDBOX_NET_INTERNAL=true.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem ATLAS targets: small models that fail on real repository work
A 14B model asked to edit a file in a real project tends to produce something plausible and wrong. ATLAS is built on the premise that the fix is not a bigger model but more machinery around the one you have. The README states the goal directly: it "puts more intelligence in the system around the model (planning, candidate generation, quality scoring, sandboxed testing, and repair) so smaller models can tackle real software work entirely on your own hardware, without a hosted API or per-token fees."
The audience is narrow and identifiable. You need a machine with a supported accelerator, GGUF weights on disk, and a willingness to run a multi-service stack rather than a single binary. In exchange you get no per-token billing, no model-provider API key, and no repository content leaving the host by design. The README is careful about that last claim: ATLAS "does not intentionally upload your repository or prompts to a hosted model or ATLAS-operated service," and it immediately notes that sandbox commands have outbound network access by default, which you disable with ATLAS_SANDBOX_NET_INTERNAL=true. That caveat matters more than the headline claim.
Where ATLAS does not fit is equally clear. It is not a chat wrapper you point at a hosted endpoint. The V3 pipeline assumes it can compile and run what the model produces, so tasks whose output cannot be executed in a sandbox get less benefit from the verification stages that justify the setup cost.
How the V3 pipeline and the Geometric Lens actually fit together
The architecture splits into two layers. The outer layer is atlas-proxy, a Go agent loop that classifies file operations by complexity tier, enforces GBNF grammars so model output lands in the expected JSON shapes, and applies turn caps, token budgets and timeouts. The inner layer is the V3 pipeline, which turns one prompt into a verified candidate.
The pipeline phases are named in the README: PlanSearch for constraint-driven structured planning, DivSampling for diverse candidates across temperature and strategy, Budget Forcing for per-phase thinking-token allocation, PR-CoT Repair for self-generated test cases used in iterative fixes, and Refinement Loops that sandbox-verify and correct before repeating. The selection step is where the Geometric Lens sits. It is described as energy-based scoring over the model's own embeddings with no external oracle, combining a C(x) cost field (a model-hidden-dim to 512 to 128 to 1 MLP) and a G(x) quality prediction built as an XGBoost ensemble. Per-step scoring runs on port 8099.
The cost model is the interesting part. Straightforward edits take a shorter path; harder tasks get more candidates, more reasoning and more validation. That is a deliberate trade of latency and tokens for pass rate, and it means ATLAS behaves badly on a machine where you cannot afford several generations per edit. The published V3 figure is 74.6% LiveCodeBench pass@1-v(k=3) on frozen Qwen3-14B, which the README defines as pass@1 with k=3 generated candidates, Lens selection and repair, explicitly not single-generation pass@1. Read it as a system number, not a model number.
Installing ATLAS with Docker Compose and running a first edit
The repository ships both a Compose path and a K3s installer at ./scripts/install.sh. The two do not share configuration: .env.example is read by docker compose and by the atlas CLI for port and model resolution, while the K3s installer reads atlas.conf. The comment in .env.example warns that values placed in atlas.conf are silently ignored by the docker path. Pick one path and stay on it.
The preferred start is the CLI's own initializer, which copies the environment file and fills in the selected model plus hardware-aware sizing. Run it in the repository root:
atlas initIf you set things up by hand instead, copy .env.example to .env and fill the required fields. Two are non-negotiable, the model filename and the model identifier:
ATLAS_MODELS_DIR=./models
ATLAS_MODEL_FILE=
ATLAS_MODEL_NAME=
ATLAS_CTX_SIZE=131072ATLAS_CTX_SIZE is the total context across all parallel slots, and llama-server divides it by ATLAS_PARALLEL_SLOTS for per-slot context. Size it for your own model and GPU with the tier command, which writes the runtime-sizing keys back into the file:
atlas tier fit --writeThe Compose file notes that the server runs with --fit off, so an oversized budget refuses to start rather than silently spilling layers to CPU and running roughly five times slower. That refusal is a feature, but it means a wrong ATLAS_CTX_SIZE looks like a broken install. Bring the stack up and launch the terminal client from any project directory:
docker compose up -d
atlasThe TUI is the canonical chat client. Slash commands such as /add, /diff, /commit and /run handle local file context and shell-out, and the live pipeline view streams V3 stages in a side pane so you can watch candidates get scored and repaired. NVIDIA is the default Compose target; docker-compose.rocm.yml, docker-compose.macos.yml, docker-compose.vulkan.yml and docker-compose.cpu.yml cover the other backends named in the README.
What the sandbox and the local-only claim do not cover
The security posture deserves a closer read than the marketing line. ATLAS does not intentionally send your repository to a hosted model, but the sandbox that verifies generated code has outbound network access by default. Setting ATLAS_SANDBOX_NET_INTERNAL=true disables it. Until you set that, code the model wrote is executing on your machine with a network path open. Verification is the feature that makes ATLAS worth running, and it is also the sharpest edge in the design.
Resource behavior is the second limitation. Every hard task multiplies generations, Lens scoring passes and repair loops. On a CPU-only setup, which docker-compose.cpu.yml supports, the pipeline is technically runnable and practically painful. The README's own framing is that ATLAS spends compute where it matters, which is honest but also a warning about what a hard task costs.
Configuration is the third. The split between .env.example and atlas.conf, combined with the note that atlas.conf values are silently ignored by the docker path, is a trap for anyone who reads one file and assumes it governs both. The .env.example comment about the K3s installer is the only place that distinction is stated plainly.
Finally, packaging. pyproject.toml declares Development Status :: 4 - Beta and classifies the operating systems as POSIX Linux and macOS. The core CLI is deliberately stdlib-only with no base dependencies, and the train extra is described as lower bounds only, with the CLI having no reason to pin exact versions. That is a defensible choice for a fast-moving project and an uncomfortable one for anyone who needs reproducible installs.
ATLAS compared with Continue and Aider
The closest comparison is not another local model runner but another coding agent. Aider and Continue both drive an assistant against your repository, and both can be pointed at a local OpenAI-compatible endpoint. The difference is where the intelligence is assumed to live.
Those tools treat the model as the source of correctness and add editing ergonomics: diff formats, repo maps, commit handling. ATLAS treats the model as one component in a pipeline. It generates several candidates, scores them with the Geometric Lens, compiles and tests them in a sandbox, and repairs the failures. Grammar enforcement through GBNF schemas and the BiasBusters mitigations exist because the proxy cannot assume the model will emit the right tool call on its own. Four composed mitigations (descriptions, grammar bans, system notes and ASA steering) push the model toward structural_edit for structural code edits.
The trade is direct. Aider-style tools are lighter, work with any endpoint including hosted ones, and give you whatever the model can do in one pass. ATLAS is heavier, needs local weights and a supported accelerator, and buys its pass rate with multiple generations plus execution. If your model is already strong, the pipeline is overhead. If you are trying to get useful work out of a 14B model on your own GPU, the scaffolding is the product.
Release cadence, licensing and the upgrade cost of a three-month-old install
The last push to the default branch was on 2026-07-06, the same day v3.1.3 "Maia" was tagged, so the repository is not archived and was touched recently. The Maia line moved quickly: v3.1.0 on 2026-05-12, v3.1.2 on 2026-06-17, v3.1.3 on 2026-07-06. Three releases in under two months is a fast cadence, and the v3.1.3 notes describe a production-platform pass that replaced Redis with a SQLite state store, added staged upgrade and rollback with auto-restore, signed artifact manifests, structured logs with correlation IDs, interactive permissions and session resume.
That changelog entry is the upgrade-cost warning. An install predating v3.1.3 has a Redis dependency that no longer exists and lacks the rollback path. If you are running an older Maia build, the state store migration is not a drop-in.
On licensing, pyproject.toml declares AGPL-3.0-or-later and the repository carries an AGPL-3.0 LICENSE file. The practical point for an engineering reader is that AGPL is a strong copyleft licence with a network-use clause, so if you modify ATLAS and expose it to users over a network, the licence's obligations follow the modified work. Whether that reaches your specific deployment depends on facts about your use that this article cannot assess, and it is worth a lawyer's read before you build a product on top of it rather than just running it internally. THIRD_PARTY_NOTICES.md exists in the repository for the bundled components.
Who should run ATLAS, and what to check before you commit
ATLAS is for the engineer who already has GGUF weights and a GPU, wants a coding agent that never calls a hosted model, and is willing to run a Compose stack with an inference container, a proxy and a sandbox. The verification loop is the reason to choose it over a thinner local assistant, and the 74.6% figure on frozen Qwen3-14B is the claim to evaluate against your own tasks.
It is the wrong tool if you work primarily on Windows, since the classifiers list POSIX Linux and macOS only. It is wrong if you need pinned dependencies, because the train extra is explicitly lower-bounded and the project is marked Beta. And it is wrong if you want a tool that never executes model output, because sandbox verification is the core mechanism.
Three things to verify before you invest a weekend. Check SUPPORT_MATRIX.md for whether your accelerator is a first-class backend or a Compose variant. Confirm that your model file exists in ATLAS_MODELS_DIR and matches ATLAS_MODEL_FILE and ATLAS_MODEL_NAME, since those are the two fields the environment file marks required. And run atlas tier fit --write before your first real task, because the server refuses to start on an oversized context rather than degrading quietly.
Editorial conclusion
Adopt ATLAS if you already run GGUF weights locally and want verification around a compact model rather than a hosted API, and if you can give it a GPU with enough VRAM for the context you size. Do not adopt it if you need a supported Windows install, a pinned dependency set, or a tool that never executes generated code; the sandbox runs commands and has outbound network access unless ATLAS_SANDBOX_NET_INTERNAL=true. Before committing, check SUPPORT_MATRIX.md for your hardware backend and confirm the model file you intend to use resolves through ATLAS_MODEL_FILE and ATLAS_MODEL_NAME.
Frequently asked questions
What is ATLAS AI?
ATLAS stands for Adaptive Test-time Learning and Autonomous Specialization, and it is a local coding agent that runs a GGUF model on your own hardware. It wraps the model in planning, candidate generation, Geometric Lens scoring, sandboxed testing and repair rather than relying on a single generation.
How do I install ATLAS on Windows 11?
The repository does not document a Windows install. pyproject.toml classifies the operating systems as POSIX Linux and macOS, and the Compose files cover NVIDIA, ROCm, macOS, Vulkan and CPU backends, none of them Windows.
How do I use ATLAS after installing it?
Run atlas in any project directory to launch the terminal UI, which the README calls the canonical chat client. Slash commands such as /add, /diff, /commit and /run handle local file context and shell-out, and a side pane streams the V3 pipeline stages.
What is in ATLAS?
The repository contains an inference service, a Go proxy agent loop, the V3 pipeline, the Geometric Lens scoring service, a sandbox, a Bubbletea terminal UI and a v3-service, alongside Compose files for NVIDIA, ROCm, macOS, Vulkan and CPU.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/itigges22-atlas)