Model or dataset
Avarok-Cybersecurity/atlas avatar
Avarok-Cybersecurity/atlas

Atlas, a 75 MB Rust and CUDA inference engine, and the four places its own files disagree

Pure Rust Inference Engine

700 stars105 forksRustAGPL-3.0

At a glance

What is it?
Atlas is an AGPL licensed Rust and CUDA inference engine that serves OpenAI, Anthropic, and Responses APIs from a single 75 MB binary, verified on the NVIDIA DGX Spark. Its own files disagree about domains, model names, state volumes, and versioning, and those disagreements are the useful part.
Who is it for?
Atlas deserves an evaluation if your target is one DGX Spark serving a 35 B mixture of experts model to coding agents, and you intend to read the flag rationale before copying a command. Three things have to be checked on your own hardware before it becomes the serving path: the conditions behind the C=1 to C=128 ladder, whether the AMD target only compiles for you or also serves, and which of the two domains your deployment should point at.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The install is one curl pipe, and the model catalogue lives in a second repository

The engine starts from a two line shell session. The install script fetches a prebuilt `atlasctl`, verifies its checksum, and installs it to `~/.local/bin`, so a machine with no Python and no Rust toolchain can end up serving a model. The first run takes a recipe slug rather than a model id:

bash
curl -fsSL https://atlascybernetics.ai/install.sh | sh
atlasctl run qwen3.6-35b-a3b-fp8-mtp

Reading source instead of piping a download into a shell is a supported alternative: `cargo install atlasctl` produces the same tool from source.

What the script does not carry is the model list. Every runnable model has to map to a recipe in the separate `atlas-recipes` repository, and the stated reason is that the catalogue cannot advertise a model the maintainers do not ship. That is a tighter guarantee than most model tables offer, and it also splits the documentation in two. The engine lives here, the recipes live next door, and a reader who follows only this page cannot enumerate what is available.

Recipe 0 mounts ~/.avarok for state, the daily driver recipe does not

The two Docker recipes differ in a way that matters once you run the server for real instead of for a demo.

Recipe 0 has no model id at all. `serve` boots into a Library where you pick a model and recipe interactively, and it needs a real terminal to render, so it stays in the foreground. It mounts two volumes: the host Hugging Face cache so weights are not downloaded again, and `~/.avarok` so recipes, benchmark records, and artifacts survive the container.

Recipe A is the daily driver, and it runs detached under `--name atlas`:

bash
sudo docker run -d --name atlas \
  --network host --gpus all --ipc=host \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  avarok/atlas-gb10:latest \
  serve Qwen/Qwen3.6-35B-A3B-FP8 \
    --port 8888 \
    --max-seq-len 65536 \
    --kv-cache-dtype fp8 \
    --kv-high-precision-layers auto \
    --gpu-memory-utilization 0.90 \
    --scheduling-policy slai \
    --enable-prefix-caching \
    --speculative \
    --num-drafts 2 \
    --tool-call-parser qwen3_coder

It mounts the Hugging Face cache and nothing else. There is no `~/.avarok` bind mount here, so the state Recipe 0 goes out of its way to persist ends up inside the container filesystem and disappears when the container is replaced. The other flags are not decoration: `--network host` makes the served port reachable on localhost with no `-p` mapping, `--gpus all` hands the GB10 to the container, and `--ipc=host` is there because the Docker default of 64 MB for `/dev/shm` is too small for CUDA.

`--kv-high-precision-layers auto` is the literal number 2

`auto` is not a policy engine. It is a fixed alias for the number `2`, defined in `serve_phases/kv_cache.rs`, sitting alongside `max` and `all`, which mean every attention layer. The choice carries a second effect that is easy to miss: because the value is non zero, it also suppresses the per dtype automatic promotion that `0` would trigger under a `turbo*` KV dtype. A flag that reads like an instruction to the engine turns out to be a number, and that number also silences another feature.

The neighbouring flags are specific. `--kv-cache-dtype fp8` halves KV memory against BF16 with no measurable quality loss in the maintainers' account, and the boundary attention blocks stay in BF16 because that is where the routing distribution is most sensitive. `--scheduling-policy slai` is not the default: `serve` falls back to `fifo`, so SLO aware ordering only happens when you ask for it. SLAi reorders concurrent sequences to keep MTP verify batches dense and prefills shortest prompt first, which targets the multi turn tool loops these recipes are built for. `--enable-prefix-caching` turns on the radix tree prefix cache, and `--gpu-memory-utilization 0.90` caps how much of the GB10 the engine will take.

Two names for one model, and two places to type them

The same weights appear under two identifiers, and picking the wrong one costs a download.

The `atlasctl` path uses a recipe slug, `qwen3.6-35b-a3b-fp8-mtp`, which encodes the architecture, the FP8 quantization, and the MTP draft head in one token. The Docker path uses the Hugging Face repository id, `Qwen/Qwen3.6-35B-A3B-FP8`. Underneath, the model is described as 35 B parameters with 3 B active, GDN plus attention plus a 256 expert MoE, and an MRoPE positioned vision tower that goes unused here because this configuration is text only. The tool call side is handled by `--speculative` with `--num-drafts 2` and `--tool-call-parser qwen3_coder`, so tool calls return in a parseable shape rather than as prose an agent has to scrape.

The repository also carries `examples/curl` and `examples/python` directories, which sits oddly beside a headline about not needing Python, and the overview does not say whether those directories are client side only. That is worth checking before you decide the project does or does not need a Python host. Speculative decoding here has two forms, MTP draft heads and DFlash block diffusion, and the recipes pick one of them per model rather than letting the engine decide.

The homepage field and the install script point at different domains

Three names and two domains appear across this project, and the one in the repository's own settings is not the one the project uses.

The homepage field on the repository points at `https://atlasinference.io`. Every other surface disagrees with it. The banner links to `atlascybernetics.ai`, the workspace manifest sets `homepage` to the same address, `documentation` points at `docs.atlascybernetics.ai`, the engineering write ups live under `blog.atlascybernetics.ai`, and the install script itself is fetched from `https://atlascybernetics.ai/install.sh`. The Docker Hub organisation is `avarok`, the GitHub owner is `Avarok-Cybersecurity`, and the product name everywhere else is Atlas.

The one line summary is a fourth variant. It calls the project a Pure Rust Inference Engine, while the opening paragraph calls it Rust and CUDA. None of this blocks a deployment, but anyone who bookmarks the homepage field and then wonders why the docs subdomain does not resolve has a wasted afternoon. It is also the first thing to settle before pointing an installer script or a container tag at production.

The vendored cudarc patch, and a vendor directory the root listing does not show

The CUDA dependency is patched rather than taken straight from crates.io, and the patch is explained in more detail than most projects manage.

`Cargo.toml` carries `[patch.crates-io]` pointing cudarc at `vendor/cudarc`. The accompanying note says the vendored copy of cudarc 0.19.2 is byte identical to the crates.io release except in one respect: its driver `Lib` loader maps CUDA driver symbols that do not exist on SCALE and gfx1151, `cuArrayGetMemoryRequirements` named as the example, onto a `__avarok_missing_cuda_sym` stub instead of panicking at init. On NVIDIA every symbol resolves normally, so the patch does nothing there. That is a reasonable trade for shipping one CUDA source to two vendors whose driver APIs do not line up.

The next comment goes further. Atlas only touches the `driver` and `nvrtc` backends, `default-features=false` drops cudarc's cublas, cublaslt, curand, and runtime backends, and the generated FFI bindings for those never compiled backends have been removed from the vendored copy entirely. `fallback-dynamic-loading` is called out as required because cudarc's build script expects it. What is not visible: the top level entries list `3rdparty_patches` and no `vendor` directory, so anyone building from source should confirm where that path gets populated before planning a fork.

The C=1 to C=128 ladder points at conditions this page does not print

The performance claim refers to a section that is not spelled out here, and the difference between a build claim and a serving claim is worth separating.

Atlas is described as out serving vLLM at every rung from C=1 to C=128 on a published GB10 concurrency ladder, with the conditions linked to a Performance section further down the page. The part of the overview that is written out stops inside Quick Start, so the ladder itself, the harness, the batch sizes, and which vLLM build acted as the baseline are not stated. What is stated: the NVIDIA DGX Spark with the GB10, Blackwell SM121, is verified today; AMD Strix Halo on gfx1151 compiles the same CUDA source through SCALE, which is a build claim rather than a serving claim; and the MLPerf Inference v6.1 submission sits in the closed edge division with results under embargo until MLCommons publishes.

There are checkable receipts next to it. The fused Qwen Gated DeltaNet kernel is merged into Hugging Face Transformers, and every release image has to pass a serve gate covering boot, coherence, tool calls, and throughput within tolerance of a committed baseline, where a release that ships slower than its baseline fails the gate. A process claim like that one is the kind you can hold a vendor to, which makes the missing benchmark conditions more noticeable rather than less.

Build numbers for releases, a beta preview string for the version

The versioning scheme has two halves that do not line up, and the workspace names are half hardware specific.

The three most recent releases are build numbers: `b468`, `b469`, and `b470`, published on 2026-09-23 and 2026-09-24. The workspace manifest declares `version = "1.0.0-beta-preview"` with edition 2024 and `rust-version = "1.85"`, and the licence is `AGPL-3.0-only`. The table of contents names both a license section and an enterprise edition, so an AGPL codebase and a commercial path sit side by side in the project's own plan, and the terms of the second are not spelled out where this page can show them.

Of the sixteen workspace members, six carry `spark-` names: runtime, comm, model, nllb, server, and storage. The same workspace also has `build-amd.sh` and `serve-amd.sh` at the root, so those names read as residue from the GB10 first target rather than as a scope limit. The root holds working files too, `scratch_llama_perf.txt`, `scratch_llama_perf2.txt`, `HANDOFF.md`, `MMQ_PORT_HANDOFF.md`, `PROGRESS_LOG.md`, and `AI_REPOSITORY.md`, next to attribution notes for four separate kernel work items. The tree was pushed on 2026-09-28, with 700 stars, 105 forks, and 196 open issues.

Editorial conclusion

Atlas deserves an evaluation if your target is one DGX Spark serving a 35 B mixture of experts model to coding agents, and you intend to read the flag rationale before copying a command. Three things have to be checked on your own hardware before it becomes the serving path: the conditions behind the C=1 to C=128 ladder, whether the AMD target only compiles for you or also serves, and which of the two domains your deployment should point at. Skip it if you need PyTorch in the process, a model outside the recipe catalogue, or a stable semantic version, because the release train here is build numbers attached to a beta preview string.

Frequently asked questions

Does the Atlas inference engine need Python on the serving machine?

No. The install script downloads a prebuilt atlasctl, verifies its checksum, and installs it to ~/.local/bin, and the engine is described as Rust and CUDA with no Python, no PyTorch, and no runtime compilation. A `cargo install atlasctl` path produces the same tool from source.

Which hardware has Atlas been verified on?

The NVIDIA DGX Spark with the GB10, Blackwell SM121. AMD Strix Halo on gfx1151 is described as compiling the same CUDA source through SCALE, which is a build claim rather than a serving claim.

What does --kv-high-precision-layers auto do in Atlas?

It is a fixed alias for the number 2 in serve_phases/kv_cache.rs, alongside max and all for every attention layer. Because the value is non zero it also suppresses the per dtype automatic promotion that 0 would trigger under a turbo* KV dtype.

How are Atlas releases numbered?

By build number rather than semantic version. The three most recent tags are b468, b469, and b470, published on 2026-09-23 and 2026-09-24, while the workspace manifest declares version 1.0.0-beta-preview.

Can Atlas serve Qwen3.6-35B-A3B from Docker?

Yes, with the avarok/atlas-gb10 image and serve Qwen/Qwen3.6-35B-A3B-FP8, using --max-seq-len 65536, --kv-cache-dtype fp8, and --ipc=host. serve defaults to the fifo scheduling policy, so --scheduling-policy slai has to be passed explicitly.

What licence does the Atlas project use?

The workspace manifest declares AGPL-3.0-only. The project's own table of contents names both a license section and an enterprise edition, so commercial terms sit alongside the AGPL grant and are not described in the overview.

Official sources

  1. Avarok-Cybersecurity/atlas on GitHub
  2. License: AGPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/avarok-cybersecurity-atlas.svg)](https://hysenlabs.com/projects/avarok-cybersecurity-atlas)