# SpatialClaw: NVIDIA's Code-As-Action-Interface Agent for Spatial Reasoning

> SpatialClaw is a training-free spatial reasoning framework from NVIDIA and KAIST that lets a VLM write one Python cell per step into a persistent Jupyter kernel. It reports 59.9% average accuracy across 20 benchmarks, and the repository is a research runtime, not a packaged library.

**NVlabs/SpatialClaw** — SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

- Repository: https://github.com/NVlabs/SpatialClaw
- Website: https://spatialclaw.github.io/
- Stars: 392 · Forks: 40
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvlabs-spatialclaw

## What SpatialClaw solves, and who the paper is arguing with

Spatial reasoning means determining where objects are, how they relate, and how they move in three dimensions. Tool-augmented agents try to solve this by giving a vision-language model specialist perception modules, but the paper's abstract argues that their effectiveness is bounded by the action interface through which those tools are invoked. That is the target. Existing spatial agents either commit to a full analysis strategy in a single code execution before any intermediate result is visible, or they use a structured tool-call interface that the authors describe as offering less flexibility for freely composing operations. SpatialClaw's answer is to make code itself the interface. The intended reader is a researcher working on agentic spatial reasoning or multimodal reasoning evaluation, not an application developer looking for an off-the-shelf component.

## The five-stage loop and the three services behind it

For every sample, the README describes a five-stage loop running on top of a persistent Python kernel. A planner drafts a strategy. The main VLM writes one Python cell per step. That cell is AST-checked and executed in the stateful kernel. Then stdout, new variables, and images produced by `show()` flow back as the next observation. The loop repeats until the agent commits with `ReturnAnswer(...)`. The kernel is pre-loaded with perception primitives (SAM3 segmentation, Depth-Anything-3 reconstruction, geometry utilities) and scientific libraries (NumPy, SciPy, Matplotlib), so a cell can compose tool outputs, inspect intermediate evidence, and revise before answering.

The runtime is split into three independent services: a vLLM backbone, a GPU perception-tool server (Reconstruct and SAM3), and the agent running Jupyter kernels. They coordinate through shared JSON registries. On a cluster, auto-restarting chain jobs keep them alive past SLURM job-time limits. The repository also states that each service is a plain Python entry point that can run on any GPU machine, which matters if you do not have SLURM. That three-process split is the main operational cost of the design: the flexibility of a live kernel is paid for in coordination.

## Installing SpatialClaw and running one experiment

The README's quickstart clones with submodules, runs a setup script, copies the environment file, and then invokes the runner. The setup script is documented as installing the agent and vLLM environments and taking roughly 15 to 30 minutes.

```bash
git clone --recursive https://github.com/NVlabs/SpatialClaw.git
cd SpatialClaw
bash spatial_agent/scripts/setup.sh
```

The `--recursive` flag is not optional here: the repository has a `.gitmodules` file and the documentation table references third-party submodules, so a plain clone leaves the perception tooling missing.

Next, create the environment file. The shipped model configs use self-hosted vLLM with `"llm_api_key": "bearer"`, so no hosted key is needed for the default path.

```bash
cp .env.example .env
```

`.env.example` states that values there are loaded by `spatial_agent.config.get_config()` and can be referenced from model config JSON via `${VAR_NAME}` expansion. It also notes that variables already set in the shell take precedence over `.env`. `HF_TOKEN` is commented out by default and is described as required for gated repos such as SAM3.1 during model download.

Then run a dataset on a single machine:

```bash
python -m spatial_agent.entrypoints.run \
    --dataset spatial_agent/config/dataset/erqa.json \
    --model   spatial_agent/config/model/gemini-3-pro.json \
    --concurrency 4
```

What you should see is the agent loop executing cells against the kernel and eventually emitting a `ReturnAnswer(...)` for each sample. Two cautions from the README: pre-downloading model weights is described as mandatory before SLURM runs, and the vLLM and SLURM setup has extra steps covered in `docs/installation.md` and `docs/running.md`. The README does not document rollback or an uninstall path.

## Where SpatialClaw is the wrong tool

The repository is the official implementation of a paper, and it reads that way. There are no retrieved releases, so there is no versioned artifact to pin. Integration is by cloning the repository and running its entrypoints, not by installing a package. If your project needs a stable API surface, this is a mismatch.

The dependency weight is the second constraint. You need GPU capacity for the vLLM backbone and for the perception tool server, plus the model weights, which the README says must be pre-downloaded. A gated HuggingFace repository (SAM3.1) sits in that download path, so an account without access stops the setup. The two-environment conda layout adds a third constraint: the setup script is a 15 to 30 minute operation, and rebuilding it is the cost of moving machines.

Finally, the evaluation harness assumes the paper's benchmark loaders and config JSON files. If your task is not one of the 20 benchmarks and you cannot express it as a dataset config plus a model config, you are writing against the runtime rather than using it. The README does not describe a plugin contract for new tools or new benchmark formats.

## How it differs from structured tool-call agents

The natural comparison is a structured tool-call agent, where the model emits a JSON object naming a tool and its arguments, the harness validates it, runs the tool, and returns a result. That design constrains each step to the tools and argument schemas the developer registered. SpatialClaw's alternative is a persistent Python kernel: the model writes arbitrary Python, and the observation is whatever the cell printed or bound to a name. The difference shows up in intermediate work. A tool-call agent that wants to crop a region, measure a distance, and plot the result before deciding needs a registered tool for each operation. In SpatialClaw, the README's description of composing tool outputs and inspecting intermediate evidence with NumPy, SciPy, and Matplotlib is the same step.

The trade is safety and reproducibility for flexibility. SpatialClaw runs cells through an AST safety check, which is a filter rather than a sandbox, and arbitrary generated Python is harder to audit than a fixed tool vocabulary. The paper reports 59.9% average accuracy, +11.2 points over the prior best spatial agent, using the same system prompt, tool set, and hyperparameters across all benchmarks and six VLM backbones from two model families (Qwen3.5/3.6, Gemma4) between 26B and 397B parameters. Those numbers come from the paper and the README; the repository does not ship a script that reproduces the headline figure in one command.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-07-18. That is recent, but it is a research codebase with no retrieved releases, so "upgrade" means pulling `main` and re-running `bash spatial_agent/scripts/setup.sh` against a possibly changed environment specification. Budget for that reinstall whenever you sync. If you cloned with `--recursive`, remember to update submodules too, since the perception wrappers live in them.

The licence field is NOASSERTION, which means GitHub could not classify the LICENSE file automatically. Read the LICENSE file in the repository root yourself before you plan any redistribution or commercial use; the README does not summarise the terms, and nothing here should be read as legal advice. The same caution applies to the model weights and third-party submodules you download, which carry their own terms independent of this repository.

## Conclusion

SpatialClaw is aimed at researchers who want to reproduce or extend the paper's results and who have GPU capacity and are comfortable with conda environments, vLLM, and optionally SLURM. It is not the right choice if you need a pip-installable library, a stable public API, or a production service, because the repository ships as an experiment runtime with no releases and a NOASSERTION licence. Before committing, verify the licence terms in LICENSE, confirm that the mandatory model weight downloads (including the gated SAM3.1 repository) are reachable with your HuggingFace token, and run the single-machine entrypoint on one dataset to see whether the two-environment setup and GPU tool server fit your hardware.

## FAQ

### Can LLMs be used for spatial reasoning?

SpatialClaw's premise is that they can, but only when augmented. The README describes a VLM-backed agent that writes Python cells into a kernel pre-loaded with perception primitives, and reports 59.9% average accuracy across 20 spatial reasoning benchmarks with consistent gains across six VLM backbones.

### What does SpatialClaw mean by code as the action interface?

Instead of emitting structured tool calls, the agent writes one executable Python cell per step into a persistent Jupyter kernel. Each cell can compose tool outputs and inspect intermediate evidence before the agent commits an answer with ReturnAnswer(...).

### Is SpatialClaw training-free, or does it need fine-tuning?

The README describes it as training-free. The same system prompt, tool set, and hyperparameters are used across all 20 benchmarks and six VLM backbones without benchmark- or model-specific adaptation.

### Do I need a SLURM cluster to run SpatialClaw?

No. The README states that each of the three services is also a plain Python entry point you can run on any GPU machine, and the quickstart shows a single-machine run of python -m spatial_agent.entrypoints.run. SLURM chain jobs are only described as the way the paper's experiments survive job-time limits.

### What is SpatialClaw's licence?

The repository's licence field is NOASSERTION, meaning GitHub could not classify the LICENSE file. The README does not summarise the terms, so read the LICENSE file in the repository root before relying on any particular permission.

## Sources

- [Issues](https://github.com/NVlabs/SpatialClaw/issues)
- [NVlabs/SpatialClaw on GitHub](https://github.com/NVlabs/SpatialClaw)
- [Project website](https://spatialclaw.github.io/)
- [README](https://github.com/NVlabs/SpatialClaw/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvlabs-spatialclaw
