# Ten prompts, 141 features, and an assembly gallery priced per item

> MAC is a Python framework that turns a design brief into editable CAD geometry through a LangGraph pipeline of planning, coding, execution and repair, and it now also generates multi-part assemblies. What the repository documents is narrow: a single-part benchmark of ten prompts and 141 features, assembly runs that cost from about one dollar to about thirteen, a requirements file that tells you not to install it, and three different version labels for one project.

**Pan-Chera/Multi-Agent-CAD** — MAC (Multi-Agent CAD): A decoupled multi-agent framework for text-to-CAD generation via constrained test-time compute

- Repository: https://github.com/Pan-Chera/Multi-Agent-CAD
- Stars: 1,016 · Forks: 94
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/pan-chera-multi-agent-cad

## The benchmark with the numbers behind it is ten prompts long

Everything quantified in the repository is scoped by one sentence: the figures belong to a documented ten-prompt, 141-feature single-part workflow benchmark, and they are explicitly not assembly success-rate or assembly-cost claims. The table holds four rows. A reproduced CAD Skill baseline on Qwen 3.7 with no in-loop verification records 103.95M tokens at $18.59 and passes 138 of 141 features. The same baseline on Qwen 3.8 with visual verification records 87.46M tokens at $29.45 and passes 132 of 141. MAC v1 on Qwen 3.7 with geometric verification records 0.90M tokens at $1.43 and passes 140 of 141. MAC v2 on Qwen 3.8 with visual plus geometric verification records 1.40M tokens at $4.85 and also passes 140 of 141.

The methodology is pointed at by name rather than embedded: the single-part workflow README, an English evaluation file, and a Chinese evaluation file under `docs/`. The project describes itself in packaging metadata as Development Status 4 - Beta, and the repository records no homepage.

## The v2 row is three times v1 on cost and one feature different

Two multiplier claims follow the table. Under the Qwen 3.7 setup, MAC v1 used 116x fewer recorded tokens and a 13x lower estimated cost than the reproduced Skill baseline. Under Qwen 3.8 with visual verification, MAC v2 used 62.7x fewer recorded tokens and a 6.1x lower estimated cost.

The two rows cannot be compared cleanly against each other, because MAC v1 ran on Qwen 3.7 with geometric verification while MAC v2 ran on Qwen 3.8 with visual plus geometric verification. Two things changed at once, and the 1.40M tokens and $4.85 of v2 against the 0.90M and $1.43 of v1 land on the same 140 of 141 features. The baseline rows show the same non-monotonic shape: adding visual verification to the Skill reproduction raised its cost from $18.59 to $29.45 while its pass rate fell from 138 of 141 to 132 of 141. Whatever visual verification costs is therefore not isolated anywhere in this table.

## A single assembly can cost more than the whole single-part benchmark

The assembly gallery publishes its own ledger, and it is a different order of magnitude from the benchmark above it. The Advanced Vision Inspection Robot Arm costs about 300,000 tokens and about $1 in total. The Three-Axis Gantry Metrology Cell runs about 178,000 tokens for planning and mating, 461,811 for part generation and repair, and about 260,157 for the Judge and other workflow, totalling about 900,000 and about $2. The Telescopic Cinema Robot Crane reaches about 1,650,000 tokens and about $3. The Heavy-Duty Mobile Manipulator runs about 520,000, then 2,678,005, then 2,602,142, for about 5,800,000 tokens and about $13.

So the most expensive assembly shown costs roughly three times the entire MAC v2 benchmark run, and its Judge column alone is larger than the whole benchmark. Two of the four assemblies are compact functional ones with no published ledger, including a hinged twin-claw gripper, a rotary fork-key tool, and a reusable five-digit hand.

## The gallery is curated output and the page says so

The section heading says Technology Preview, and the paragraph under it sets the terms: the assembly examples are curated outputs, not a measured prompt-success rate, the assembly workflow is not presented as outperforming CAD-agent skills, and generated geometry and joints should be validated before manufacturing or simulation use. Assembly generation is also kept out of the benchmark table entirely, with MAC v1 and MAC v2 defined as single-part workflows and assembly described separately.

What the pipeline produces is inspectable rather than a render. Structured briefs, geometry plans, generated Python, measurements, QA reports, and repair history are all listed as auditable output, and recovery runs through bounded feedback loops for what it calls recoverable failures. Exports are STEP, STL, and GLB, with a URDF handoff added for assemblies. The URDF demonstrations run an AI-generated hand assembly through object rotation and pick-and-place in simulation, and the part gallery is described as editable CAD results rather than image-only generations, with a print-in-place mechanism example carrying separate bodies and functional clearances.

## The requirements file opens by telling you not to install it

`requirements.txt` is labelled a reference list and says so in its own comment header: do not run `pip install -r requirements.txt` directly. The reason given is a version conflict. `aider-chat` pins `numpy==1.26.4` while `build123d>=0.8` requires `numpy>=2,<3`, which makes a pure pip install fail with ResolutionImpossible. The recommended path is conda, `conda env create -f environment.yml` followed by `conda activate multi_agent_cad`, so native C extensions for numpy, scipy, trimesh, and rtree come from conda-forge prebuilt wheels. On Windows the file says to prefer conda because native wheels are unreliable.

One gap sits between the two dependency lists. `aider-chat>=0.50` appears under the heading Autonomous code repair engine in `requirements.txt`, and `aider` appears in the project keywords, but the installable dependencies in `pyproject.toml` are build123d, langgraph, langgraph-checkpoint, pydantic, openai, and anthropic, with no aider entry. The repair engine is therefore not something `pip install multi-agent-cad` pulls for you.

## Orchestration is held on an old LangGraph for Python 3.11

The orchestration stack is pinned narrowly and the reason is written next to it: `langgraph>=0.2,<0.3`, pinned for Python 3.11 compat, with `langgraph-checkpoint>=2.0,<3.0` alongside it. Both ranges come from different release lines, and only the first carries an explanation. The package requires Python 3.11 or newer while its classifiers name 3.11 and 3.12 only, so 3.13 is outside the tested set even though the constraint would accept it.

Model choice is per stage. The client dependency is an OpenAI-compatible one, with the comment naming DashScope, Qwen, OpenAI, and DeepSeek, and `anthropic>=0.30` is present as an optional fallback provider. The README makes the same promise in feature terms, that different OpenAI-compatible models can be configured for planning, geometry, coding, and repair. Extras split the rest: a mesh extra for STL analysis with trimesh and rtree, a science extra for numpy, scipy, and scikit-learn, a dev extra for ipython, pytest, and ruff, and a web extra pairing fastapi with uvicorn, which is what backs the browser UI with 3D preview.

## Three version labels and one tag that is a video

The benchmark table names MAC v1 and MAC v2, the packaging metadata records version 1.0.0, and the narrative calls this release the one that adds visual verification, an independent Judge, and assembly generation. Those are three labels for one project, and the tags make it four. Two releases exist: v1.0.0 on 2026-09-24, and ui-demo-v1 on 2026-08-04, whose release name is a Web UI walkthrough video rather than a build. The last recorded push to the repository is 2026-09-24.

The tree shows how the two halves of the project are kept apart. `multi_agent_cad/` holds the original single-part workflow whose documentation covers configuration, execution, caching, and QA, `mac_assembly/` holds the assembly workflow with its own README describing implementation details and current limitations, and `legacy_refs/` sits beside them. `docs/`, `tests/`, `packages/`, and `assets/` complete the layout, with a `.lfsconfig` and `.gitattributes` indicating that the gallery imagery is tracked through Git LFS.

## Conclusion

MAC is worth a look if you generate mechanical parts from language and want to inspect the work rather than receive a render, because the pipeline keeps structured briefs, geometry plans, generated Python, measurements, QA reports, and repair history, and it exports STEP, STL and GLB instead of an image. Two things to settle first. The headline comparison moves from one model and one verification mode to another between its v1 and v2 rows, so the cost of visual verification is not isolated anywhere in the table. Second, the benchmark is ten prompts and the assembly gallery is curated output, so neither number is a success rate, and the page's own advice stands: validate geometry and joints before manufacturing or simulation. Install through conda rather than pip, since the pinned NumPy ranges in the project and its requirements file cannot resolve together.

## FAQ

### What does the MAC Multi-Agent-CAD framework generate?

Editable CAD parts and multi-part assemblies from natural language input. Outputs are exported as STEP, STL, and GLB rather than a render-only result, with a URDF handoff for assemblies, and the pipeline keeps structured briefs, geometry plans, generated Python, measurements, QA reports, and repair history so a run can be inspected.

### How much does a Multi-Agent-CAD assembly run cost in tokens?

The published complex-gallery ledger runs from about 300,000 tokens and about $1 for the Advanced Vision Inspection Robot Arm to about 5,800,000 tokens and about $13 for the Heavy-Duty Mobile Manipulator, whose part generation and repair alone records 2,678,005 tokens. Those figures are labelled curated outputs rather than a measured prompt-success rate.

### What installation problems does Multi-Agent-CAD have?

The requirements file states directly that `pip install -r requirements.txt` must not be run, because `aider-chat` pins `numpy==1.26.4` while `build123d>=0.8` requires `numpy>=2,<3`, which makes pure pip fail with ResolutionImpossible. The recommended path is `conda env create -f environment.yml` then `conda activate multi_agent_cad`, and on Windows conda is preferred because native wheels are unreliable.

### Which models can Multi-Agent-CAD use at each stage?

Stages are model-flexible, with an OpenAI-compatible client named for DashScope, Qwen, OpenAI, and DeepSeek, and `anthropic>=0.30` included as an optional fallback provider. Planning, geometry, coding, and repair can each be pointed at a different model, and the benchmark rows were run on Qwen 3.7 and Qwen 3.8.

### Is Multi-Agent-CAD assembly generation ready for manufacturing?

The repository presents it as a technology preview: the gallery entries are curated outputs rather than a measured prompt-success rate, the workflow is not presented as outperforming CAD-agent skills, and generated geometry and joints should be validated before manufacturing or simulation use. Assembly figures are excluded from the ten-prompt single-part benchmark numbers.

## Sources

- [Issues](https://github.com/Pan-Chera/Multi-Agent-CAD/issues)
- [License: MIT](https://github.com/Pan-Chera/Multi-Agent-CAD/blob/main/LICENSE)
- [Pan-Chera/Multi-Agent-CAD on GitHub](https://github.com/Pan-Chera/Multi-Agent-CAD)
- [README](https://github.com/Pan-Chera/Multi-Agent-CAD/blob/main/README.md)
- [Releases](https://github.com/Pan-Chera/Multi-Agent-CAD/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pan-chera-multi-agent-cad
