DREAMPlace: GPU-Accelerated VLSI Placement via Deep Learning Toolkit
Deep learning toolkit-enabled VLSI placement
At a glance
- What is it?
- DREAMPlace is an open-source C++/Python VLSI placement tool that frames chip placement as a differentiable optimization problem and uses PyTorch as the compute backend to run global and detailed placement on NVIDIA GPUs, achieving over 30x speedup on ISPD 2005 benchmarks compared to the CPU-based RePlAce placer.
- Who is it for?
- DREAMPlace is the right choice for researchers and EDA teams who need a programmable, GPU-accelerated placement engine and are comfortable with academic software that requires manual dependency management. The strict GCC version requirement (7.5 recommended, 9 and later not recommended) and the PyTorch 1.6 to 2.0 constraint are real friction points for teams on modern Linux distributions.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 74 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What DREAMPlace does and who it is for
Placing millions of cells onto a chip floorplan is one of the most compute-intensive steps in VLSI design. The traditional approach treats placement as a continuous nonlinear optimization problem and solves it with custom C++ solvers that run on CPU. DREAMPlace reframes global placement as an optimization problem analogous to training a neural network: cell positions are continuous variables, the density and wire length penalties are differentiable functions, and a GPU can compute gradients across millions of cells simultaneously.
The result is that existing deep learning toolkits, specifically PyTorch, can serve as the optimization engine without requiring a custom GPU kernel for every placement operation. DREAMPlace exposes the placement state as PyTorch tensors and uses PyTorch's automatic differentiation to drive the solver. The C++ layer handles the computationally intensive primitives like density and routing estimation.
The intended audience is EDA researchers who want to experiment with placement algorithms, teams benchmarking against academic placers like RePlAce and NTUPlace3, and anyone who needs a GPU-accelerated placement baseline that produces ISPD-compatible results. The tool runs on both CPU and GPU; on a machine without a GPU, it falls back to multi-threaded CPU execution. The reported 30x speedup over RePlAce in global placement and legalization on ISPD 2005 contest benchmarks was measured on an Nvidia Tesla V100 GPU.
The analogy between placement and neural network training
Global placement in DREAMPlace minimizes a combination of total wire half-perimeter length and cell overlap density. Both objectives are expressed as differentiable loss functions over the cell position tensor. The optimizer performs gradient descent on this loss, nudging cells toward positions that reduce wire length while spreading density across the chip.
The analogy to neural network training is direct: the cell positions are parameters, the wire length and density terms are the loss function, and the GPU-accelerated PyTorch autograd machinery computes the gradient updates. This design means researchers can swap optimizer algorithms (SGD, Adam, or a nonlinear conjugate gradient through the ncg_optimizer package) without rewriting the C++ core.
The electrostatic analogy in DREAMPlace models cell density as a charge distribution on the placement canvas. The electric potential and field derived from that distribution form the gradient signal that pushes overlapping cells apart. This approach was introduced in the ePlace method and extended through DREAMPlace 4.0's timing-driven weighting and multi-electrostatics formulation.
Publications spanning DAC 2019 through ASPDAC 2026 describe successive versions: DREAMPlace 1.0 (global placement on GPU), 2.0 (global and detailed), 3.0 (multi-electrostatics with region constraints), 4.0 (timing-driven with momentum-based net weighting), 4.1 (second-order information for mixed-size), and 4.3 (static timing analysis integration via HeteroSTA).
Building DREAMPlace with Docker
DREAMPlace has a non-trivial dependency chain. The README lists Python 3.5 through 3.9, PyTorch 1.6 through 2.0, GCC 7.5 with C++17 support, Boost 1.55.0 or later, and Bison 3.3 or later. It explicitly does not recommend GCC 9 or later due to backward compatibility issues. Several C++ components are integrated as git submodules: Limbo, Flute, OpenTimer (a modified fork), and CUB.
The Docker image defined in the repository's Dockerfile encapsulates the dependency setup and is the most reliable starting point. The base image is pytorch/pytorch:1.7.1-cuda11.0-cudnn8-devel. The Dockerfile installs system libraries, Bison via conda, and a set of Python packages:
apt-get install -y flex libcairo2-dev libboost-all-devconda install -y -c conda-forge bisonThe Python dependencies installed inside the Docker image match the repository's requirements.txt:
pip install pyunpack>=0.1.2 patool>=1.12 matplotlib>=2.2.2 \
cairocffi>=0.9.0 pkgconfig>=1.4.0 setuptools>=39.1.0 \
scipy>=1.1.0 numpy>=1.15.4 shapely>=1.7.0After the Python dependencies are in place, DREAMPlace itself is built with CMake. The repository root contains CMakeLists.txt, and the build follows the standard CMake out-of-source pattern. The README separates the build into two paths: building with Docker and building without Docker, each with its own section.
For researchers working without Docker on a modern Linux machine, the GCC version constraint is the main obstacle. GCC 7.5 is from 2017 and is not the default on Ubuntu 22.04 or later. It must be installed separately or the build must be configured to use an older toolchain.
ABCDPlace: batch-based concurrent detailed placement
Global placement produces cell locations that minimize wire length but leave many cells overlapping. A detailed placer resolves these overlaps while making small adjustments to preserve wire length quality. DREAMPlace integrates ABCDPlace, a GPU-accelerated detailed placer described in a separate TCAD 2020 paper.
ABCDPlace organizes cells into independent batches that can be moved concurrently without conflicting with one another, then processes each batch on the GPU in parallel. On million-size benchmarks, the README reports approximately 16x speedup over NTUPlace3, the widely used sequential CPU detailed placer. The speedup figure comes from published benchmarking on ISPD 2005 contest benchmarks; the exact hardware used is described in the cited TCAD paper.
The practical consequence is that detailed placement, which is typically the slower sequential phase following global placement, no longer limits the overall tool throughput when a GPU is available. On a machine without a GPU, ABCDPlace falls back to multi-threaded CPU execution, the same as the global placement stage.
Dependency constraints and CPU-only fallback
The GCC constraint is the most significant barrier to adoption on current systems. GCC 7.5 predates several language and standard library changes that GCC 9 introduced. The README does not document a workaround for building with GCC 9 or later; it only notes that those versions are not recommended. Teams on Debian 11 or Ubuntu 22.04 will encounter this immediately, as both ship with GCC 10 or later as the default.
The PyTorch version range (1.6 through 2.0) is narrow relative to the current PyTorch release cadence. The Dockerfile pins pytorch/pytorch:1.7.1-cuda11.0-cudnn8-devel, which corresponds to CUDA 11.0. Running on more recent CUDA versions (12.x) requires building the Docker image from a different base or adjusting the build configuration manually.
On a CPU-only machine, DREAMPlace compiles without GPU support and uses multi-threading for the placement kernels. This mode is documented but produces slower results than the GPU path. For researchers who only have access to CPU hardware, the tool remains functional but loses the speedup that motivates its adoption.
Limitations and comparison with OpenROAD
DREAMPlace is a placement-only tool. It does not include synthesis, floorplanning, routing, sign-off, or any other step in the physical design flow. A team needs to bring their own netlist, benchmark, or synthesis output and pipe DREAMPlace's placed output into a separate routing step. The repository includes a benchmarks directory for obtaining ISPD 2005 contest benchmarks, which are the primary reference for the reported speedup numbers, but it does not bundle a router or a complete flow script.
OpenROAD, the open-source RTL-to-GDS flow from DARPA's OpenROAD project, includes a placement engine called RePlAce (on CPU) and a GPU-accelerated version in development. OpenROAD provides a complete flow from synthesis to sign-off, which is architecturally different from DREAMPlace's single-step role. Researchers who want a placement-only component they can embed in custom flow experiments will find DREAMPlace more modular; teams who want a complete open-source tapeout flow will find OpenROAD more practical.
The README also explicitly recommends HeteroPlace for industrial designs that need a holistic GPU-accelerated placement and routing solution. HeteroPlace is a separate project hosted at heteroplace.pkueda.org.cn.
The repository has no GitHub releases. Versions correspond to the publication history (4.0, 4.1, 4.3) rather than tagged release artifacts on GitHub.
Editorial conclusion
DREAMPlace is the right choice for researchers and EDA teams who need a programmable, GPU-accelerated placement engine and are comfortable with academic software that requires manual dependency management. The strict GCC version requirement (7.5 recommended, 9 and later not recommended) and the PyTorch 1.6 to 2.0 constraint are real friction points for teams on modern Linux distributions. For production industrial designs, the README points to HeteroPlace as a more complete GPU-accelerated placement and routing solution. The project is BSD-3-Clause licensed, with the last push on 2026-07-18.
Frequently asked questions
Does DREAMPlace require a GPU to run?
No. The README states that DREAMPlace runs on both CPU and GPU. On a machine without a GPU, only CPU support is enabled and multi-threading is used instead of GPU parallelism.
Which PyTorch and Python versions does DREAMPlace support?
The README lists Python 3.5 through 3.9 and PyTorch 1.6, 1.7, 1.8, and 2.0 as tested versions. Other versions may work but are not tested. The included Dockerfile uses pytorch/pytorch:1.7.1-cuda11.0-cudnn8-devel as the base image.
Why does DREAMPlace recommend GCC 7.5 instead of a newer compiler?
The README recommends GCC 7.5 for its C++17 support and explicitly notes that GCC 9 or later is not recommended due to backward compatibility issues with the C++ components in the project and its submodules.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/limbo018-dreamplace)