# Coral NPU: Google's Open Source RISC-V Accelerator IP for Wearables

> Coral NPU is open source hardware IP for ML inference in ultra-low-power SoCs, built on RV32IMF_Zve32x. It is RTL and a simulator, not a board, and the README says nothing about buying silicon.

**google-coral/coralnpu** — A machine learning accelerator core designed for energy-efficient AI at the edge.

- Repository: https://github.com/google-coral/coralnpu
- Stars: 2,585 · Forks: 339
- Language: Emacs Lisp
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-coral-coralnpu

## What Coral NPU actually is, and who it is for

Coral NPU is not a product you can buy. It is an Open Source IP block designed by Google Research and offered for integration into ultra-low-power System-on-Chips aimed at hearables, AR glasses and smart watches. The deliverable is hardware description, not a board. The repository's top level confirms this: hdl/, verilog/, fpga/, platforms/, toolchain/ and tests/ sit alongside sw/ and examples/, and there is no packaging directory, no BOM, no schematic.

The intended reader is a silicon or FPGA engineer who is choosing an accelerator core for a wearable SoC and wants to evaluate the RTL before licensing anything. That is a narrow audience. If you write application code, this repository has nothing for you except the two example programs in examples/, which exist to prove the toolchain and the simulator work, not to demonstrate an application.

The README describes Coral NPU as a neural processing unit, also called an AI accelerator or deep-learning processor, based on the 32-bit RISC-V ISA. That framing matters: this is a processor core with a matrix unit attached, not a fixed-function inference engine with a register interface bolted on. The programming model is instruction-driven.

## Three processors in one core: matrix, vector and scalar

The architecture is split into three cooperating processor components, and the README's own data-flow diagram is the primary reference for how they connect. The scalar component is a four-stage, in-order dispatch pipeline with out-of-order retire, dispatching four ways. The vector component is SIMD, dispatching two ways, with a 128-bit datapath and a 256-bit pipeline listed as future. The matrix component handles the dense arithmetic that dominates inference.

The instruction set is the concrete commitment: rv32imf_zve32x_zicsr_zifencei_zbb, described in the feature list as RV32IMF_Zve32x. The zve32x extension is what makes the vector unit a standard RISC-V vector target rather than a proprietary ISA, and zbb adds the bit-manipulation instructions. That means an existing RISC-V compiler can target the core, subject to the toolchain supporting those extensions.

Memory is tightly coupled rather than cached: 8 KB of ITCM for instructions and 32 KB of DTCM for data, both single-cycle-latency SRAM. The README argues this is more efficient than cache memory. The trade-off is capacity. A model that does not fit in 32 KB of data memory has to stream over the AXI4 interface, and the README does not describe a DMA engine or a prefetch scheme for that case. AXI4 is used both as manager and subordinate, so external CPUs can configure the NPU and the NPU can reach external memory.

## Building Coral NPU and running the hello-world on the simulator

The system requirements are stated exactly: Bazel 8.6.0 and Python 3.9 to 3.13. The README points to utils/coralnpu.dockerfile for the detailed list and notes that CI runs most builds and tests using that image, which is the practical way to match the environment. The repository pins .bazelversion and MODULE.bazel, so Bazel will fetch its own dependencies.

The quick start has four steps. First, confirm the test suite passes:

```bash
bazel run //tests/cocotb:core_mini_axi_sim_cocotb
```

That target runs the cocotb RTL simulation. If it passes, the toolchain and the RTL are consistent. Next, build one of the example binaries:

```bash
bazel build //examples:coralnpu_v2_hello_world_add_floats
```

The README labels this a non-RVV build path, meaning it does not exercise the vector unit. The output is an ELF under bazel-out. Then build the simulator itself:

```bash
bazel build //tests/verilator_sim:core_mini_axi_sim
```

Finally, run the binary on it:

```bash
bazel-bin/tests/verilator_sim/core_mini_axi_sim --binary bazel-out/k8-fastbuild-ST-dd8dc713f32d/bin/examples/coralnpu_v2_hello_world_add_floats.elf
```

Note that the path in the README's last command contains a build-configuration hash. If your Bazel configuration differs, that segment will differ too, and you should take the path from the build output rather than copying the line verbatim. The second example, examples/rvv_add_intrinsic.cc, is the one that exercises the vector extension.

## What you cannot learn from the README

Several things a hardware integrator needs are simply absent. There is no area figure, no power figure, no clock-frequency target and no process node mentioned. For an IP block whose entire purpose is energy efficiency in wearables, the absence of any published power or area number is the largest gap, and it means you cannot compare Coral NPU against another accelerator on the evidence here. You would have to synthesize it yourself and measure.

The README also does not document rollback, versioning policy or long-term support for the RTL interface. The release tags (m3-initial and M3-2026-04-27) suggest a milestone-based scheme, but the README does not explain what a milestone guarantees or whether the register interface is frozen between them. An integrator who tapes out against m3-initial has no stated contract about what changes in M3-2026-04-27.

The primary language shown for the repository is Emacs Lisp, which is almost certainly an artifact of how the languages are counted rather than a description of the design. The real content is in hdl/, verilog/ and the test directories. Treat that label as noise.

Finally, the verification story is split across two testbenches with separate READMEs: cocotb for RTL and netlist simulation, and UVM for co-simulation. The top-level README defers to both. That is reasonable, but it means the quality of the verification evidence depends on documents you have to read separately, and the top-level page gives no summary of coverage.

## Coral NPU compared with a fixed-function NPU and with a TPU

The natural alternative is a fixed-function NPU IP block: a register-mapped accelerator with its own command queue, a compiler that lowers a model to that queue, and no user-visible instruction set. Those blocks typically win on area and power for a fixed set of operators, because the control logic is smaller than a general pipeline. Coral NPU takes the opposite bet. It exposes a RISC-V core with vector and matrix units, so the same block can run non-ML code and the ML kernels are written against a standard ISA.

That is a real difference in approach, not a marketing one. With a fixed-function block, an operator your compiler does not support is a hard failure. With Coral NPU, an unsupported operator is a kernel you write in RV32IMF_Zve32x and compile with a RISC-V toolchain. The cost is that you now own a compiler target and a kernel library, and the README does not describe either.

Against a TPU, the distinction is scale and deployment. Coral NPU is aimed at wearables and ships as synthesizable RTL. The repository's fpga/ and platforms/ directories indicate FPGA prototyping paths, which is the usual way to evaluate this class of IP before committing to silicon. Anyone searching for Coral NPU vs TPU should understand that these are different tiers of hardware, and the README makes no comparison.

## Licence, maintenance and the cost of tracking upstream

The repository is licensed Apache-2.0, which is the permissive licence most silicon teams expect for integration IP. It permits commercial use and modification, and it includes a patent grant. It does not give legal advice, and it does not address whether the RISC-V extensions implemented here require anything further from you. Read LICENSE and CONTRIBUTING.md before you build a product plan on it.

On maintenance, the last push to main was on 2026-09-18, and the repository is not archived. The most recent release tag is M3-2026-04-27, dated 2026-04-27. Those two facts are the only maintenance signals available here, and they are consistent with a project that is still moving.

The upgrade cost is the part that is easy to underestimate. Coral NPU is a Bazel workspace with pinned module dependencies, a custom toolchain directory and two simulation environments. Every time you move to a new milestone you are re-running the cocotb and UVM suites against your own modifications, because you will almost certainly have modified the RTL for your SoC. There is no stated backport policy and no LTS branch, so plan on tracking main or on forking at a tag and carrying your own patches. Budget for the Verilator build too: the README itself notes that the non-RVV build is chosen for shorter build time, which implies the full vector build is slower.

## Conclusion

Adopt Coral NPU if you are integrating an NPU into an ultra-low-power SoC and your team already works in SystemVerilog, Bazel and cocotb or UVM. Do not adopt it if you want a board to plug in or a finished driver stack: there is no packaged hardware and the README points only at RTL simulation. Verify the RV32IMF_Zve32x toolchain end to end before committing, because the quick start builds a hello-world binary for a simulator, not for silicon.

## FAQ

### What is Coral NPU?

It is an open source hardware accelerator for ML inference, designed by Google Research and offered as IP for integration into ultra-low-power SoCs such as hearables, AR glasses and smart watches. It is a neural processing unit based on the 32-bit RISC-V ISA, with matrix, vector and scalar processor components.

### How do I use Coral NPU?

The README's quick start is a Bazel workflow: run the cocotb test target to confirm the suite passes, build an example binary, build the Verilator simulator, then run the binary on the simulator. Bazel 8.6.0 and Python 3.9 to 3.13 are the stated requirements.

### How does Coral NPU compare with a TPU?

The README makes no comparison. Coral NPU is synthesizable RTL for wearable-class SoCs with 8 KB ITCM and 32 KB DTCM, while a TPU is a different tier of hardware; the repository only describes Coral NPU's own architecture and its AXI4 interfaces.

### What can I use instead of Coral NPU?

The README does not name alternatives. The closest design choice it implies is a fixed-function NPU IP block with a register interface rather than a RISC-V core with vector and matrix units, which trades the ability to write your own kernels for smaller control logic.

### How much does Coral NPU cost?

The repository does not state a price. Coral NPU is published as Open Source IP under Apache-2.0, so there is no listed cost for the design itself, and the README describes no packaged hardware or board that could be purchased.

### Is Coral NPU owned by Google?

The README states that Coral NPU is an Open Source IP designed by Google Research. The repository sits under the google-coral organization and is licensed Apache-2.0.

## Sources

- [google-coral/coralnpu on GitHub](https://github.com/google-coral/coralnpu)
- [Issues](https://github.com/google-coral/coralnpu/issues)
- [License: Apache-2.0](https://github.com/google-coral/coralnpu/blob/main/LICENSE)
- [README](https://github.com/google-coral/coralnpu/blob/main/README.md)
- [Releases](https://github.com/google-coral/coralnpu/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-coral-coralnpu
