# Erigon: an Ethereum execution layer built around immutable snapshot files

> Erigon is an Ethereum execution client that stores most chain data in immutable segment files and ships its own embedded consensus layer. It is aimed at node operators who care about disk layout and sync behaviour more than about drop-in familiarity with Geth.

**erigontech/erigon** — Ethereum implementation on the efficiency frontier 

- Repository: https://github.com/erigontech/erigon
- Website: https://docs.erigon.tech
- Stars: 3,583 · Forks: 1,551
- Language: Go
- License: LGPL-3.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/erigontech-erigon

## What Erigon is for, and who ends up running it

Erigon describes itself as "an implementation of Ethereum (execution layer with embeddable consensus layer), on the efficiency frontier." The wording matters: this is a full execution client, not a library, and the consensus layer is embedded by default rather than bolted on as a separate process. Caplin, Erigon's own consensus layer, is enabled by default, and `--internalcl` is on unless you pass `--externalcl`.

The audience is narrower than "anyone who wants an Ethereum node." Erigon is for operators who have opinions about where bytes live. The Erigon3 design stores most data in immutable files, which the README calls segments or snapshots, and keeps `chaindata` comparatively small: the README gives 22 GB on an archive mainnet node as an example. That split lets you symlink or mount latest state on a fast drive and history on a cheap one. If you just want a node that starts and answers RPC, the extra layout decisions are overhead you did not ask for.

The README is explicit that it is written for people working on the repository, and that everything about operating a node lives on docs.erigon.tech. That is an unusual choice for a project of this size, and it means the repository you clone and the documentation you operate from are deliberately separate artifacts.

## How the Erigon3 storage model changes sync and history queries

The load-bearing change from Erigon2 to Erigon3 is that initial sync does not re-execute from block zero. According to the README, the node downloads 99% LatestState and History instead. That is a different starting assumption from a client that replays the chain, and it is why snapshot files carry so much of the weight.

History granularity moved from per-block to per-transaction. The README lists the consequences directly: you can execute a single historical transaction without executing its block, and if an account changes V1 to V2 and back to V1 within one block across different transactions, `debug_getModifiedAccountsByNumber` returns it. That last detail is a correctness property, not a performance one, and it is the kind of thing that only shows up when someone queries historical state for an audit.

Erigon3 does not store logs, which the README also calls receipts. Historical transactions are re-executed on demand instead, and the README claims this is cheaper than storing them. Treat that as a design trade: you trade disk for CPU on historical log queries, and the balance point depends on your query pattern.

The execution stage absorbed several Erigon2 stages: stage_hash_state, stage_trie, log_index, history_index and trace_index now live inside it. Restart behaviour also changed, with `--sync.loop.block.limit=5_000` enabled by default so a restart does not discard as much partial progress.

## Installing Erigon and running a first node

The README gives build instructions for people working on the repository. The toolchain is Go 1.26 or newer, GCC 11+ or Clang 13+, on a 64-bit architecture, with Linux kernel above v4. On x86-64 the build targets the `x86-64-v2` baseline, which needs SSE4.2 and POPCNT, so Intel Nehalem (2008) and AMD Bulldozer (2011) or newer qualify. Older hardware is not supported.

Clone and build the node:

```bash
git clone https://github.com/erigontech/erigon.git
cd erigon
make erigon
```

Binaries land in `./build/bin`. `make all` builds the full suite there, and `make help` lists every target. Use `-j<n>` to parallelise. If you want a specific release rather than `main`, check out its tag before building.

For anything other than a source build, the README points at the installation page, which covers packaged binaries, Docker and per-platform steps for Linux, macOS, Windows and ARM.

The repository also ships `docker-compose.yml`, which is an example of running Erigon's services as separate processes rather than one binary. Its header warns not to split services "unless you have clear reason for it: resource limiting, scale, replace by your own implementation, security." The compose file documents the default datadir as `/home/erigon/.local/share/erigon`, default UID and GID as 1000, and lists the ports: 9090 execution engine private API, 9091 sentry, 9092 consensus engine, 9093 snapshot downloader, 9094 TxPool, 8545 JSON-RPC, 8551 consensus JSON-RPC, 30303 eth p2p, 42069 bittorrent. `.env.example` exposes `ERIGON_USER`, `DOCKER_UID` and `DOCKER_GID` for the Docker path.

Once running, the RPC surface is documented under interacting with Erigon, covering JSON-RPC, GraphQL and gRPC namespaces. `rpcdaemon` can run in-process or standalone against a local or remote Erigon.

## Pruning modes and the disk trade you are actually making

Pruning is where Erigon asks you to make a decision you cannot easily undo. The `--prune` flags changed in Erigon3 and are now expressed through `--prune.mode`, whose default is `full`. The other documented modes are `archive`, and `minimal`, which the README ties to EIP-4444.

The default matters more than it looks. A node started without thinking about pruning is in `full` mode, not archive, so historical state queries that work on an archive node will not behave the same way. The pruning modes page is the place where the semantics are defined; the README only names the modes.

The immutable-file design interacts with this. Because `chaindata` holds tens of gigabytes rather than the whole chain, the README notes it is acceptable to `rm -rf chaindata`, and recommends `--batchSize <= 1G` to stop it growing. That is a real recovery option, but it is also a reminder that the directory is treated as disposable in a way that a monolithic database file is not. Anyone who has built backup scripts around a single database file will need to rethink them here.

## Where Erigon is the wrong tool

The clearest limitation is operational: the README does not document rollback. It points to an upgrading page for versions and snapshot formats, but nothing in the repository README describes how to move a running node back to a previous release or a previous snapshot layout. If your environment requires a tested downgrade path before you deploy, that path is not established by the documentation here.

Second, the embedded consensus layer is a commitment, not a convenience. Erigon's stated reason for building Caplin rather than relying on the Engine API is architectural: the Engine API delivers blocks one at a time, which does not suit Erigon's bulk model of handling many blocks simultaneously and sorting and processing data in batches. That reasoning is sound for Erigon, but it means the consensus side is maintained by the same project, and it is on by default. If your operational model already standardizes on an external consensus client, you are running against the grain by disabling `--internalcl`.

Third, the hardware baseline excludes older machines outright. The `x86-64-v2` requirement is a hard floor, not a recommendation, and pre-2008 Intel or pre-2011 AMD hardware simply will not run a build from this repository.

Finally, the split between README and docs means you cannot evaluate operational fit from the repository alone. Flag semantics, ports, monitoring and staking live on the documentation site, and the README says so plainly.

## Erigon versus Geth and Reth: different bets on storage

The comparison people actually search for is Erigon against Geth, and the honest answer is that they disagree about where the chain lives. Geth's long-standing model is a database-backed node in which state and history sit in the same store; Erigon3 pushes most data into immutable segment files and keeps a comparatively small mutable `chaindata` directory alongside them. That difference is what makes the Erigon layout advice (fast drive for latest state, cheap drive for history) possible in the first place, and it is also what makes Erigon's backup and recovery story different.

Reth is the other comparison worth naming, because it is also a from-scratch rethinking of execution client architecture in a systems language. Both projects reject the assumption that the chain must be replayed from genesis on first sync, and both invest heavily in how data is laid out on disk. The documentation here does not give enough detail on Reth's internals to compare mechanisms fairly, so the useful distinction is narrower: Erigon bundles its own consensus layer and treats that bundling as a deliberate architectural position, while Erigon's README frames the Engine API's one-block-at-a-time delivery as fundamentally mismatched to its batch processing model.

Nethermind is the third name that comes up. Without documentation for it in this repository, the only defensible statement is that all four are independent implementations of the same execution-layer specification, and the practical differences for an operator show up in storage layout, pruning semantics and how much of the stack ships in one binary.

## Licence, maintenance and what upgrades cost you

Erigon is licensed under LGPL-3.0. That is a copyleft licence with a linking exception tradition, and it sits differently from the permissive licences many Go infrastructure projects use. If you embed Erigon components into a larger product, the obligations around modified library code are worth reading with your own counsel; nothing here is legal advice, and the COPYING file in the repository is the authoritative text.

The repository is not archived, and the last push was on 2026-09-23. Recent releases are v3.6.1 on 2026-09-09, v3.6.0 on 2026-08-24 and v3.5.5 on 2026-08-12. Those dates show a steady release cadence rather than a frozen tree, but cadence is not the same as upgrade safety.

The upgrade cost is where Erigon differs from clients that treat a release as a binary swap. The README links an upgrading page that covers "versions and snapshot formats," which tells you the snapshot format is versioned and therefore something an upgrade can change. Combined with the fact that `chaindata` is designed to be removable, the practical implication is that an upgrade plan needs to account for snapshot format changes, not just the binary. That page is the first thing to read before moving a production node between minor versions.

## Conclusion

Adopt Erigon if you want a node whose bulk data lives in immutable snapshot files you can mount on separate disks, and you accept that operating it means reading the documentation site rather than the README. Do not adopt it if you need a node that behaves exactly like Geth, or if you are unwilling to plan disk layout up front. Before committing, verify the current pruning mode semantics on the pruning modes page, the hardware requirements page, and the upgrading page for the snapshot format your release uses.

## FAQ

### How does Erigon compare with Geth?

The documentation here does not describe Geth's internals, so the comparison is limited to what Erigon states about itself: Erigon3 stores most data in immutable segment files, keeps chaindata comparatively small, and downloads 99% LatestState and History instead of re-executing from block zero on initial sync.

### How does Erigon compare with Reth?

The README does not document Reth's internals, so no mechanism-level comparison is possible from this repository. What can be said is that Erigon ships an embedded consensus layer called Caplin, enabled by default, and frames the Engine API's one-block-at-a-time delivery as mismatched to its own batch processing model.

### How does Erigon compare with Nethermind?

The repository material does not cover Nethermind, so no direct comparison can be made here. Both are independent implementations of the Ethereum execution layer, and the README describes Erigon's own differentiators as immutable segment files, per-transaction history granularity and an embedded consensus layer.

## Sources

- [erigontech/erigon on GitHub](https://github.com/erigontech/erigon)
- [License: LGPL-3.0](https://github.com/erigontech/erigon/blob/main/LICENSE)
- [Project website](https://docs.erigon.tech)
- [README](https://github.com/erigontech/erigon/blob/main/README.md)
- [Releases](https://github.com/erigontech/erigon/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/erigontech-erigon
