distcc, and the one requirement its pump mode adds
distributed builds for C, C++ and Objective C
At a glance
- What is it?
- distcc distributes C and C++ compilation by sending preprocessed source to worker machines that need nothing but a compiler daemon, which is why it works across mixed operating systems with no shared filesystem and no clock synchronisation. Its faster pump mode moves preprocessing onto the workers too, and the price is that client and server must have identical system headers.
- Who is it for?
- Adopt distcc when you are compiling C or C++ on a machine that is not keeping up and you have spare machines that can host a daemon and a compiler, because the default mode needs almost nothing from them. Do not adopt pump mode unless client and workers really do have the same system headers, since that is the requirement it adds and a mismatch there produces a build that is wrong rather than one that fails.
- Can I use it commercially?
- Yes, with conditions. GPL-2.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 83 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What distcc actually sends across the network
distcc is a program to distribute compilation of C or C++ code across several machines on a network, and the mechanism is simpler than most distributed build systems make it sound.
The client is a front-end. It is not a compiler itself but sits in front of the GNU C/C++ compiler, or another compiler of your choice, and all the regular compiler options and features work normally. When make invokes what it thinks is the compiler, the front-end does the preprocessing locally, sends the complete preprocessed source code for that job to a worker, and returns the object file.
That design has one large consequence, and the README states it plainly: all distcc requires of a volunteer machine is that it be running the `distccd` daemon and have an appropriate compiler installed. No source tree, no build directory, no configuration. The worker never sees your project structure, only a stream of preprocessed translation unit and a request for an object file back.
The economics follow from that. The README notes that shipping files across the network takes time but few cycles on the client machine, so anything built remotely is essentially free in client CPU. The cost is bandwidth and latency, and the reason distcc is designed to be used with GNU make's parallel build is that a make job queue already has the property you need, which is many independent translation units compiled at once. Scale the make parallelism and you scale the number of machines working.
The payoff the project claims is two or more times faster than a local compile, with near-linear scaling for small numbers of machines, the specific figure given being three machines at 2.6 times the speed of one. It also says distcc has been used in environments with hundreds of servers supporting dozens of simultaneous compiles, and that it can successfully compile the Linux kernel, rsync, KDE, GNOME through GARNOME, Samba and Ethereal. Those are the project's own figures and its own list, not measurements you can check, and they are worth repeating as claims rather than as results.
What is not in dispute is the design's logic. A C++ build spends most of its time in the compiler, and the compiler is a pure function of its input, so moving it somewhere else cannot change the answer as long as the same compiler and the same preprocessed input are used.
The invariant: the same results as a local compile
The first claim in the README is that distcc should always generate the same results as a local compile. That sentence carries the entire correctness argument, and it is worth understanding what it depends on.
In default mode, the preprocessor runs on the client. That means every macro definition, every conditional compilation branch and every header inclusion has already been resolved before anything crosses the network. What arrives at the worker is a flat translation unit with no `#include` lines left and no conditional branches, so the worker does not need your headers, your include paths, or your compiler's predefined macros. The only thing the worker has to agree with the client about is the compiler itself and the target binary format.
That is why the list of non-requirements in the README is as long as it is. Unlike other distributed build systems, distcc does not require all machines to share a filesystem, to have synchronised clocks, or to have the same libraries or header files installed. Machines can be running different operating systems, as long as they have compatible binary formats or cross-compilers.
Each of those three non-requirements is load-bearing. A shared filesystem is unnecessary because the worker never reads your tree. Synchronised clocks are unnecessary because nothing in the protocol depends on timing, which is a genuine difference from schedulers that lease work by timestamp. Matching headers are unnecessary in default mode precisely because the preprocessor already flattened them. And cross-compilation works because the worker is running a compiler, not your build system, so a worker can be a Linux machine producing a Windows object file if the cross-compiler is installed there.
The failure mode follows from the same logic, and it is the one to test for. If a worker's compiler differs from the client's in a way that changes codegen, you do not get an error, you get an object file that behaves differently. Nothing in the design prevents that, which is why the projects distcc claims to compile are all large and well-tested: a kernel build or a Samba build will fail loudly somewhere else if the toolchain disagrees. For a small project, the sensible first step is to compile it both ways and compare.
Pump mode, and the identical-headers requirement it introduces
Pump mode is the feature that makes distcc faster than its own design otherwise allows, and it is also the feature that changes the requirements on the workers.
The README describes it as functionality added in distcc 3.0 that distributes not only compilation but also preprocessing to the servers. The pump functionality is credited to Fergus Henderson, Nils Klarlund, Manos Renieris and Craig Silverstein of Google, which explains both why it exists and why it arrived when it did: preprocessing is a substantial part of a C++ build's cost, and moving it off the client is a large win if you have idle machines.
The payoff is stated as equal results, faster. In non-pump mode the client preprocesses every file locally, which burns client CPU and produces a large stream, because a preprocessed translation unit has every system header inlined into it. Pump mode sends the actual source and headers to the worker and preprocesses there, so the client spends almost nothing and the data transferred is smaller. The README's summary of the whole feature is that the preprocessor no longer runs locally.
The cost is a hard requirement that default mode does not have: the server and the client must have the same system headers, with the client taking responsibility for transmitting application-specific headers. That sentence is doing all the work. System headers are the compiler's own, from the same distribution or the same build, and if a worker has a different glibc, a different C++ standard library version, or a different kernel header set, the preprocessed output it produces from identical source will differ. In default mode that difference is impossible because the client already made the decisions. In pump mode the worker makes them, and it needs the same inputs to make them the same way.
This is why the repository contains an `include_server` directory. A mode in which the worker preprocesses needs a way to obtain the client's headers, and a directory with that name is the mechanism. It is the concrete cost of the feature, sitting in the tree next to the code that benefits from it.
The practical advice follows directly. Default mode is nearly free to adopt because a worker needs only a daemon and a compiler, and any spare machine qualifies. Pump mode is faster and considerably more fragile, and the fragility is invisible until a header version drifts. The README points to a separate `README.pump` for the details, and that is the file to read before turning it on.
What the repository layout says about the era this code comes from
The top-level file list is a better description of this project than the README is, because it records what the maintainers considered worth shipping.
It starts with the complete classic free software set: `AUTHORS`, `COPYING`, `ChangeLog`, `INSTALL`, `NEWS`, `README`, `TODO`, plus `autogen.sh`, `configure.ac`, `Makefile.in` and an `m4` directory for autotools. That is a GNU-project layout from an era when autotools was the default and hand-written configure scripts had not yet gone out of fashion. The build system is generated rather than modern, and a first-time contributor should expect `./autogen.sh` rather than a package manager.
There are four README files, which is unusual. `README.md` is the one the repository presents, `README` is the original plain-text version, `README.pump` covers the pump mode, and `README.packaging` is aimed at distribution packagers rather than users. Having a packaging-specific document at all is a sign of a project that has been packaged by several distributions and has opinions about how.
Two implementation details in the tree are worth naming. There is an `lzo` directory, so the compression library is vendored into the repository rather than being a build dependency, which means the data stream between client and daemon is compressed and the code for it travels with distcc. And there is a `find_c_extension.sh` script, whose name tells you that distcc has to decide which files are compilable translation units, which is a smaller problem than it sounds until you consider a project with generated sources and unusual extensions in its build.
The infrastructure directories are more revealing about the project's history. There is a `gnome` directory, which is the GARNOME integration the README mentions as a project distcc can compile. There is a `docker` directory, so containerised builds are anticipated. There is a `bench` directory and a `test` directory, so both benchmarking and testing are part of the repository rather than something done ad hoc. And there is a Python script named `update-distcc-symlinks.py`, which is a reminder of how the tool gets used: distcc is installed so that make invokes it instead of the compiler, and symlinks are how that substitution is arranged.
One more detail deserves attention because it is a maintenance signal rather than a feature. The continuous integration configuration is a `.travis.yml`, and there is no other CI configuration in the top-level entries. Travis was the default hosted CI for open source projects for most of the 2010s and has since been superseded. A project whose only CI configuration names a service that is no longer the default is a project whose test suite is probably not running on every push, and that is worth knowing before you send a patch.
A lexer release in 2021 and commits in 2026
The release history is short, whimsically named, and stopped more than five years ago.
Version 3.3.3, titled Charlie the Unicorn, was released on 2019-08-14. Version 3.3.5, titled Charlie's Kidney, followed on 2021-01-04. Version 3.4, titled Lax lexer, was published on 2021-05-11. The last push to the repository was on 2026-07-08, and the repository is not archived.
The title of the most recent release is the one worth thinking about. distcc's traditional job includes working out how to split a translation unit so that the pieces can be handled as jobs, and a hand-written lexer is the classic place for a distributed compiler to be conservative, because a mis-split produces an object file that is subtly wrong rather than a build that stops. A release called Lax lexer is a deliberate relaxation of that strictness, which almost always trades a class of failures where distcc refuses to handle your code for a smaller class of failures where it handles it incorrectly.
That is a defensible trade and it is also the kind of change you want to read the notes for rather than discover. Combined with the gap since 2021, the practical picture is a project that is still being worked on but has not cut a release in five years. If you install from a distribution package you get whatever that distribution packaged, which may be 3.3 or 3.4 or a patch on top. If you build from the repository you get code that no tag describes.
The documentation has also moved. The README points to distcc.github.io as the current documents and notes that the project was formerly at distcc.org. That is a normal migration for a project that has outlived its original host, and it is a reminder that anything you find describing distcc at the old address may predate the current state.
The licence has not changed. distcc is distributed under the GNU General Public Licence version 2, with the text in the `COPYING` file, which is the same licence as the autoconf and automake tools it is built with.
distcc against a shared build server, a local cache, and a scheduler
Three alternatives, and they are not all competing for the same job.
A shared build machine with a network filesystem is the traditional answer, and it is simpler in one way and worse in several. You do not have to think about which machine ran which compile. In exchange, every machine preprocesses locally and then competes for one server, the server becomes a bottleneck exactly when the build is most parallel, and you inherit a shared filesystem that has to be fast, writable by many users, and identical in content across all of them. distcc's specific advantage over that arrangement is that the client stops doing compiler work at all, so a build on a laptop that has no local capacity can be distributed rather than queued.
A local compilation cache such as ccache is not really an alternative to distcc, it is a complement. ccache makes a rebuild of unchanged files nearly free by keeping object files keyed on the preprocessed source. distcc makes a first build fast by moving the work elsewhere. In a project where you rebuild a few files hundreds of times, ccache is the bigger win and costs nothing. In a project where every build is from clean, ccache does nothing and distcc still does. Running both is a normal arrangement, and they do not conflict because they operate at different points in the pipeline.
A scheduler-based distributed compiler, of which Icecream is the well-known example, takes a different approach again. Instead of the client choosing where a job goes, a central service hands out jobs and the machines poll for them. That is better at keeping a heterogeneous pool busy and at failing over a dead machine, and it is worse to operate, because the scheduler is a piece of infrastructure that has to be running, reachable and correct. distcc's model has no such component: the client connects to a list of machines, and a machine that is down is simply not in the list.
The wrong-tool cases are worth naming too. distcc is for C, C++ and Objective-C and it is a compiler front-end, so it does nothing for a project whose build time is dominated by linking, by a code generator, or by a test suite. And the projects it is documented to compile, the Linux kernel, rsync, KDE, GNOME, Samba, are all large C or C++ codebases, which is a fair description of the tool's intended scope.
Editorial conclusion
Adopt distcc when you are compiling C or C++ on a machine that is not keeping up and you have spare machines that can host a daemon and a compiler, because the default mode needs almost nothing from them. Do not adopt pump mode unless client and workers really do have the same system headers, since that is the requirement it adds and a mismatch there produces a build that is wrong rather than one that fails. Verify first by compiling the same project with and without distcc and diffing the objects, reading man/ in the repository for the invocation, and checking the ChangeLog, because the last tagged release is v3.4 from 2021 and the last commit is from 2026.
Frequently asked questions
What is distcc?
distcc is a program to distribute compilation of C or C++ code across several machines on a network. It is not a compiler itself but a front-end to gcc or another compiler of your choice, with all the regular compiler options working normally, and it is designed to be used with GNU make's parallel build feature.
What does a distcc worker machine need?
By default only the distccd daemon and an appropriate compiler, because distcc sends the complete preprocessed source code for each job. The worker never reads your source tree, needs no shared filesystem and no synchronised clocks, and can run a different operating system as long as the binary formats are compatible or a cross-compiler is used.
What is pump mode in distcc?
Pump mode, added in distcc 3.0, distributes preprocessing as well as compilation, so the preprocessor no longer runs on the client. The README says it yields the same results as non-pump mode but faster, at the cost of requiring the server and client to have the same system headers, with the client responsible for transmitting application-specific headers.
How much faster is distcc, according to the project?
The README says distcc is often two or more times faster than a local compile, and that it is nearly linearly scalable for small numbers of machines, giving the specific figure of three machines being 2.6 times faster than one. It also reports use in environments with hundreds of servers supporting dozens of simultaneous compiles.
When was the last distcc release?
Version 3.4, titled Lax lexer, was published on 2021-05-11, after 3.3.5 in January 2021 and 3.3.3 in August 2019. The repository's last push was on 2026-07-08 and it is not archived, so commits have continued without a new tag, and the current documentation has moved to distcc.github.io from the former distcc.org.
What licence is distcc released under?
The GNU General Public Licence version 2, with the text in the COPYING file at the repository root. The build system is autotools based, with autogen.sh, configure.ac, Makefile.in and an m4 directory, and the repository also ships man pages under a man directory and a packaging-specific README.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/distcc-distcc)