# rpmalloc: a single-file, lock-free thread caching allocator for C

> rpmalloc is a public domain, cross platform memory allocator in one C source file, with 16-byte alignment and lock-free thread caches. It is aimed at C and C++ projects on Windows, Linux, macOS, iOS and Android that want to replace malloc without adopting a build system.

**mjansson/rpmalloc** — Public domain cross platform lock free thread caching 16-byte aligned memory allocator implemented in C

- Repository: https://github.com/mjansson/rpmalloc
- Stars: 2,510 · Forks: 217
- Language: C
- License: MIT
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/mjansson-rpmalloc

## What rpmalloc replaces, and for whom

The general-purpose allocator shipped with a C runtime is designed for correctness and broad compatibility, not for throughput under many threads. Every call into it has to coordinate with shared state, and the cost of that coordination grows as threads are added. rpmalloc is an alternative implementation of the same interface: it exposes rpmalloc and rpfree plus the usual aligned variants, and it can also take over the standard malloc family entirely. Allocations come back 16-byte aligned, which removes a common source of alignment bugs when code assumes malloc returns memory suitable for SIMD types or for structures with strict alignment requirements.

The audience is narrow and specific. This is for C and C++ projects that already build from source, that run on Windows, Linux, macOS, iOS or Android, and whose author is willing to compile a third-party C file into the binary. The README states the code should be portable to any platform with atomic operations and an mmap-style virtual memory API, and that the page mapping and page size can be configured at runtime to a custom implementation. That matters for embedded or console targets where the default OS-backed mapping is not what you want. It is not a library you add as a binary dependency and forget; the intended integration is source-level.

## Inside the allocator: thread caches and lock-free frees

The name describes the design. Each thread gets its own cache of memory, so a typical allocation is served from thread-local state without touching a shared lock. The README describes the library as a lock free thread caching allocator, and notes that caches are released at thread exit for reuse by other threads on platforms that support thread destructors. Cross-thread frees are explicitly part of the benchmark workload, which is the case that hurts naive thread-caching schemes: when thread A allocates and thread B frees, the block has to find its way back somewhere useful without a global lock on the hot path.

Memory is mapped from the operating system in pages. The configuration surface exposed by rpmalloc_initialize_config covers page size, huge page or transparent huge page usage, decommit behaviour, page naming, and whether to unmap on finalize. Those knobs are the difference between an allocator that returns memory to the OS aggressively and one that holds it for reuse. The README presents peak resident memory as a separate chart from throughput, which is the honest way to frame this: a thread cache that improves allocation speed also holds memory that the OS cannot reclaim. If your process has a hard memory ceiling, the decommit and unmap-on-finalize settings are the ones that matter, and the README does not give a worked example of tuning them.

There is also a separate heap API. The rpmalloc_heap_* functions provide explicit first class heaps, but they are compiled out by default: RPMALLOC_FIRST_CLASS_HEAPS defaults to 0, and the README says enabling it imposes a very slight performance hit in the deallocation path from an extra conditional instruction. That is a fair description of the trade-off. If you want per-subsystem heaps with independent lifetimes, you pay a branch on every free.

## Adding rpmalloc.[h|c] to a project and overriding malloc

The README gives the simplest integration path first: add rpmalloc.h and rpmalloc.c to your project and compile them with your sources. The allocator initializes itself with default configuration on the first allocation request, so a project that only calls rpmalloc and rpfree needs no init or fini calls at all.

If you want the standard malloc family replaced, define ENABLE_OVERRIDE to non-zero. The default is already 1, and the README says this includes malloc.c in the compilation of rpmalloc.c.

The README warns about a linker trap here, and it is the kind of thing that wastes an afternoon. If you build rpmalloc as a separate static library and call only plain malloc and free, never any rp* function, the linker may not pull the rpmalloc object out of the archive, and the overrides silently do nothing. The fix the README gives is to include rpmalloc.h in at least one source file and call rpmalloc_linker_reference, a dummy empty function provided for exactly this purpose.

The README states this is not needed if you call rpmalloc or rpfree directly, link the dynamic library, or use LD_PRELOAD or DYLD_INSERT_LIBRARIES. For C++ operator overrides the behaviour is platform specific. On Windows you must include rpnew.h in exactly one source file, because it defines the new and delete operators with external linkage and has no include guard, so including it in more than one translation unit produces duplicate symbol errors. On other platforms the operators are overridden automatically and rpnew.h must not be included. The README is direct about the limits of this feature: the list of replaced libc entry points may not be complete, and libc or stdc++ replacement should be treated as a convenience for testing on an existing code base rather than a final solution.

If you supply a custom memory interface or configuration, the rules change. The README says you must call rpmalloc_initialize or rpmalloc_initialize_config before any other call into the allocator, and that you should call rpmalloc_finalize before terminating use, to release caches, unmap virtual memory, and prepare for cleanup at process exit or library unload.

## Where rpmalloc is the wrong choice

The README's own framing of its performance section is a limitation worth reading carefully. It says the project believes rpmalloc is faster than tcmalloc, hoard, ptmalloc3 and others, and that the numbers in the charts should not be interpreted as absolute performance figures but as relative comparisons. The benchmark is a single workload: randomly sized blocks in the [16, 8000] byte range with a linear falloff distribution and cross-thread frees, on a 13th Gen Intel Core i7-13800H running Ubuntu 26.04. That is one machine, one OS, one distribution. A workload dominated by large allocations, or by long-lived objects, or by a single thread, is not what that chart measures.

The override path carries a second risk. Replacing malloc and free process-wide means every library in the process, including ones you did not build, now allocates through rpmalloc. The README acknowledges the replaced entry point list may be incomplete and calls libc replacement a testing convenience rather than a final solution. If a third-party binary expects the platform allocator's behaviour, you find out at runtime.

There is no documented security policy in the README, no fuzzing or hardening story, and no stated ABI stability guarantee for the first class heap API. For a component that sits under every allocation in a process, that silence is a real consideration. The README also does not document rollback, so a team that switches and hits a problem has the usual option of reverting the source change, with no migration notes to help. Finally, the first class heap feature is off by default and the README does not walk through a complete example of using it, so anyone who needs per-heap isolation is working from the header and the source rather than from a tutorial.

## rpmalloc against mimalloc and jemalloc

The obvious comparison points are mimalloc and jemalloc, and the difference is mostly about packaging and scope rather than raw allocation strategy. jemalloc is a large, long-lived allocator with extensive platform tuning and a configuration surface aimed at production servers; adopting it usually means building and linking a library, and its behaviour is tuned through environment variables and build-time options. mimalloc is a compact allocator with its own thread-local design and a documented set of build options.

rpmalloc's distinguishing choice is that the core is a single source file of roughly 3300 lines of C, according to the README, which also argues the implementation is easier to read and modify than the alternatives. That is the real trade: you get something you can read end to end and vendor into a tree without a build system, in exchange for a smaller feature surface and less accumulated operational history. The README points to BENCHMARKS.md for a full comparison across the mimalloc-bench suite, including peak memory, so the project does invite the head-to-head comparison rather than avoiding it. If your decision hinges on measured peak memory under your own workload, that file and the script in the benchmark directory are where to start, not the README's summary chart.

## Licence, releases and what upgrading costs

The licence is unusually permissive. The README states the library is in the public domain and can be redistributed and modified without restrictions, or used under the MIT license if you prefer, declared as the SPDX expression Unlicense OR MIT. For a component compiled into a product, that removes the attribution and source-disclosure questions that come with copyleft allocators. It does not remove the need to record which licence you are relying on in your own notices; that is a decision for your legal process, not something this article can settle.

The release history is thin and the current line is young. 2.0.1 was published on 2026-07-15, 2.0.0 on 2026-07-10, and before that the previous release was 1.4.5 on 2024-04-01. The last push to the develop branch was on 2026-07-15, so the repository is not archived and the branch has moved recently. The gap between 1.4.5 and 2.0.0 is over two years, which means a project pinned to the 1.x line has been sitting on an old release for a long time and should expect real changes when it moves. The CHANGELOG file at the repository root is the place to read before that jump; the README does not summarise what changed between 1.4.5 and 2.0.0.

Upgrade cost is otherwise low by construction. Because the recommended integration is copying rpmalloc.h and rpmalloc.c into your tree, an upgrade is a file replacement plus a rebuild. The risk is not the build, it is the behaviour change in an allocator that everything in your process depends on. Pinning to a specific commit and re-running your own tests after each replacement is the only real safeguard, and the benchmark directory in the repository gives you the harness the project itself uses.

## Conclusion

Adopt rpmalloc if you ship C or C++ on Windows, Linux, macOS, iOS or Android, want 16-byte aligned allocations by default, and value a single readable source file over a dependency. Do not adopt it if you need a documented security posture, a stable ABI for plugins, or a project with many years of public issue history behind it. Before committing, verify three things in your own build: that rpmalloc_linker_reference is called when you link the static library and rely on malloc overrides, that rpmalloc_finalize is invoked at exit if you call rpmalloc_initialize yourself, and that your platform is one of the five the README lists.

## FAQ

### Is jemalloc still used?

The README does not discuss jemalloc's adoption. It does list jemalloc among the allocators rpmalloc is compared against in the mimalloc-bench suite referenced by BENCHMARKS.md, and it frames the project's own performance charts as relative comparisons between allocators rather than absolute figures.

### What is TCMalloc?

The README names tcmalloc as one of the popular memory allocators the project believes rpmalloc is faster than, alongside hoard and ptmalloc3. It does not describe how tcmalloc works internally.

### Is alloca faster than malloc?

The README does not compare alloca with malloc. It covers rpmalloc's own allocation path, including the statement that all allocations have a natural 16-byte alignment, and the rptest benchmark that measures throughput against other allocators.

### How does kmalloc work?

The README does not describe kmalloc. It documents rpmalloc's page mapping, which uses an mmap-style virtual memory management API that can be configured at runtime to a custom implementation, with the memory page size also configurable.

## Sources

- [Issues](https://github.com/mjansson/rpmalloc/issues)
- [License: MIT](https://github.com/mjansson/rpmalloc/blob/develop/LICENSE)
- [mjansson/rpmalloc on GitHub](https://github.com/mjansson/rpmalloc)
- [README](https://github.com/mjansson/rpmalloc/blob/develop/README.md)
- [Releases](https://github.com/mjansson/rpmalloc/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mjansson-rpmalloc
