# Tailslayer: hedged reads across DRAM channels to cut refresh-induced tail latency

> Tailslayer is a C++ header library that replicates a value across independent DRAM channels and reads whichever replica answers first, targeting the latency spikes caused by DRAM refresh. It is aimed at latency-sensitive engineers, not general application developers.

**LaurieWired/tailslayer** — Library for reducing tail latency in RAM reads

- Repository: https://github.com/LaurieWired/tailslayer
- Stars: 2,846 · Forks: 164
- Language: C++
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/lauriewired-tailslayer

## The problem Tailslayer targets: refresh stalls, not average latency

Most latency work optimises the mean. Tailslayer is aimed at the opposite end of the distribution. The README states that it reduces tail latency in RAM reads caused by DRAM refresh stalls. DRAM has to be refreshed periodically, and while a bank is being refreshed a read that lands on it waits. That wait is rare relative to the total number of reads, which is exactly why it shows up in the tail and not in the average.

The intended audience is narrow. This is not a general-purpose memory allocator or a caching layer. It is a library for code where a single slow read is expensive, where the developer controls thread placement, and where the machine has multiple independent DRAM channels. If your service is bound by network round trips or by lock contention, Tailslayer does not address your bottleneck. The README also describes the mechanism as relying on channel scrambling offsets that it explicitly labels undocumented, which should tell you the project is written for people comfortable working close to the memory subsystem.

## How hedged reads across channels work

The mechanism is replication plus racing. According to the README, Tailslayer replicates data across multiple independent DRAM channels with uncorrelated refresh schedules. Because the channels refresh on schedules that do not line up, a read that stalls on one channel can still complete on another. When a request arrives, the library issues hedged reads across all replicas and lets the work proceed on whichever result responds first.

The README notes that each insert copies the element N times, where N is the number of replicas, and that the address calculation happens on the backend. That is the trade: you spend memory and some insert cost to buy down the tail. The library acts, in the README's phrasing, as a hedged vector that uses logical indices, so callers keep working with ordinary indices while the replication stays behind the interface.

Thread placement is part of the design, not an afterthought. The README states that each replica is pinned to a separate core and will spin on that core according to the signal function until the read happens. That is a real resource commitment: replicas cost cores, and those cores are busy-waiting rather than sleeping. The README also notes the library currently works with two channels, with full N-way usage available in the benchmark under discovery/benchmark.

## Installing Tailslayer and running the example

There is no package manager step. The README says to copy include/tailslayer into your project and include the header. The repository ships a Makefile that builds the example with g++ and C++17.

Build and run the example from the repository root:

```bash
make
./tailslayer_example
```

The Makefile compiles tailslayer_example.cpp with -O3 -std=c++17 -pthread -D_GNU_SOURCE -Iinclude, so the include path already points at include/. You should see the binary build without errors and then run.

To use the library, provide the value type and two functions as template parameters. The signal function waits for your external signal and returns the index to read; the read fires immediately. The final work function receives the value right after it is read.

```cpp
#include <tailslayer/hedged_reader.hpp>

[[gnu::always_inline]] inline std::size_t my_signal() {
    return index_to_read;
}

template <typename T>
[[gnu::always_inline]] inline void my_work(T val) {
    // Use the value
}
```

Wiring it together means pinning to a core, constructing the reader, inserting values, and starting the workers. The README's example uses tailslayer::pin_to_core(tailslayer::CORE_MAIN), then HedgedReader<T, my_signal, my_work<T>>, then insert and start_workers. Arguments reach either function through tailslayer::ArgList, and the constructor optionally accepts a channel offset, a channel bit, and the number of replicas.

## The benchmark and refresh probe in discovery/

The discovery/ directory is where the project characterises the hardware behaviour it depends on. The README lists discovery/benchmark/ as a channel-hedged read benchmark and discovery/trefi_probe.c as a spike timing probe for measuring the refresh cycle.

Running the benchmark takes a channel bit argument:

```bash
cd discovery/benchmark
make
sudo chrt -f 99 ./hedged_read_cpp --all --channel-bit 8
```

The README gives --channel-bit 8 in this example. sudo and chrt -f 99 put the process in the real-time scheduling class, which is consistent with a benchmark that needs to avoid being preempted while measuring microsecond-scale stalls. If you run this on your own hardware, the channel bit is the value you need to confirm for your platform; the README presents the benchmark as the way to exercise full N-way usage, since the library itself currently supports two channels.

## Where Tailslayer is the wrong tool

The most important limitation is stated by the README itself: the channel scrambling offsets are undocumented. That means the project is building on behaviour that is not part of a published interface and could differ across platforms or firmware. The README claims it works on AMD, Intel, and Graviton, but the offsets and the channel bit are things you should verify on your target hardware rather than assume.

Second, the core cost is real. Each replica is pinned to a separate core and spins there until the read happens. On a machine with few cores, or in a container with a CPU quota, dedicating cores to spinning replicas is likely to cost more than the tail latency it removes. This is not a library for oversubscribed environments.

Third, the copy amplification is explicit. Each insert copies the element N times. If your values are large or your insert rate is high, the memory bandwidth and write cost can outweigh the read tail benefit. And the library currently works with two channels, so the N-way path lives in the benchmark rather than in the shipped header. A workload whose latency is dominated by anything other than DRAM refresh stalls gets nothing from this.

## How it differs from a general-purpose memory library

The obvious comparison is to a memory allocator such as jemalloc or a cache library. Those change how memory is laid out, reused or reclaimed, and they are designed to be dropped into an existing program with little or no change to the calling code. Tailslayer does not manage allocation at all. It replicates a specific value across channels and races the reads, and it requires you to supply the signal and work functions as template parameters and to pin threads to cores.

The difference in approach matters for adoption. An allocator swap is a build-flag decision. Tailslayer is a code-structure decision: you restructure the read path around a signal function and a work function, and you accept a spin-waiting core per replica. The README's own framing, calling the library a hedged vector that uses logical indices, makes the scope clear. If you want a drop-in improvement to allocation behaviour, look elsewhere. If you have already measured refresh stalls in your tail and you control the machine, this is the narrower tool that matches.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-04-11. No releases are listed, and no version tags appear in the README. That means there is no published versioning scheme to upgrade against; the README's install instruction is to copy include/tailslayer into your project, which makes the header itself the unit of upgrade. You will be tracking the main branch by hand rather than pulling a tagged release.

The README also notes that the library currently works with two channels and that updates are to come, with full N-way usage available in the benchmark. Treat that as a statement about the shipped header rather than a roadmap commitment. If your design assumes N-way replication in the library, that assumption is not supported by the README today.

The licence is Apache-2.0, per the repository's LICENSE file and the badge in the README. Apache-2.0 is a permissive licence with an explicit patent grant, which is generally straightforward for commercial use, but the header depends on undocumented channel scrambling offsets rather than on licensed technology. That is a technical and portability risk, not a licensing one. Nothing here is legal advice; read the LICENSE file for the actual terms.

## Conclusion

Tailslayer suits engineers who already know their workload is dominated by DRAM refresh stalls and who can pin threads to dedicated cores, because the library spins on a core per replica and depends on channel scrambling offsets that the README itself flags as undocumented. It is the wrong choice for ordinary application code, for anything that cannot spare a core per replica, or for machines where the channel offset and channel bit have not been verified on that specific hardware. Before adopting it, build the example with make, run the discovery benchmark with the --channel-bit value that matches your machine, and confirm the refresh spike timing with discovery/trefi_probe.c on the hardware you actually deploy on.

## FAQ

### What is meant by tail latency, in the context of Tailslayer?

It is the slow end of the latency distribution rather than the average. Tailslayer targets tail latency in RAM reads caused by DRAM refresh stalls, which are rare enough not to move the mean but slow enough to be visible at the tail.

### What does memory latency mean for a library like Tailslayer?

It is the delay between issuing a read and getting the value back. Tailslayer's approach is to replicate data across independent DRAM channels with uncorrelated refresh schedules and issue hedged reads across all replicas, so the work proceeds on whichever result responds first.

### How do I install Tailslayer?

There is no package manager step. The README says to copy include/tailslayer into your project and include <tailslayer/hedged_reader.hpp>, and the repository's Makefile builds the bundled example with g++ and C++17.

### Does Tailslayer work with more than two DRAM channels?

The README states the library currently works with two channels and that full N-way usage is available in the benchmark under discovery/benchmark.

### What does Tailslayer cost in memory and CPU?

Each insert copies the element N times, where N is the number of replicas, so memory use scales with the replica count. Each replica is also pinned to a separate core and spins on that core until the read happens.

## Sources

- [Issues](https://github.com/LaurieWired/tailslayer/issues)
- [LaurieWired/tailslayer on GitHub](https://github.com/LaurieWired/tailslayer)
- [License: Apache-2.0](https://github.com/LaurieWired/tailslayer/blob/main/LICENSE)
- [README](https://github.com/LaurieWired/tailslayer/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lauriewired-tailslayer
