soundcloud/roshi: A CRDT-Based Distributed Event Store Over Redis
Roshi is a large-scale CRDT set implementation for timestamped events.
At a glance
- What is it?
- Roshi is a Go library and server for storing timestamped events across distributed Redis instances, using a Last Writer Wins element-set CRDT to converge replicated data without requiring consensus, originally built for the SoundCloud activity stream.
- Who is it for?
- Roshi suits engineering teams who need a high-performance, eventually-consistent index of timestamped events across a distributed Redis infrastructure and who are comfortable operating a stateless Go server tier. It is not the right tool for use cases that require strong consistency, transactional writes, or richer query semantics than newest-first range scans.
- Can I use it commercially?
- Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem Roshi Solves
Building a high-performance activity feed or event timeline at scale requires storing and querying large sets of timestamped items efficiently. A naive approach using Redis sorted sets gives you ordered storage by score, but it provides no mechanism for replicating that data across multiple Redis instances with automatic conflict resolution when replicas diverge due to network partitions or node failures.
Roshi addresses this by implementing a Last Writer Wins element-set CRDT on top of Redis. The originating use case, as noted in the repository's README, is the SoundCloud stream: a feed of events ordered by external timestamp where the system needs to handle insertions and deletions from multiple sources without requiring a global coordinator. The README describes Roshi as a high-performance index designed to sit in the critical request path of an application or service.
The project is for backend engineers who are already running Redis infrastructure and need a replicated, fault-tolerant, eventually consistent layer for ordered event data. It is not a general-purpose database; it offers exactly three operations and stores each key as an ordered set of timestamped values.
The LWW-Element-Set Mechanism
A CRDT (conflict-free replicated data type) is a data structure where applying the same set of operations in any order produces the same final state. Roshi implements the Last Writer Wins element-set variant, which uses two internal sets per key: an add set and a remove set. Each entry is a tuple of a value and a timestamp.
To add an element, the system places the (value, timestamp) tuple into the add set. To remove an element, it places the same tuple into the remove set. An element is considered present in the logical set when its most recent entry is in the add set rather than the remove set.
Roshi extends this with inline garbage collection. When writing an element that already exists in either the add or remove set, the system compares the timestamps. If the incoming timestamp is higher than the stored one, the existing entry is removed and the new (value, timestamp) pair is written to the add set. If the incoming timestamp is lower or equal, the write is a no-op. This ensures each element appears in exactly one of the two sets at any time, keeping storage compact.
Because all operations are associative, commutative, and idempotent, Roshi replicas can receive operations in different orders and converge to the same state without coordination.
Cluster, Farm, and Read-Repair Architecture
Roshi organizes Redis storage into two levels. A cluster is a set of sharded Redis instances that together hold one copy of the dataset; each key is assigned to one shard within a cluster. A farm is a collection of clusters, typically three, that each hold an independent copy of the entire dataset.
Writes are sent to all clusters in the farm simultaneously. The write operation returns success when a user-configurable number of clusters confirm the write. Clusters that are slow or offline may miss the write, but because operations are idempotent, they can accept the write later without corrupting the state.
Reads query one or more clusters depending on the configured read strategy. When multiple clusters return results that disagree, Roshi returns the union of those sets to the client and triggers a read-repair operation in the background. Read-repair sends the missing writes to each replica that was behind, lazily converging the clusters without requiring a separate reconciliation process.
Roshi instances themselves are stateless; they hold only transient per-request state. If a Roshi instance crashes, the Redis data is unaffected. Clients reconnect to another Roshi instance and retry their operation. Unresolved in-flight read-repairs are lost, but the same repair will be triggered again on the next read that observes the inconsistency.
Building and Running the Roshi Server
Roshi is a Go module. The go.mod declares the module path and minimum Go version:
module github.com/soundcloud/roshi
go 1.14The repository contains two runnable components in roshi-server/ and roshi-walker/. The roshi-server package exposes the Insert, Delete, and Select API over HTTP. The roshi-walker package performs periodic background scans of the dataset.
The repository's README does not include explicit build or installation commands. Standard Go build tooling applies: clone the repository and run go build in the roshi-server/ directory to produce a binary, then configure it with the addresses of your Redis shards. The README links to the farm package documentation for details on read strategies, replication factors, and the read-repair pipeline configuration.
The three API methods are Insert(key, timestamp, value), Delete(key, timestamp, value), and Select(key, offset, limit) which returns a slice of TimestampValue pairs ordered newest-first. The API is minimal by design; there is no filtering, secondary indexing, or bulk operation interface beyond what these three methods provide.
Where Roshi Is the Wrong Tool
Roshi is eventually consistent, not strongly consistent. If your application cannot tolerate reading a state that is slightly behind because a read-repair has not yet completed, Roshi is not suitable. The README is explicit about this: it offers partition tolerance and high availability at the cost of eventual consistency.
The query surface is narrow. Select returns a range of items by offset and limit in newest-first order. There is no time-range filtering, no secondary index, no aggregation, and no key-scanning interface. If you need to answer questions like 'all events of type X between timestamps A and B', Roshi cannot answer that query directly; you would need a separate index or a different storage system.
The CRDT semantics also mean that deletions win over additions when the deletion timestamp is higher. If your event stream requires exactly-once delivery guarantees or event ordering guarantees stronger than last-write-wins, Roshi's semantics will not satisfy those requirements. The system is designed for scenarios where the most recently timestamped operation is always the authoritative one.
Comparison with Redis Sorted Sets and Project Status
Redis sorted sets provide an ordered data structure keyed by score, which can represent timestamped events when you use the event timestamp as the score. The difference from Roshi is that Redis sorted sets are single-node or primary-replica data structures with no built-in CRDT semantics. If two nodes write conflicting scores for the same member, the result depends on which write arrives last at the primary; there is no automatic convergence across independent replicas.
Roshi adds CRDT convergence semantics, a multi-cluster replication model, and the inline garbage collection that keeps the add and remove sets compact. The trade-off is added operational complexity: you must run the Roshi server tier in addition to Redis, and you must size a farm of at least three clusters to benefit from fault tolerance.
The Go module requires go 1.14 and the project carries a BSD-2-Clause license. The last push to the repository was on September 11, 2026. The repository is not archived.
Editorial conclusion
Roshi suits engineering teams who need a high-performance, eventually-consistent index of timestamped events across a distributed Redis infrastructure and who are comfortable operating a stateless Go server tier. It is not the right tool for use cases that require strong consistency, transactional writes, or richer query semantics than newest-first range scans. Before adopting it, verify that Go 1.14 or newer is in your toolchain and that your Redis cluster topology fits the cluster-and-farm model Roshi expects: the design requires at least three independent clusters to provide meaningful fault tolerance.
Frequently asked questions
What Redis topology does roshi require?
Roshi organizes Redis into clusters and farms. A cluster is a set of sharded Redis instances holding one copy of the data. A farm is a group of clusters, typically three, providing replication. The README recommends at least three clusters in a farm for meaningful fault tolerance.
Does roshi guarantee strong consistency across replicas?
No. Roshi is eventually consistent. It is partition tolerant and highly available, using read-repair to lazily converge replicas when they disagree. Applications that require strong consistency need a different storage system.
What operations does roshi's API expose?
Roshi exposes exactly three operations: Insert(key, timestamp, value), Delete(key, timestamp, value), and Select(key, offset, limit). Select returns results in newest-first order. There is no filtering, aggregation, or secondary index support.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/soundcloud-roshi)