rudderlabs/rudder-server: a self-hosted, Segment-compatible CDP you run yourself
Privacy and Security focused Segment-alternative, in Golang and React
At a glance
- What is it?
- RudderStack's open source server collects events, routes them to warehouses and tools, and keeps the pipeline inside your own infrastructure. Here is what the repository actually shows, how to run it with Docker, and where it stops being the right choice.
- Who is it for?
- Adopt rudder-server if you want Segment-compatible event collection that terminates in your own PostgreSQL and warehouse, and if you can run the transformer container alongside it. Do not adopt it if you need a purely managed pipeline, or if you expect the GitHub repository and the Docker images to move in lockstep, since the README says the images receive bug fixes far more often.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem rudder-server solves, and who ends up running it
Most product analytics stacks start the same way: an SDK fires an event, a vendor's servers receive it, and your warehouse gets a copy some hours later, if at all. RudderStack inverts that. The README describes the project as "the leading open source Customer Data Platform (CDP)" whose pipelines "collect data from every application, website and SaaS platform, then activate it in your warehouse and business tools." The repository you are looking at, rudderlabs/rudder-server, is the piece that does the collecting and routing.
The intended audience is stated plainly in the README's own tagline: "The Customer Data Platform for Developers." That is not marketing noise. The architecture section says the backend is written in Go with a React UI, and that it is "an independent, stand-alone system with a dependency only on the database (PostgreSQL)." A team that already runs Postgres and a data warehouse can stand this up without adopting a new managed service.
The privacy angle is the reason many teams arrive here. The README frames it as collecting and storing customer data "without sending everything to a third-party vendor," with control over "what data to forward to which analytical tool." If your constraint is that raw event data must not leave infrastructure you control, that constraint is the whole product decision, and the rest of the feature list matters less.
How events move through the Go backend
The repository layout is the clearest description of the data flow. Top-level directories include gateway/, processor/, router/, warehouse/, schema-forwarder/, archiver/, regulation-worker/ and jobsdb/. Read together, they describe a pipeline rather than a monolith: something receives events, something processes and transforms them, something routes them outward, and something loads them into a warehouse.
jobsdb/ is the interesting one. RudderStack's retry and delivery guarantees depend on persisting jobs in PostgreSQL rather than holding them in memory, which is why the README can claim a system that "ensures that your data will be delivered even in the event of network partitions or destinations downtime." That claim is only as good as the database behind it. The docker-compose.yml sets JOBS_DB_HOST=db for the backend service, which is the direct evidence that the job store is the same Postgres the stack starts.
The transformer is a separate service with its own image, rudderstack/rudder-transformer:latest, exposed on port 9090. That split matters operationally: JavaScript-based event transformations, which the README calls a "JavaScript-based event transformation framework," do not execute inside the Go binary. If the transformer container is unhealthy, transformations stop even though the rest of the server is up.
The go.mod file confirms the destination surface is broad and compiled in rather than bolted on as plugins. Direct dependencies include cloud.google.com/go/bigquery, github.com/ClickHouse/clickhouse-go, github.com/apache/pulsar-client-go, and a set of AWS SDK v2 clients covering s3, firehose, kinesis, glue, lambda, eventbridge and personalizeevents. Warehouse and streaming targets are first-class Go dependencies, not external adapters.
Installing rudder-server with Docker and sending a first event
The README points to three setup paths: Docker, Kubernetes and a developer machine setup, each with a documentation link. It also carries a warning worth repeating: for production it "STRONGLY recommend[s] using our Kubernetes Helm charts," and notes that Docker images are updated with bug fixes much more frequently than the GitHub repository. If you deploy from the repository's docker-compose.yml, you are deploying the source tree, not the freshest image.
The compose file in the repository defines the services you need. The db service runs postgres:15-alpine and maps host port 6432 to container 5432, the backend service builds from the local Dockerfile and listens on 8080, and the transformer runs on 9090. Two more services, minio and etcd, sit behind the storage and multi-tenant profiles and do not start by default.
Start the default stack from the repository root:
docker compose upThe backend waits for Postgres using the entrypoint sh -c '/wait-for db:5432 -- ./rudder-server', so the first start takes longer than subsequent ones. When it settles, the UI is reachable on port 8080.
Configuration comes from build/docker.env, which every service loads via env_file. The compose file also shows how to mount a workspace configuration if you keep it as a file rather than in the UI:
# Uncomment the following lines to mount workspaceConfig file
# volumes:
# - <absolute_path_to_workspace_config>:/etc/rudderstack/workspaceConfig.jsonOnce the UI is up, the README's Step 2 is to "send test events" to verify the setup, with a link to the sending-test-events documentation. The image also ships helper scripts: the Dockerfile copies scripts/generate-event and scripts/batch.json into the final image, so the container carries the tooling for generating a test payload without you writing one.
If you would rather build from source, the Dockerfile shows the path it takes: go mod download, then make build with version metadata injected through LDFLAGS, then separate builds of cmd/devtool and cmd/rudder-cli. The rudder-cli binary is installed to /usr/bin/rudder-cli in the runtime image.
Where rudder-server is the wrong tool
The strongest argument against self-hosting here is the one the README makes for you. It recommends Kubernetes Helm charts for production and warns that Docker images move faster than the repository. That is a maintenance signal, not a bug: you are taking on the operational surface of a system whose own documentation steers production users toward an orchestrated deployment. A team without Kubernetes experience, or without someone who will own the Postgres behind jobsdb, should treat the docker-compose path as an evaluation environment rather than a production plan.
The dependency on PostgreSQL is absolute. The architecture section says the system depends "only on the database (PostgreSQL)," which is a strength until that database becomes the bottleneck. Every event that needs retrying is a row somewhere. Sizing, backups and failover for that instance are your responsibility, and the README does not discuss them.
The licence is the other boundary. The repository's own badge identifies the licence as ELv2, while the repository metadata reports the licence as NOASSERTION. Those two signals disagree, and the LICENSE file in the repository root is the only thing that settles it. Anyone planning to redistribute the software, or to offer it as a hosted service, needs to read that file rather than the badge. This is not legal advice; it is a pointer to the one document that matters.
Finally, consider the case where you do not actually need a CDP. If you have one destination and one event type, a queue and a small loader will do. rudder-server earns its place when you have many sources, many destinations, and a requirement that the routing rules live somewhere you control.
RudderStack Cloud versus running the server yourself
The most direct alternative is RudderStack's own hosted offering. The README opens by pointing at RudderStack Cloud and a free tier, and the Get started section says the easiest way to experience RudderStack is to sign up for it. That is the same event pipeline, operated by the vendor.
The difference is where the data plane lives. With the cloud tier, the vendor runs the gateway, processor and router, and you configure sources and destinations in their UI. With rudder-server, those components run in your environment against your Postgres, and the README's privacy claim holds literally: nothing is forwarded to a tool you did not configure. You also absorb the upgrade cadence, the transformer container, and the database.
A second alternative is Segment itself, and the README addresses it directly: RudderStack is "fully compatible with the Segment API," so "you don't need to change your app if you are using Segment; just integrate the RudderStack SDKs into your app." The practical difference is not the wire format but the commercial model and the data path. Segment receives your events; rudder-server, self-hosted, does not.
The README's pricing argument is worth quoting because it is the project's own framing: "Event volume-based pricing of most of the commercial systems is broken. With RudderStack Open Source, you can collect as much data as possible without worrying about overrunning your event budgets." That is the trade in one line. You pay in operational work instead of per-event fees.
Maintenance, releases and what an upgrade actually costs
This is a live repository. The last push was on 2026-09-22, and the most recent release listed is v1.88.0, tagged the same day, preceded by v1.88.0-rc.1 on 2026-09-21 and v1.87.1 on 2026-09-18. Release candidates appear before the final tag, so a team that wants to test ahead of a cut has something to test.
The upgrade cost is not the binary. It is the schema. The repository has a top-level sql/ directory and a schema-forwarder/ component, which together suggest that moving between versions can involve database migrations rather than a container swap. The README does not document a rollback procedure for those migrations, and it does not describe version compatibility between the server and the transformer image. That gap is the thing to resolve before you schedule an upgrade window.
The build itself is pinned. go.mod declares go 1.27.1, and the Dockerfile takes GO_VERSION=1.27.1 with a matching GO_VERSION_SHA256, plus ALPINE_VERSION=3.24 with its own digest. A Makefile comment says GO_VERSION "is updated automatically to match go.mod." Building from source therefore gives you reproducible base images, at the cost of tracking Go toolchain releases.
On licensing, the badge says ELv2 and the repository metadata says NOASSERTION. Read LICENSE in the repository root before you plan anything that involves distributing the software or exposing it to third parties. The README itself makes no licensing claims beyond the badge.
Editorial conclusion
Adopt rudder-server if you want Segment-compatible event collection that terminates in your own PostgreSQL and warehouse, and if you can run the transformer container alongside it. Do not adopt it if you need a purely managed pipeline, or if you expect the GitHub repository and the Docker images to move in lockstep, since the README says the images receive bug fixes far more often. Before committing, verify three things: that your Segment SDK calls survive a change of write key and endpoint, that you have sized the PostgreSQL instance behind JOBS_DB_HOST, and that the ELv2 licence terms in the LICENSE file fit how you intend to distribute or host the software. The decisive test is whether events land in your warehouse with the schema you expect after you run the docker-compose stack and send one test event.
Frequently asked questions
How much does RudderStack cost?
The README does not publish pricing. It points to RudderStack Cloud, including a free tier, and argues that self-hosting avoids the event volume-based pricing of commercial systems, since you can collect as much data as your own infrastructure allows.
Who owns RudderStack?
The README does not name an owner or parent company. It identifies RudderStack as an open source project hosted at github.com/rudderlabs/rudder-server and links to rudderstack.com as the project website, with a Slack community, blog and Twitter account.
What is rudder analytics used for?
The README describes pipelines that collect data from applications, websites and SaaS platforms and then activate it in a warehouse and business tools. Warehouse destinations are treated as first-class, with near real-time sync, and over 90 tool and warehouse destinations are listed as supported.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/rudderlabs-rudder-server)