Trench: a self-hosted Kafka and ClickHouse event pipeline behind a Segment-compatible API
Trench — Open-Source Analytics Infrastructure. A single production-ready Docker image built on ClickHouse, Kafka, and Node.js for tracking events. Easily build product analytics dashboards, LLM RAGs, observability platforms, or any other analytics product.
At a glance
- What is it?
- Trench packages Kafka and ClickHouse into one Docker image and exposes a Segment-compatible tracking API. It suits teams who want raw SQL access to their own event data, and it is not a finished analytics product.
- Who is it for?
- Adopt Trench if you have Docker Compose experience, want event data in your own ClickHouse, and are comfortable writing SQL instead of clicking through a dashboard product. Skip it if you need a maintained client SDK, a query UI, or a dashboard builder out of the box; the newest published release listed here is [email protected] from 2024-12-04.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 177 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Trench solves, and who it is actually for
Most product analytics tools make you choose: a hosted service that owns your event data, or a self-hosted stack you assemble from Kafka, ClickHouse, a collector, and a query layer. Trench takes the second path and collapses it into one Docker image. The README describes it as an event tracking system built on Apache Kafka and ClickHouse that "can handle large event volumes and provides real-time analytics", and the project states it was built to scale the real-time event tracking pipeline at Frigade.
The intended user is a backend or data engineer who wants events landing in a ClickHouse table they control, with the option to run arbitrary SQL against them. The README lists the audience indirectly through its feature list: Segment API compatibility, a single production-ready Docker image, webhooks to forward data to other destinations, and MIT licensing. There is no mention of a hosted dashboard product, a visual funnel builder, or a charting interface. That is the boundary of the project, and it is worth reading the feature list as a scope statement rather than a roadmap.
Two claims in the README deserve separate treatment. The first is compliance: Trench is described as no-cookie, GDPR, and PECR compliant, with users able to access, rectify, or delete their data. Because events are keyed by userId and stored in your own ClickHouse instance, the data residency argument is real, but the compliance claim is about architecture, not about a certification the repository ships. The second is throughput: "process thousands of events per second on a single node". The README states it; nothing in the repository layout lets a reader verify it, and the hardware assumptions behind it are not spelled out beyond a general recommendation of at least 4GB of RAM and 4 CPU cores for production.
Kafka in, ClickHouse out: the data flow Trench exposes
The architecture is visible from the deployment shape rather than from a diagram. The quickstart starts a Trench server that, in the words of the README, "includes a local ClickHouse and Kafka instance", listening on port 4000. So the HTTP layer is the front door, Kafka is the buffer, and ClickHouse is the store.
Events arrive at POST /events as a JSON body with an events array. Each entry carries userId, type, event, and a properties object. The README's example uses type "track" and event "ConnectedAccount", matching the Segment Track, Group, and Identify vocabulary that the feature list claims compatibility with. Authentication is a bearer token, and the README distinguishes a public key from a private key: the public key is used for sending events, the private key for reading them back.
Reads go through GET /events, which accepts an event filter as a query parameter and returns a JSON envelope with results, limit, offset, and total. The README shows limit 1000 and offset 0 in its sample response. That envelope is the shape of the read API, not a general query language. For anything analytical you drop to POST /queries, which takes a queries array of raw SQL strings and runs them against the events table. The README's example is a COUNT(*) filtered by userId.
There is a real consequence to that split. The /events endpoint is a record fetch with pagination; the /queries endpoint is arbitrary SQL against ClickHouse. The first is safe to expose to an application, the second is not, which is why the README's own query example uses the public key while the event read uses the private key. If you deploy this, that key separation is the thing to get right before anything else.
Installing Trench with Docker Compose and sending a first event
The README states the only prerequisite is a system with Docker and Docker Compose installed, and recommends at least 4GB of RAM and 4 CPU cores for a production environment. Clone the repository, move into the app directory, copy the environment file, and bring the stack up. The dev compose file is layered on top of the base one, and the flags force a rebuild and discard anonymous volumes, which matters when you are iterating on configuration.
git clone https://github.com/frigadehq/trench.git
cd trench/apps/trench
cp .env.example .env
docker-compose -f docker-compose.yml -f docker-compose.dev.yml up --build --force-recreate --renew-anon-volumesWhen the containers are up, open http://localhost:4000 in a browser. The README says you should see the message "Trench server is running". The README also notes that the default public and private API keys live in the .env file and that you should edit that file to change configuration options. The key strings shown in the examples are the defaults from that file, not values you generate.
With the server running, send an event using the public key. The payload is an events array; here it is the README's ConnectedAccount example.
curl -i -X POST \
-H "Authorization:Bearer public-d613be4e-di03-4b02-9058-70aa4j04ff28" \
-H "Content-Type:application/json" \
-d \
'{
"events": [
{
"userId": "550e8400-e29b-41d4-a716-446655440000",
"type": "track",
"event": "ConnectedAccount",
"properties": {
"totalAccounts": 4,
"country": "Denmark"
},
}]
}' \
'http://localhost:4000/events'Read it back with the private key and an event filter. The README shows the response returning the event with a uuid, the original properties, a timestamp, and a parsedAt field, wrapped in an envelope with limit 1000, offset 0, and total 1.
curl -i -X GET \
-H "Authorization: Bearer private-d613be4e-di03-4b02-9058-70aa4j04ff28" \
'http://localhost:4000/events?event=ConnectedAccount'For anything beyond record lookup, post SQL to the queries endpoint. The README's example counts events for one userId, and the sample response is a results array containing a count field.
curl -i -X POST \
-H "Authorization:Bearer public-d613be4e-di03-4b02-9058-70aa4j04ff28" \
-H "Content-Type:application/json" \
-d \
'{
"queries": [
"SELECT COUNT(*) FROM events WHERE userId = '550e8400-e29b-41d4-a716-446655440000'"
]
}' \
'http://localhost:4000/queries'If your Kafka brokers need authentication, the README documents environment variables rather than a config file: KAFKA_SSL_ENABLED, KAFKA_SSL_REJECT_UNAUTHORIZED, KAFKA_SSL_CA, KAFKA_SSL_CERT, KAFKA_SSL_KEY, KAFKA_SASL_MECHANISM (plain, scram-sha-256, or scram-sha-512), KAFKA_SASL_USERNAME, and KAFKA_SASL_PASSWORD. The README notes that certificate variables take PEM content directly, and that KAFKA_SSL_REJECT_UNAUTHORIZED defaults to true, so self-signed development certificates require setting it to false.
Where Trench stops short: dashboards, SDK maturity, and operational load
The most important limitation is the one the README implies by omission. Trench is infrastructure, not a product. There is no dashboard builder, no charting interface, and no visual funnel editor described anywhere in the README. The demo video walks through building a basic version of Google Analytics using Trench and Grafana, which tells you exactly what the intended pattern is: you point a separate visualization tool at the data. If your team wants a product analyst to answer questions without writing SQL, Trench is the wrong layer of the stack.
The client SDK situation is the second constraint. The releases listed for this repository are [email protected] and [email protected], both dated 2024-12-04, plus [email protected] from the same day. The 0.0.x versioning is a signal about API stability, and the release cadence visible here is not a stream of frequent updates. The last push to the repository was on 2026-04-06, so the repository is not archived, but the published package releases are considerably older than that push. Anyone adopting trench-js should pin a version and read the changelog rather than assuming the client surface is settled.
Operationally, the single-image claim is about packaging, not about eliminating moving parts. You are still running Kafka and ClickHouse, with the memory and disk footprint that implies. The README's own guidance of at least 4GB of RAM and 4 CPU cores for production is a floor for a small deployment, and it says nothing about retention, disk growth, or ClickHouse table maintenance. The self-hosted and cloud split in the README is worth reading as a statement about who should run this: Trench Cloud is described as fully managed with autoscaling and 99.99% SLAs. If you do not want to own a Kafka and ClickHouse cluster, the open-source image does not remove that responsibility, it just removes the assembly work.
Trench compared with Plausible and Matomo
The repository topics list matomo, matomo-analytics, plausible-analytics, and posthog alongside clickhouse and kafka, so the project itself frames those as the adjacent tools. The difference in approach is worth stating plainly.
Plausible and Matomo are analytics products with a defined data model and a user interface. You install them, they collect pageviews and events, and they present charts. You can query their underlying storage, but the schema and the reporting layer belong to the product, and extending them means working within their plugin or API surface. Trench inverts this. It gives you an ingestion endpoint and a SQL endpoint, and the reporting layer is whatever you connect. That is more work up front and considerably more flexibility later, particularly if your events are not pageviews at all.
The README's feature list points at use cases that the pageview-oriented tools handle poorly: product analytics dashboards, LLM RAGs, and observability platforms. Those share a property, which is that the interesting queries are custom. A RAG pipeline wants to log retrieval and generation events and then ask questions about latency or hit rate. A pageview tool will store that data if you push it hard enough, but its query model is not built for it. Trench's /queries endpoint is. The trade is that you write the SQL, and you own the Grafana or custom front end that renders it.
Licence, upgrade cost, and what the MIT terms do not cover
The repository is MIT licensed, and the root package.json carries "license": "MIT" with the author listed as Frigade Inc. MIT is permissive: you can use, modify, and redistribute the code, including commercially, provided the copyright notice and permission notice are retained. The licence file sits at the repository root as LICENSE.
Two things the licence does not settle, and this is not legal advice. First, running Trench means running ClickHouse and Kafka, which carry their own licences and their own operational terms; the MIT grant on Trench says nothing about those components. Second, the README describes a separate Trench Cloud offering with autoscaling and SLAs. The relationship between the open-source image and the hosted service is not spelled out in the README, so if you are evaluating this for a commercial deployment, check what, if anything, is reserved for the cloud product.
On upgrade cost, the repository is a pnpm and Turborepo monorepo. The root package.json declares packageManager [email protected] and workspaces for apps/* and packages/*, with scripts routed through turbo: build, test, dev, lint, clean, format. Versioning and publishing run through Changesets, with changeset, version-packages, and release scripts. That matters for anyone consuming trench-js or analytics-plugin-trench: releases are cut through the Changesets flow, so the changelog entries in the .changeset directory are the record of what changed between published versions. The devDependency list is notably old, with eslint ^7.32.0 and prettier ^2.5.1 against turbo ^1.10.14. That combination is workable, but it means a contributor picking up the repository is working with a toolchain from a different era than a fresh 2026 project, and a dependency refresh is a nontrivial change rather than a version bump.
Editorial conclusion
Adopt Trench if you have Docker Compose experience, want event data in your own ClickHouse, and are comfortable writing SQL instead of clicking through a dashboard product. Skip it if you need a maintained client SDK, a query UI, or a dashboard builder out of the box; the newest published release listed here is [email protected] from 2024-12-04. Before committing, clone the repository and confirm that the .env.example keys, the default API key format, and the docker-compose.dev.yml service set match what your deployment expects.
Frequently asked questions
What is Trench and what does it do?
Trench is an open-source event tracking system built on Apache Kafka and ClickHouse, MIT licensed and distributed as a single production-ready Docker image. It accepts events over a Segment-compatible API and lets you query them with raw SQL.
How do I install and run Trench self-hosted?
The README's quickstart requires Docker and Docker Compose, then clones the repository, copies .env.example to .env inside apps/trench, and starts the stack with docker-compose.yml plus docker-compose.dev.yml. The server comes up on http://localhost:4000 with a local ClickHouse and Kafka instance.
Does Trench include a dashboard or a user interface?
The README does not describe a dashboard or charting interface. Its demo video shows building a basic version of Google Analytics using Trench and Grafana, which means visualization is handled by a separate tool pointed at the data.
Can Trench connect to a Kafka cluster that requires authentication?
Yes. The README documents optional environment variables for SASL and SSL, including KAFKA_SSL_ENABLED, KAFKA_SSL_REJECT_UNAUTHORIZED, KAFKA_SSL_CA, KAFKA_SSL_CERT, KAFKA_SSL_KEY, KAFKA_SASL_MECHANISM, KAFKA_SASL_USERNAME, and KAFKA_SASL_PASSWORD.
What licence is Trench released under?
Trench is MIT licensed, with the LICENSE file at the repository root and "license": "MIT" in the root package.json, authored by Frigade Inc. The permissive terms apply to Trench itself, not to the ClickHouse and Kafka components it runs alongside.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/frigadehq-trench)