# BetterDB Monitor: the ring buffer is not a history

> A slowlog is a bounded rotating buffer, so it can answer exactly one question, which is what was slow in the last few minutes. BetterDB's argument is that this is the wrong shape for a database you are going to be paged about, and its answer is to copy that data out continuously into storage you can query by time range.

**BetterDB-inc/monitor** — Real-time monitoring, slowlog analysis, and audit trails for Valkey and Redis

- Repository: https://github.com/BetterDB-inc/monitor
- Website: https://betterdb.com
- Stars: 1,304 · Forks: 83
- Language: TypeScript
- License: NOASSERTION
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/betterdb-inc-monitor

## What the database throws away

The positioning line in this README is unusually specific: it persists what Valkey throws away, so you can debug what happened at three in the morning rather than only what is happening now.

That is a real gap and it is worth being precise about why it exists. A slowlog is a fixed-size rotating list. The database writes an entry when a command exceeds a latency threshold, overwrites the oldest when the list is full, and you can only ever ask it about the recent past. Set the threshold high and you miss things; set it low, as the test compose file in this repository does, and you drown. Either way the history is gone.

So the first thing this tool does is copy that out continuously into storage it controls, which turns an ephemeral diagnostic into a queryable time series. The historical analytics feature is described in exactly those terms: slowlogs, command patterns, client activity and latency across any time range, described as the data that used to disappear after a log rotation.

Two newer sources make that history much more useful on Valkey specifically. Command log support is called out as exclusive to Valkey 8.1 and above, and the distinction it draws is the important one: it captures large requests and large replies, not just the slow ones. A slowlog is a latency filter, so a command that returns forty megabytes in two milliseconds never appears in one. That is the class of incident that looks like a memory problem in the database and a payload problem in your application.

And there is live traffic capture, which records real sessions on demand with a tail, filtering, replay and export, cross-referenced against connection history. That is closer to packet capture than to metrics, and it is the feature people reach for when a slowlog has already failed to explain anything.

The rest of the feature set follows the same logic: hot key tracking with rank movement over time, cluster topology and slot statistics, per-thread CPU and I/O metrics, client attribution by name and pattern, and an access control list audit trail kept for compliance and post-incident work. Everything in the list answers a question about the past.

## Two images, because the language model is not a flag

Two container variants are published, both multi-architecture for the two Linux architectures that matter, and the difference between them is not a feature flag.

One is the default image, tagged as the latest release and with a no-ai suffix. It carries every monitoring feature and none of the dependencies for the experimental AI helper. The other adds that helper, which expects you to bring your own Ollama instance, and it is disabled by default through an environment variable.

Most projects would ship one image and let the helper fail to start if you had not configured a model. Splitting the image means the default artefact contains no language model runtime at all, which matters for three separate reasons. Size, because an Ollama client and a model server in one image is a large download for something most deployments will never use. Attack surface, because a tool with network access to your database credentials should not also be a general-purpose model client. And air-gapped installs, which the Kubernetes documentation explicitly covers and which would otherwise need a build of the whole stack.

The build enforces the split from both ends. The install step excludes the proprietary workspace, and a comment in the Dockerfile notes that the development dependencies are included for the build and then stripped again in a cleanup stage for the no-ai variant. That is the difference between a tag that claims to be lighter and one that is.

The layering elsewhere is careful in the same way. The package files are copied before the source so the dependency install is cached, the install uses a frozen lockfile, and the toolchain itself is pinned through a package manager shim with an explicit version rather than whatever the base image ships. The base image is a current Node release, while the CLI's own documented requirement is a floor rather than a match, which is the usual relationship between the two.

## An SSH key directory, and the invariant that makes it safe

The example environment file contains the best piece of security design in this repository, and it is about a monitoring tool rather than about monitoring.

A monitoring tool holds credentials for the database it watches, and sometimes that database is only reachable through a bastion, which means an SSH private key. So this tool can be pointed at an SSH key. That is the most sensitive file it can touch, and the failure mode of a tool that takes a file path is that the path is whatever somebody put in it.

The configuration comment states the rules. There is a directory variable containing the private keys a connection may reference. A connection's private key path must resolve inside that directory. Leave the variable unset and server-side key files are disabled entirely. And inline, pasted keys always work regardless of that setting.

So the invariant is containment, expressed as a rule rather than implied: the tool will not open a key outside the directory you nominated. That is the difference between a feature that is dangerous by default and a feature that is dangerous only if you misconfigure it, and it is documented in the file people copy rather than in a security page nobody reads.

The last clause is the subtle one. Inline keys bypass the directory check entirely, which sounds like an inconsistency and is not. An inline key is held in memory for one connection and never persisted as a file on disk, so there is no path to traverse and no file for a later compromise to read. The two mechanisms have different threat models and the documentation says so.

The general lesson is worth carrying to any tool that accepts a path to a credential: constrain the path, and say in the configuration file what the constraint is.

## Three documented behaviours for one database URL

The rest of the example file is a short security manual, and the comments do the work rather than the values.

There is a flag for connecting to the monitored database over TLS, off by default, with a note to turn it on for managed providers that require encryption. There is a variable for pointing storage at a certificate authority, accepting either a file path or a trusted URL, and a separate flag for connecting over TLS without verifying the chain. That flag's comment explains why you might need it: some managed providers present their own certificate authority and inject a requirement on the connection, so your careful verification settings are overridden by the server's.

And then the hosted edge database, where three different URL shapes get three different behaviours, each explained.

An authenticated scheme URL requires a token. A hosted endpoint over TLS also needs a token, but startup does not enforce it there, and the comment says the consequence plainly: an endpoint with no token starts cleanly and fails on the first query. And a token on an unencrypted URL is rejected at startup, because it would cross the network in cleartext.

That third behaviour is the one worth admiring. The alternative would be to accept it with a warning, which would mean the configuration works in development and silently leaks a credential in production. Failing at startup is the better failure, and it costs nothing except the inconvenience of knowing why.

The failing-on-first-query case is a different trade and it is documented as such, which is the minimum you can ask. A hosted endpoint cannot be checked at boot without a round trip, so the check moves to the first query and the comment says so rather than leaving you to discover it.

Also in that file: a poll interval for the audit trail, and a retention setting measured in days with a note that leaving it unset means keeping everything, plus a daily sweep that deletes older rows.

## Open core, enforced in the build file

This project has a commercial boundary, and the most interesting thing about it is how visible the boundary is.

The licence badge reads as the permissive licence plus commercial. The workspace definition includes a directory named for the proprietary code alongside the application and package directories. There are dedicated scripts to develop, build and start that one workspace, separate from everything else.

And then the part that makes it honest: the container build explicitly filters it out. The install step excludes that workspace by name, so the published open-source image does not contain the code that implements the paid features. The exclusion is in the build recipe rather than in a policy document, which means it is enforced on every build by something that runs in CI.

The feature list is honest in the same way. The features that are paid are labelled as such in the README, and several are described as free during an early-access period. Anomaly detection with automatic baselines and more than twenty detectors is one. Capacity forecasting is another. Inference latency percentile alerts and the semantic cache analysis with an approve-and-reject workflow are others. Key analytics, which adds type, time-to-live and size distributions from live sampling, is a third.

Whether you should read that as a good sign is a matter of taste. Some tools would hide the boundary so you only find out at checkout. This one names the paid features in the same list as the free ones and marks them, which means you can evaluate exactly what you would be paying for. That is a more useful position than a feature list that turns out to be half a product, and it also means the free version has a clear shape: history, capture, hot keys, cluster visibility and the audit trail.

## Six compose files, each testing one thing

The root of this repository has six Compose files, and their names are a test plan.

There is one for a single instance, one for a cluster, one for a Redis cluster specifically, one for testing failover behaviour, one for testing the migration path between the two databases, and one general test file. Each pins different image versions and, more importantly, different server configuration.

Look at what the single-instance service is configured with. The slowlog threshold is set to zero, so every command is logged rather than only the slow ones, which is the only way to test a tool that reads slowlogs without generating an artificial load. The command log thresholds are set to a hundred bytes, so large requests and replies get captured. The maximum lengths are capped so the ring buffers are small and the rotation behaviour is exercised. The failover service turns on the access control log and the latency monitor. A second service runs a different bundle version on a different port, and a third runs Redis rather than Valkey with persistence enabled, so the compatibility path is covered by the same suite.

That is how you test a tool whose input is the database's own diagnostics: you configure the database to emit everything, in a small buffer, and then assert on what your tool did with it.

The build file has the same discipline. Dependencies are installed from a frozen lockfile, the toolchain is pinned explicitly, and the copy order puts package manifests ahead of source so an edit to a file does not invalidate the dependency layer. There is even a comment explaining a subtle dependency chain: the application depends on a memory package that in turn pulls two others, and all three manifests have to be present before the workspace install will resolve.

## The two things that will cost you an afternoon

Two practical notes, both of which are the kind of thing a README buries and a first deployment discovers.

The first is the container-to-host address. The documentation includes a note about connecting to a database on your host machine, and it is worth reading slowly. Inside a container, the name for the host is not the loopback address, because loopback is the container itself. You are told to use a host alias instead, that this works out of the box on the desktop application for macOS and Windows, and that on Linux you have to add a host mapping so the name resolves.

That is the single most common reason a first run of this class of tool appears broken: the dashboard connects, sees nothing, and the database is fine. The starting point is a container with one published port and no database configuration at all:

```bash
docker run -d --name betterdb -p 3001:3001 betterdb/monitor:latest
```

Point the browser at that port and you get the dashboard, which is the fastest way to see the tool working before you point it at anything you care about. The mitigation shipped for the address problem is good and worth naming: there is a one-click connect button in the dashboard that detects the situation and pre-fills the right host. Somebody got burned by this and built the fix into the product.

The second is the Helm example, which sets the database password as a command line value. That value then lives in your shell history, in whatever CI log recorded the command, and in the Helm release history in plain text. It works and it is the documented path, which means people will copy it, and for a long-lived cluster that is a credential you will rotate eventually.

The rest of the Kubernetes path is careful: a namespace created for you, a chart repository with a documented URL, and a port-forward command for local access with an ingress alternative. And the two things people usually forget are called out: persistent history in a separate database, and secrets supplied rather than generated.

## Conclusion

Adopt BetterDB Monitor if you have been paged about something a slowlog could not tell you, and start with the no-ai image variant so you are not carrying a local language model runtime you have not enabled. Configure the SSH key directory before you need it and read the invariant stated in the example file, because that constraint is what stops a key path escaping the directory. Expect the open-core boundary, since the anomaly detection and capacity forecasting you may want are marked Pro and the entitlement workspace is excluded from the open-source image at build time. And pass a Helm database password through a values file rather than a command line flag.

## FAQ

### What does BetterDB Monitor do that Redis or Valkey monitoring does not?

It copies out the data the database discards. Slowlogs are a fixed-size rotating list, so they only describe the recent past; BetterDB persists slowlogs, command patterns, client activity and latency into storage it controls so you can query any time range. On Valkey 8.1 and above it also captures large requests and replies, not just slow commands.

### How do I run BetterDB Monitor?

As a container image, as a Helm chart on Kubernetes, or through npm. The container needs the database host, port and password passed as environment variables and exposes a port you point a browser at. The npm path runs an interactive wizard on first run covering the database connection, the storage backend and server settings, saving the result to a configuration file in your home directory.

### What is the difference between the two BetterDB Monitor image variants?

One contains every monitoring feature without the dependencies for the experimental AI helper; the other adds that helper and expects you to bring your own Ollama instance, with the feature disabled by default through an environment variable. Both are published for the two common Linux architectures, and the build excludes the proprietary workspace and strips development dependencies from the lighter variant.

### How does BetterDB Monitor handle SSH keys for tunnelled database connections?

There is a setting for a directory of SSH private keys that connections may reference, and a connection's key path must resolve inside that directory. Leaving the setting unset disables server-side key files entirely, while inline pasted keys work regardless of the setting.

### Which parts of BetterDB Monitor are paid?

The README labels them: anomaly detection, capacity forecasting, key analytics with size and time-to-live distributions, inference latency percentile alerts, and semantic cache intelligence, several of which are described as free during early access. The free version covers history, live traffic capture, hot key tracking, cluster visibility and the access control audit trail.

### Why can I not connect to my database from the BetterDB Monitor container?

Inside a container, loopback is the container itself rather than your host. Use the host alias for your platform instead, which works out of the box on the desktop application for macOS and Windows, and on Linux add a host mapping to the run command so the name resolves. The dashboard also has a one-click connect button that detects this and pre-fills the right host.

## Sources

- [BetterDB-inc/monitor on GitHub](https://github.com/BetterDB-inc/monitor)
- [Issues](https://github.com/BetterDB-inc/monitor/issues)
- [Project website](https://betterdb.com)
- [README](https://github.com/BetterDB-inc/monitor/blob/master/README.md)
- [Releases](https://github.com/BetterDB-inc/monitor/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/betterdb-inc-monitor
