# counterscale: analytics where the query window is a vendor limit

> The most important number in this project is not a feature count, it is ninety days. Cloudflare's Analytics Engine, which stores the hot data, caps retention at ninety days as of February 2025, so anything older has to live somewhere else, and the answer here is a bucket of Apache Arrow files in R2, written by default and disableable from the CLI.

**benvinegar/counterscale** — Scalable web analytics you run yourself on Cloudflare

- Repository: https://github.com/benvinegar/counterscale
- Website: https://counterscale.dev
- Stars: 2,155 · Forks: 133
- Language: TypeScript
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/benvinegar-counterscale

## The store has a ninety-day ceiling

The limitations section is the first substantive thing in this README, and it is worth reading before the installation instructions.

Counterscale is powered primarily by Cloudflare Workers and by Workers Analytics Engine, which is the piece that stores queryable events. As of February 2025, that engine has a maximum ninety days of retention, so the dashboard can only show the last ninety days of recorded data. That number is not the author's choice and not a setting; it is a limit of a product that is still described as a beta.

The response to it is the interesting architectural decision. Counterscale provides long-term storage of your data in an R2 bucket using Apache Arrow files. It is enabled by default and can be disabled through the command line tool.

So the system has two stores with different characteristics. The hot one answers questions about last Tuesday quickly and forgets anything older. The archive keeps everything as columnar files in object storage, and columnar is the right format for the job: a dashboard that groups by page and counts rows is exactly the query shape a columnar file is good at, so the archive is queryable rather than a pile of JSON you would have to download and process yourself.

That choice has a consequence for anyone extending this. A query about recent data and a query about last year's data are not the same code path, because one hits an engine with a ninety-day memory and the other hits files in a bucket. The README does not describe how the two are unified for the dashboard, which means either the dashboard simply does not show older data, or there is a path worth reading in the source before you rely on year-over-year comparisons.

What it does mean is that the ninety-day limit is not a limitation of this project so much as a limitation it inherited and worked around, and the workaround is enabled by default, which is the part most people would get wrong by assuming it was opt-in.

## Enabling a beta, and the Worker you must create first

The preparation steps describe a dependency chain that is easy to get wrong, and the README handles it well.

You need a Workers subdomain, and then you need Analytics Engine switched on for the account. The instruction is to go to Storage and Databases, then Analytics Engine, and click Enable. Then it tells you something you would not guess: a Create Dataset dialog pops up, and you can ignore it and exit. That is a dead end the README closes for you.

Then the real trap. If this is your first time using Workers, you cannot enable Analytics Engine at all until you have created a Worker, which means navigating to Workers and Pages, clicking Create Worker, and creating a Hello World one. The name does not matter and you can delete it afterwards, which the README says explicitly, so it is a workaround rather than a component you need to keep.

That ordering is imposed by Cloudflare's control panel rather than by this project, and it is the kind of dependency that costs an afternoon when you hit it unprepared.

The API token comes next, and it needs account-level analytics permissions at a minimum. There is a warning about it that deserves repeating: copy it somewhere safe before you close the window, because if you close it you cannot get the token back and have to create another one. That is a one-way action in a dashboard, and it is the step most likely to make you start the whole sequence over.

So the order is: subdomain, then a throwaway Worker, then enable the engine, then the token. Following it in that order means the token is the last thing you create and the first thing you need at the installer.

## The dashboard password is one keystroke

The install step is two commands and a series of prompts, and one of those prompts decides whether your traffic data is public.

```bash
npx wrangler login
```

authorises the Cloudflare command line tool, and then:

```bash
npx @counterscale/cli@latest install
```

runs the installer. You are asked for the API token you just created, and then you are asked whether you want to protect the dashboard with a password. The answer yes means you set a password that is required to reach the dashboard, and the README recommends it for public deployments. The answer no means the dashboard is publicly accessible without authentication.

That is the whole security model of the deployment, and it is one question with two answers. The deployment target is a subdomain on Cloudflare's own workers domain, which is public by construction, so a dashboard with no password is a URL that anybody who guesses or crawls can read. Analytics is not user data, but a year of traffic to a site is still a map of when your visitors arrive and what they read, and it is the kind of thing you would rather not publish by accident.

The choice being a prompt rather than a documented default is defensible, since self-hosters know what they are doing and the installer cannot tell a public deployment from a local experiment. It is also the step most likely to be answered with muscle memory rather than thought, because the prompt appears in the middle of a sequence where every other answer is a fact you just looked up.

One more small operational note from the same section: a first deployment can take a few minutes before the subdomain is live, so an immediate failed request is not necessarily a broken install.

## Three ways to report a pageview

Once it is deployed, something has to send events, and there are three documented routes with different trade-offs.

The first is a script served from your own deployment, which you paste into your pages. It is a script element with a site identifier and a defer attribute, which is the correct shape for an analytics tag: it does not block rendering, and it does not hold up the page if the analytics endpoint is slow or down.

The second is an npm module, installed under a scoped name and initialised with your site identifier and the URL of your reporting endpoint. Its surface is small and worth reading because the shapes tell you what it does. Initialisation creates a global client instance if one does not already exist. There is a predicate to check whether that instance exists, and a getter to reach it. A pageview call requires the client to have been initialised first and will detect the URL and referrer itself if you do not pass them, which is the right default because getting the referrer right by hand is a common source of wrong data. And there is a cleanup function that removes event listeners and resets the global, which tells you the library attaches listeners rather than only exposing a function to call.

The third is the server module, for people who would rather not run anything in the browser. Its options are the interesting part:

```typescript
Counterscale.init({
    siteId: "your-unique-site-id",
    reporterUrl:
        "https://{subdomain-emitted-during-deploy}.workers.dev/collect",
    reportOnLocalhost: false, // optional, defaults to false
    timeout: 2000, // optional, defaults to 1000ms
});
```

Two defaults worth praising. Localhost reporting is off unless you ask for it, so developing a site does not fill your dashboard with your own hits. And the timeout exists, defaults to one second, and the example raises it to two, which is the correct shape for an analytics call on a request path: analytics must never be able to hang the thing it is measuring, so a ceiling is part of the design rather than an afterthought.

The server call takes the URL and hostname explicitly, and a referrer and campaign parameters, which is what makes the campaign reporting possible without a client-side script.

## A tracker is a write path into your account

It is worth being explicit about the security surface, because the documentation is thin on it and the shape is easy to underrate.

The collect endpoint lives on your own deployment, on a public subdomain. Anything that can reach it can post to it. That is inherent to how a client-side tracker works, and it is not a flaw, but it means the write path is public and the question is what a successful write costs you.

The read side is better controlled, because the dashboard and the API token both sit behind credentials you created during setup. So the shape is an unauthenticated write endpoint feeding a service that can query your whole history, guarded by a token with at least account-level analytics permission.

The README does not describe rate limiting on the collect endpoint, payload validation, or whether the endpoint verifies that a request came from a site you registered. It also does not describe what the site identifier does beyond grouping events, and whether an unguessable identifier is the only thing separating your traffic from somebody else's. Those are the questions to answer in the source before you point this at a site whose traffic pattern you care about.

The wider framing is that this whole class of self-hosted analytics has a different threat model from the hosted equivalent. A commercial product spends money on abuse prevention because its customers depend on it. A self-hosted one gives you a deployment you control and a code path you can read, which is the trade, and the trade is only good if somebody on your team actually reads it.

For the same reason, the 90-day limitation section being first in the README is a good sign about how this project communicates. It leads with the constraint rather than the feature list.

## The build tool is pinned to latest

The root package manifest is short, and one line in it is worth pulling out.

The project uses a task runner to orchestrate the workspace, with build, dev and lint scripts delegating to it. The formatting and linting tools are declared with caret ranges, so they can float within a major version. The platform-specific bundler binary, for one Linux architecture, is pinned to an exact version. And the task runner itself is declared as a version literally spelled as the latest release.

That last one is a real, if small, problem. The repository has a lock file and a workspace file, so the intention is reproducible installs, and everything except one build-critical dependency honours that. A floating task runner means a fresh install on a clean machine can resolve a different version than yours, and because the difference is in the orchestration layer, the symptom is a build that behaves oddly rather than a resolution error.

The rest of the configuration is conventional and worth copying. The package is marked private so it cannot be published by accident, the package manager is pinned to a version, and the engines block requires a Node version and a package manager version. There is a four-space tab width for the formatter, which is a matter of taste but at least it is written down.

Two other files in the root tell you how the project works day to day. A shell script for bumping versions, and a coverage configuration, which pairs with the coverage badge and the continuous integration workflow.

There is also a stale field worth mentioning, since both workspace mechanisms are present: a workspaces field pointing at a directory of packages, alongside a workspace file for the package manager. The file is the one the package manager reads; the field is the older convention and appears to be left over.

## Where counterscale is the wrong tool

Five cases, three of them about what analytics you actually need.

If you need to know who your visitors are, this is the wrong tool. The documented events are pageviews with a URL, a hostname, a referrer and campaign parameters. That is traffic measurement. There is no session, no identity, no retention beyond the raw event, and no way to attribute anything to a person. If your question is about users rather than pages, you need something else.

If you need real time, the ninety-day hot store is not the constraint; the ingestion model is. Events go into an analytics store designed for querying what happened, not for triggering, so do not build alerting on top of it.

If you need more traffic than the free tier allows, the README is explicit about the shape of the claim: near-zero operating cost, with the free tier described as hypothetically supporting up to a hundred thousand hits a day. That is an order-of-magnitude figure attached to a platform whose limits change, so treat it as a sizing hint rather than a guarantee.

And if you develop on Windows, the stated requirements are a macOS or Linux environment. That is a real constraint given that the installer is a command line tool.

On maintenance: the release history is three versions, with two patch releases a day apart in December 2025, and the last push is in September 2026 with no release since. That is a project still being worked on rather than one that has stopped, but if you are deploying it, deploy the commit you tested rather than assuming the newest tag is the newest state.

## Conclusion

Adopt counterscale if you want traffic data you own and can afford to run, and treat the dashboard-password prompt as the most consequential step in the install, because declining it leaves an analytics dashboard publicly reachable on a workers.dev subdomain. Check that you can live with a ninety-day hot window and that you are prepared to query the R2 Arrow files yourself for anything older, since that path is not described in the README. Pin your tool versions rather than following the floating build dependency, and note the stated requirements are macOS or Linux.

## FAQ

### What is counterscale and where does it run?

It is a self-hosted web analytics tracker and dashboard that runs on Cloudflare Workers, using Cloudflare Analytics Engine to store queryable events. The design goal is near-zero operating cost, with the Cloudflare free tier described as hypothetically supporting up to 100k hits a day.

### How much history can counterscale show?

Only the last 90 days from the hot store, because Cloudflare's Analytics Engine had a maximum 90-day retention as of February 2025. Longer-term storage is written to an R2 bucket as Apache Arrow files, enabled by default and disabled through the CLI.

### How do I deploy counterscale?

Authorise the Cloudflare command line tool with `npx wrangler login`, then run `npx @counterscale/cli@latest install` and follow the prompts. You will be asked for the Cloudflare API token you created during preparation, and whether to protect the dashboard with a password. A first deployment can take a few minutes before the subdomain is live.

### Do I need to do anything special in my Cloudflare account first?

Yes. You need a Workers subdomain, the Analytics Engine beta enabled under Storage and Databases, and an API token with account analytics permissions. If you have never used Workers you must create a throwaway Worker first, because Analytics Engine cannot be enabled before one exists; you can delete it afterwards.

### How do I add the counterscale tracker to my site?

Either paste the script element served from your deployment into your HTML with a site identifier, or install the published npm module and initialise it with a site identifier and your reporting endpoint. A server-side module is also available for tracking without running anything in the browser.

### Can I record analytics from server-side code with counterscale?

Yes, through the server entry point of the tracker package. Its initialisation takes a site identifier, a reporting endpoint, an optional flag for reporting from localhost that defaults to false, and an optional timeout defaulting to 1000 milliseconds, so the analytics call cannot hold up the request that triggered it.

## Sources

- [benvinegar/counterscale on GitHub](https://github.com/benvinegar/counterscale)
- [License: MIT](https://github.com/benvinegar/counterscale/blob/main/LICENSE)
- [Project website](https://counterscale.dev)
- [README](https://github.com/benvinegar/counterscale/blob/main/README.md)
- [Releases](https://github.com/benvinegar/counterscale/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/benvinegar-counterscale
