# polis: an OpenID Connect simulator, a local certificate authority, and a published method

> This is a large-scale feedback platform for public consultation, and the two things that make it worth reading are not its feature list. The first is that local development runs against a real OpenID Connect simulator with a proper local certificate authority, including a certificate for a hostname that only exists inside the container network. The second is that the clustering method is a published paper rather than a black box.

**compdemocracy/polis** — :milky_way: Open Source AI for large scale open ended feedback

- Repository: https://github.com/compdemocracy/polis
- Website: https://pol.is
- Stars: 1,210 · Forks: 262
- Language: Python
- License: AGPL-3.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/compdemocracy-polis

## An OpenID Connect simulator, a local certificate authority, and a restart instruction

The first step of the quick start is installing a tool for creating a local certificate authority, and the reason the project needs one is that it does not stub out authentication in development. It runs an OpenID Connect simulator instead.

That is a significant choice for a platform of this kind. A participant-facing consultation tool has to know who somebody is, has to keep their answers associated across a session, and in a real deployment would delegate that to an identity provider. The easy path in development is to fake it: a test user, a hardcoded identity, an unsigned token. The consequence is that the identity code path is the one part of the system never exercised locally, so the first time it runs is in front of real people.

This project runs the real thing. The simulator speaks the protocol, so the application code that talks to an identity provider is the same code locally and in production, differing only in the endpoint. That is the right architecture and it is a lot of extra work.

The certificate setup is where the cost shows. Locally trusted certificates have to be generated for a list of names, and the list is not just localhost. It includes the loopback address and the IPv6 loopback, a hostname for the simulator itself, and a name that only resolves inside the container network. That last one is the detail that tells you how much thought went into this: a certificate for a hostname that does not exist on your machine, so that a request from inside the network to the host resolves to something the certificate trusts.

Then the root certificate is copied to a known location, with a comment saying it is needed for server-to-server communication. That is the part people forget. Trusting a certificate in a browser is one thing; a service making an outbound request to another service has its own trust store, and it has to be told about the local authority separately. The setup step exists because the failure without it is a TLS error somewhere in the middle of a request chain.

And then the instruction to completely restart the browser. Locally trusted certificates are installed into a system trust store, and every running browser process has a snapshot of that store from when it started. Not restarting produces an error that looks like a misconfiguration and is not. That sentence is in the documentation because somebody lost an afternoon to it, and it is the clearest signal in the whole readme that these steps have been run by the person who wrote them.

## The Makefile treats environment parsing as expensive, because make thinks it is

There is a comment in the build file that reads as housekeeping and is actually the most interesting engineering in the repository.

It says, in effect, that parsing the environment is lazy, evaluated only when needed. The macro is three lines:

```
define parse_env_value
$(shell grep -e ^$(1) ${ENV_FILE} | awk -F'[=]' '{gsub(/ /,"");print $$2}')
endef
``` And the machinery to achieve it is a three-line macro that greps a line out of the configuration file, splits it on the equals sign, strips spaces, and takes the second half. There is a second macro that takes a value, lowercases it, and reports true unless it is exactly the word false.

Why bother? Because of how make works. Any expression that shells out is run every single time it is expanded, and a variable that appears in ten places is expanded ten times. If each expansion runs a pipeline of three processes, then a make invocation that touches that variable thirty times has spawned ninety processes before it has done any work. On a container-backed development machine that is not a rounding error, and it is the difference between a make target feeling instant and feeling like it is thinking.

Wrapping the expensive work in a definition that expands a variable once, into an evaluated assignment, is the fix. The comment is the kind of thing that gets written by somebody who profiled the build and then left a note so the next person does not undo the optimisation. That is worth more than the optimisation.

The boolean macro is worth a second look for a different reason. The convention is that a variable is true unless it is literally the word false, which means a typo, an empty value, or a stray space produces true rather than an error. That is the right default for a flag that controls whether a database runs in a container, and it is also the kind of convention that surprises people. The space stripping in the value macro is there precisely because of that: somebody typed a space in the configuration file, and rather than making it an error they made the parser tolerant. That is a defensible trade and it is worth knowing about before you spend an hour wondering why a flag you set to false is still true.

The header of the same file documents the command surface in four examples, and they all share a shape. A target, optionally preceded by a word that selects a configuration, optionally followed by a word that selects an environment file. So the same start target runs against the development file, the production file, or a file you name. That is a small design language for a build file, and it is legible in the four lines of its own documentation.

## A production clone cannot share a volume, because the volume name comes from the mode

Among the variables the build file computes, two are worth pulling out, and neither is obvious from its name.

One of them selects where the database's data lives. The volume name is not a constant. It is derived from a mode flag, so that one mode uses a volume named for production and the other uses a volume named for development. Two names, one machine, no overlap.

That is a data-safety mechanism disguised as configuration, and it is the right way to do it. The failure it prevents is the specific and common one where a developer makes a copy of production data to debug against, runs the stack, and the development database writes into the production volume. Nothing about that is likely to be noticed until the data is wrong. A volume name that is a function of the mode makes the mistake impossible rather than merely discouraged, and it does so without asking anybody to remember a convention.

The second derived value is how the database is initialised, and it also branches on the same mode. So a production clone is not just a different volume, it is a different initialisation path. That is consistent: if the data is a copy of something real, it should not be re-initialised, and if it is a fresh development database it should be.

The same flag appears in the compose file's own header comment, which explains that the database service is included by default unless a variable says otherwise, and that you can omit the profile entirely to run against a database on your host machine or a remote server. So the profiles in the compose command are not a build convenience. They are the mechanism for swapping infrastructure, and the build file picks which profiles to include based on the same configuration the compose file reads.

That is a two-layer arrangement and it is worth naming, because it means the compose file stays free of environment-specific branching. The compose file declares services and profiles. The build file decides which profiles to activate, based on the environment file. So there is one place where the decision is made and one place where the options are defined, and neither duplicates the other.

The commit hash is also exported at parse time, from the version control system, which means the image can identify exactly which revision it was built from without anybody having to remember to set a variable. For a platform where a bug report needs to be reproducible, that is a small thing that removes a whole class of ambiguity.

## Four compose overlays, one of which is for recovery

There are four compose files. One is the base. One is for development. One is for tests. And one is for recovery, and that fourth one is the interesting entry.

Compose overlays work by merging, so each file only has to state its differences. A base file defines the services, their networks, their ports and the configuration they need. A development overlay says run the local dependencies and use the development settings. A test overlay says run against a throwaway database. That is a normal arrangement and it scales well to about three overlays.

A recovery overlay is not normal, and its existence tells you something. It means somebody hit a state, in production or in a long-running development environment, where the normal start did not get you back to a working system, and they wrote down the sequence that does. That sequence is a small piece of operational knowledge, and the fact that it is version controlled rather than living in a support ticket or in one person's shell history is the difference between knowledge and folklore.

It also suggests the system has state worth recovering. Which it does: a consultation platform accumulates responses, and a database that has accumulated months of them is not something you recreate from an image. The recovery overlay being checked in rather than improvised means the worst day of somebody's year is reproducible, which is the whole point.

The base compose file deserves a mention for a documentation choice that is better than it first looks. Its first twenty-five lines are a comment block, and they are not a description of the services. They are the operating instructions: what to copy before the first run, how to build, how to run afterwards, how to force a full rebuild, what to do if you changed only the configuration file, and how to stop. Six commands, each with the condition under which you need it, in the order you will need them.

That placement matters more than it would in a readme. A readme is a document you choose to open. A compose file is a document you scroll past. Putting the recovery and the quick-rebuild commands in the place where the reader is already looking, rather than three sections away, is a small act of writing for the actual reader rather than for the file's author.

## Three deployment targets across two infrastructure tools and a hosted platform

The deployment surface is broader than most projects of this size, and reading the file names tells you what it consists of.

There is a specification for a managed deployment service on one cloud provider, a directory for the same cloud provider's infrastructure-as-code tool, and a configuration file for a well-known hosted application platform. Plus the compose files, which are the fourth option and the one most deployments will actually use.

Two infrastructure-as-code approaches for the same cloud is unusual and worth a moment. It is either a legacy path being kept alive alongside a new one, or two deployment models that genuinely differ, which happens when some deployments are long-lived and hand-tuned while others are rebuilt from scratch on every change. The readme does not say which, and the presence of a changelog and a version file at the root suggests the project has been running long enough for both situations to have arisen.

The hosted platform configuration is the interesting one for a different reason. A configuration file for a hosted application platform implies that at some point this ran as a service on somebody else's infrastructure, which for a project that describes its main deployment as free for nonprofits and government means the economics were at some stage somebody else's problem. Whatever the current arrangement, the file is a record of how the platform was made runnable as a service rather than as a process on somebody's laptop, and that is the harder half of deploying a multi-service system.

The number of distinct services in the base compose file is itself notable. There is a front-end build for the participation interface, another for an in-progress version of it, another for administration, another for reports. There is a server, a service whose name suggests it runs the clustering, an identity simulator, and a file server. Plus the database, behind a profile.

So this is not a monolith. It is roughly eight services with their own build contexts, and the reason the deployment tooling is broad is probably that the team has had to deploy this several different ways for several different users. A system used by a small nonprofit on a single virtual machine and a system used by a city have very different requirements, and a project that supports both ends up supporting several deployment targets because it cannot choose.

The recovery overlay from the previous section is the same story from the other direction. Eight services with a database means a lot of ways to end up half-started, and a checked-in way back is what keeps that from being a crisis.

## The method is a paper, which is what makes the central claim checkable

The readme's second line points at a methods paper, and for a project in this category that link matters more than any feature.

The platform's central claim is not a feature claim. It is a claim about method: that you can ask an open question to a large group of people and get back something more useful than a poll, with less effort than a focus group. That claim is either true or it is not, and it is not something a demonstration can establish, because the thing being claimed is about how people respond to a process rather than about whether the software works.

That makes the paper load-bearing in a way it would not be for most repositories. A closed implementation of the same idea would be a marketing claim with a user interface attached. A published method with a specified algorithm is a claim that a third party can evaluate, reimplement, and disagree with, which is the only form of legitimacy available to a civic technology project that is asking public bodies to change how they consult.

The other structural choice supports the same goal. The main deployment is described as free for nonprofits and government, and the licence is the copyleft variant that requires modified versions to be published. For a platform whose credibility comes partly from being auditable, and whose users are public institutions that are themselves subject to transparency rules, an open licence is not merely a preference. It is the mechanism by which a city can be sure what software is collecting their residents' responses.

That is a real constraint, and it is worth being honest about its cost, because the deployment section in the previous part is partly a consequence of it. Self hosting is the only way to use a strongly copyleft platform with full feature access, so the project has to make self hosting work for a small nonprofit with one server and for a city with a platform team. Eight services, four overlays and three deployment targets is what that costs. It is a lot of operational surface, and it is the direct price of the licence and of the refusal to make the method proprietary.

The remaining pieces of the readme are smaller but consistent with the same posture. Questions go to a discussion forum and well-defined bugs go to the issue tracker, and the distinction is stated. There is a direct email address for anyone applying the platform in a high-impact context who needs more help than the public channels give them. And the project board is described as somewhat incomplete but still useful, with the reason given: the team stopped updating it when the forge's replacement tool arrived and have not migrated. Naming your own stale tooling, and saying why, is a small thing that makes everything else in the readme easier to trust.

## Conclusion

Polis is worth evaluating if you are running public consultation and want a mechanism that produces something other than a survey summary, and you are willing to host it yourself, which for a nonprofit or a public body means owning the operational surface. It is a poor fit if you need something you can deploy in an afternoon, because the local setup involves generating a certificate authority, issuing certificates for container-internal hostnames and restarting your browser afterwards. Read the methods paper before you commit, because the clustering is the product and the platform is a delivery mechanism for it, and configure the production-clone mode deliberately so your development and your live data cannot end up sharing a database volume.

## FAQ

### What is the Polis platform?

It is an AI powered platform for gathering sentiment at scale, positioned as more organic than a survey and less effort than a focus group. The main deployment is described as free for nonprofits and government, and the clustering method that distinguishes it is published as a methods paper.

### Why does setting up Polis locally require a local certificate authority?

Because local development runs against an OpenID Connect simulator rather than a stubbed identity provider, so the real authentication code path is exercised. Certificates are generated for the loopback names, the simulator, and a hostname that only resolves inside the container network, and the root certificate has to be placed for server-to-server requests as well as in the browser.

### What does the USE_PRODCLONE setting do in the Polis build?

It changes both where the database data lives and how the database is initialised, deriving the volume name from the mode rather than using a constant. That makes it impossible for a development stack and a production clone to share a volume on the same machine, which is the specific data-loss mistake the setting is guarding against.

### Why does the Polis Makefile avoid parsing the environment file more than once?

Because any make expression that shells out is re-evaluated on every expansion, so a variable used in many places would spawn a pipeline of processes many times over. The build file wraps the parsing in definitions expanded into evaluated assignments, with a comment marking the intent, so the cost is paid once.

### How many deploy paths does Polis support?

Several: compose files with a base and development, test and recovery overlays, a specification for one cloud provider's managed deployment service, infrastructure-as-code configuration for the same cloud, and a configuration file for a hosted application platform. The breadth reflects a system used both by small nonprofits on one server and by larger organisations.

## Sources

- [compdemocracy/polis on GitHub](https://github.com/compdemocracy/polis)
- [Issues](https://github.com/compdemocracy/polis/issues)
- [License: AGPL-3.0](https://github.com/compdemocracy/polis/blob/edge/LICENSE)
- [Project website](https://pol.is)
- [README](https://github.com/compdemocracy/polis/blob/edge/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/compdemocracy-polis
