The Flawless repository ships a product called CISRE, on one compose service
AI SRE AgenticOps for Kubernetes and cloud infrastructure.
At a glance
- What is it?
- An SRE platform for Kubernetes where the model plans and a typed action chain does the changing, with recovery defined by new evidence rather than a 200 response. The project file is in Chinese, the deployment is a single container, and the kernel is four ports wide.
- Who is it for?
- CISRE is worth reading if your problem is an agent that claims a fix worked when it did not, because that is the specific failure the whole design is built around, and the same-target readback plus recovery verifier is the part most tools skip. It is not a general operations platform yet.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 52 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The repository is Flawless and the product is CISRE
The repository is named Flawless, and the product described inside it is called CISRE, expanded as Cloud Infrastructure Site Reliability Engine, at version 5.6.0. Only one project file exists at the root and there is no English variant among the top-level entries.
The older name is still everywhere else. The compose file declares its project name as flawless, every build argument in it is prefixed FLAWLESS_, and the state paths point at a flawless directory:
EFFECTIVENESS_STORE_PATH: /var/lib/flawless/effectiveness-state.json
KNOWLEDGE_STORE_PATH: /var/lib/flawless/knowledge-base.json
MODEL_PROFILES_STORE: /var/lib/flawless/model-profiles.json
OPS_SKILL_ROOT: /var/lib/flawless/ops-skillsThe environment template keeps it too, pointing the knowledge store at a temporary flawless path. There have been no published releases, so 5.6.0 exists only as a version string in the project file and the last push was 2026-08-14. Anyone searching for this project should try both names, because a person who has read the documentation is looking for CISRE and a person looking at a cluster is looking for Flawless.
A 200 response is not a recovery
The central rule is stated before anything else, and it is worth quoting the shape of it. Success from an execution API does not mean recovery. A model saying the task succeeded does not mean recovery. An old instance still being healthy does not mean recovery. A task closes only when new evidence from the real target satisfies the recovery contract.
The pipeline that enforces it is nine steps:
发现 → 取证 → 诊断 → Skill 路由 → 变更预览 → 人工审批
→ 执行 → 同目标回读 → 稳定性验证 → Records / Skill 成效
↘ 未恢复:保留证据并换策略继续Discovery, evidence collection, diagnosis, skill routing, change preview, human approval, execution, readback on the same target, stability verification, then records. The branch underneath is the important part: if the target has not recovered, the evidence is kept and the strategy changes rather than the task being closed.
The chain a real change has to travel is written out separately:
typed action → policy / blast radius → human approval → executor
→ same-target readback → recovery verifier → recordAnd the security section forbids the three shortcuts that make the rest moot: no treating an HTTP 2xx as success, no treating a command exit code of zero as success, and no treating a model conclusion as success.
Contract ready does not mean connected
The capability table has six rows and only one of them is complete. Kubernetes is listed as a full closed loop, reachable through Rancher, through an uploaded or pasted kubeconfig, or through an in-cluster service account. Database, virtual machine and host, storage, middleware and cloud resources, and network are all listed as extension contract ready.
The document then defines what that phrase means, which is the most useful sentence in the file: contract ready means the interfaces, permissions, audit and closed-loop semantics are in place, and does not mean any specific product has been connected. The access column says as much in each row, naming a read-only provider, a typed action executor, an array, CSI or storage platform provider, stable adapters and harness service contracts, and providers for switching, routing, load balancing, DNS, access control and link management.
Two constraints go with it. Pages must not fabricate resource or health data, so a dashboard showing a green service means something was queried. And an external plugin cannot obtain the mutate, execute or secrets read permissions directly, nor inject arbitrary shell, SQL or HTTP mutations into the API process. Those three capabilities sit inside the core, which is why the plugin route requires human approval at the end of its chain.
Everything new goes in plugins, and application.py is still shrinking
The stated architecture is plugin first, and the document is careful not to oversell it. The plugin runtime and the cross-team contracts are usable, the Kubernetes closed loop still runs through a compatibility service, and one file, backend/app/application.py, is described as still being reduced. The instruction to other teams is blunt: do not keep adding vendor SDKs or new business branches to that file.
The target is stated as submitting only plugins and skills with zero core changes, and a roadmap document is referenced for the migration boundary. The layout backs this up, with plugins/, agents/, mcp_servers/, manifests/, charts/ and deploy/ all sitting beside the backend rather than inside it.
Underneath sits a kernel contract, cisre.kernel.ports/v1, fixing four replaceable ports: an append-only event journal, a distributed lease with a fencing token, a durable job queue, and a snapshot store with compare and swap. Agents, plugins and skills depend on those contracts only, and explicitly not on PostgreSQL, Redis or Kafka. The agents directory is annotated as holding model inference and no write permission, which is the same separation the approval chain enforces at runtime.
One compose service, seven loopback ports
The compose file declares a single service. There is no separate container for the model gateway, the observability layer, the agents or the MCP server. Instead the one container is told where each of them lives, and every address is loopback:
ADAPTER_URL: http://127.0.0.1:8200
OBSERVABILITY_URL: http://127.0.0.1:8100
HEALING_AGENT_URL: http://127.0.0.1:8101/a2a/tasks
INCIDENT_AGENT_URL: http://127.0.0.1:8102/a2a/tasks
POSTMORTEM_AGENT_URL: http://127.0.0.1:8103/a2a/tasks
MCP_SERVER_URL: http://127.0.0.1:8105/mcp
CMDB_URL: http://127.0.0.1:8300That arrangement lines up exactly with the stated scaling story. The default backend is file and in-process, which suits a single replica, and the page says so by displaying a single replica label rather than reporting distributed readiness. There is an endpoint for that check, GET /api/harness/scalability, and the document says a multi-replica deployment still on the local backend produces explicit violations from it.
The swap path is described as replacing the same four ports with a transactional event store, a distributed lease, a durable queue and transactional snapshots, then scaling the stateless API and worker horizontally. The semantics that must not move are named too: an idempotency key per change, one fencing token per target, a bounded queue that applies backpressure instead of pretending to be overloaded, replayable events and rebuildable read models.
The nginx stage installs packages as root before dropping to uid 101
The container build has three image arguments, and two of them are floating while one is not. Node defaults to node:24-slim and Python to python:3.13-slim, both without a digest, while the nginx image is pinned to a specific sha256 digest on the stable alpine line.
The frontend build stage is pinned to the build platform rather than the target platform, with a comment explaining why: running TypeScript and Vite under emulation makes multi-architecture builds unstable, so the output is architecture independent static files. The install step branches on whether a lock file exists, choosing a clean install when it does and a normal install when it does not.
The runtime stage is where the detail sits. The base image is the unprivileged nginx build, so the file starts as root, installs curl and libcurl with the package manager, and then switches to user 101 before copying its configuration and the built frontend. The pattern works, but it means the unprivileged guarantee is established after an install step rather than being inherited.
One more mismatch: the compose file builds target backend-runtime, while the stages visible in this file are frontend-builder and frontend-runtime, so the stage the compose file depends on sits further down.
Gray release requires Prometheus that the template leaves blank
The environment template is where the project's assumptions become checkable, and it has a few values worth reading before anything else. The gray release switch is on, Prometheus is required, and the verify timeout is 3600 seconds, while the Prometheus URL itself is empty and both default queries, one for error rate and one for p99 latency, are also empty:
GRAY_RELEASE_ENABLED=true
GRAY_RELEASE_REQUIRE_PROMETHEUS=true
GRAY_RELEASE_VERIFY_TIMEOUT_SECONDS=3600
PROMETHEUS_URL=
GRAY_RELEASE_DEFAULT_ERROR_RATE_PROMQL=So out of the box, the gray release path asks for a metrics source it has not been given. Set the URL and the two queries, or expect that check to have nothing to read.
The rest of the template makes the local-first intent explicit. The model endpoint points at a local service on port 11434 with qwen2.5:7b selected and authentication set to none, the embedding model is nomic-embed-text with embedding use switched off, Langfuse tracing is disabled with empty keys, and the template header states that every component runs locally or on a private network with no cloud service dependency.
Cluster credentials are handled with care worth copying. Web-uploaded kubeconfigs are encrypted with Fernet before being persisted, and the comment says production belongs in a Kubernetes secret, while a single instance with the key left blank generates a local key with owner-only permissions beside the database.
Local setup is one stack script and two test gates
The backend is started with a conventional sequence, and the last line is the part that matters, since a single command brings up the whole set of loopback services rather than the API alone:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
python -m pytest tests
python scripts/run_local_stack.py --host 127.0.0.1 --api-port 8080The frontend is the usual three commands from its own directory, and the pre-commit minimum is two gates:
python -m pytest tests
cd frontend/modern && npm run buildThe dependency list explains the shape of the system. FastAPI, uvicorn and pydantic for the service, the official Kubernetes client for cluster work, a cryptography pin, then a model stack built on langchain and langgraph with checkpoints and an SDK, alongside the model context protocol package, server sent events for streaming, PDF parsing, Langfuse for tracing and the OpenAI client. A lock file sits next to the requirements file at the root, so the range pins in the requirements file are meant to be resolved once rather than floating.
Editorial conclusion
CISRE is worth reading if your problem is an agent that claims a fix worked when it did not, because that is the specific failure the whole design is built around, and the same-target readback plus recovery verifier is the part most tools skip. It is not a general operations platform yet. Only Kubernetes is described as a complete closed loop; database, virtual machine, storage, middleware, cloud and network are all at contract ready, which the project itself defines as interfaces existing rather than products connected. Two practical notes before you try it. The product is named CISRE while the repository, the compose project, the environment variable prefix and the state paths are all still Flawless, so search for both names. And the default compose file runs every component as loopback services inside one container against a single replica file backend, which is a local development arrangement rather than a deployment shape.
Frequently asked questions
Why is the Flawless repository described as CISRE?
They are two names for one project at two stages. The repository, the compose project name, the FLAWLESS_ environment prefix and the state paths under /var/lib/flawless all use Flawless, while the product described in the project file is CISRE, Cloud Infrastructure Site Reliability Engine, at version 5.6.0.
Does CISRE support more than Kubernetes?
The capability table lists Kubernetes as a full closed loop and marks database, virtual machine and host, storage, middleware and cloud resources, and network as extension contract ready. The project defines that phrase as interfaces, permissions, audit and closed-loop semantics being in place, not a specific product being connected.
How does CISRE decide that a change actually fixed something?
Not by the execution API returning success, by the model saying so, by an old instance still being healthy, by an HTTP 2xx, or by a command exit code of zero. A change has to run a typed action, a policy and blast radius check, human approval, an executor, a readback on the same target and a recovery verifier before the task records a result.
What does a team have to deliver to add a new resource domain to CISRE?
Six items: a manifest declaring the identifier, semantic version, provides and requires, permissions and events; a read-only provider service for discovery, evidence and verification; a skill file per incident holding trigger, evidence, root cause, action, rollback and success criteria; an action catalog of typed actions with arbitrary shell and SQL forbidden; contract tests covering success, timeout, permission denial, rollback and verification; and a README covering scope, limits, on-call ownership and compatibility.
Does CISRE need a cloud service or an API key to run locally?
The environment template points the model endpoint at a local service on port 11434 with a 7b model selected and authentication set to none, disables Langfuse tracing, and states that all components run locally or on a private network without depending on any cloud service. Kubernetes credentials are the exception, since production expects them injected from a Kubernetes secret.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/william-lu-stack-flawless)