Model or dataset
coze-dev/coze-loop avatar
coze-dev/coze-loop

Coze Loop: a self-hosted platform for prompt debugging, agent evaluation and trace observability

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.

5,746 stars797 forksGoApache-2.0

At a glance

What is it?
Coze Loop is the open-source edition of Coze's agent operations platform, written in Go and licensed under Apache-2.0. It bundles a prompt Playground, an evaluation module and trace observability behind one Docker Compose or Helm deployment, but model configuration is a manual YAML edit and rollback is not documented.
Who is it for?
Adopt Coze Loop if you already run OpenAI or Volcengine Ark models and want prompt debugging, evaluation sets and trace reporting in one self-hosted service rather than three separate tools. Do not adopt it if you need a managed SaaS with no cluster to operate, or if you cannot give it a model API key at deploy time, because the README's deployment path stops until model_config.yaml is edited.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Coze Loop is for, and who it is not for

Coze Loop targets a specific gap: the tooling around an agent after the prompt is written. The README describes it as a "developer-oriented, platform-level solution focused on the development and operation of AI agents", covering prompt engineering, evaluation, and post-deployment monitoring. That is three jobs that teams usually solve with three tools, and the pitch is that one deployment covers all three.

The intended user is a developer or platform team running agents on their own infrastructure. The repository is Go, the licence is Apache-2.0, and the deployment artifacts are Docker Compose and a Helm chart, so the assumption is that you have somewhere to run containers and someone who can edit a YAML file. If you want a hosted product with a signup form, this is the wrong shape of project. The README also states that the open-source edition is derived from a commercial version and offers "core foundational feature modules", which is an explicit warning that the open-source build is a subset, not the whole product.

The three modules and how they connect

The feature table splits the platform into prompt debugging, evaluation, observation and model integration. Prompt debugging means a Playground for interactive testing plus prompt version management. Evaluation means evaluation sets, evaluators, and experiments that score prompt or agent output on dimensions such as accuracy, conciseness and compliance. Observation means an SDK reports traces, and the platform stores and displays them.

The data flow implied by the README is: a prompt is authored and versioned in the platform, tested in the Playground against a configured model, then scored by an evaluation experiment, and once the agent runs in production its SDK trace reports land in the same system. The observability description is the most specific part: it records "every stage from user input to AI output, including key stages such as prompt parsing, model invocation, and tool execution", and captures intermediate results and exceptions automatically.

Model access is a configuration concern rather than a code concern. The README states support for OpenAI, Volcengine Ark and other models, and both deployment paths route through a model_config.yaml file where api_key and model are set. That file is the seam between Coze Loop and whatever model you actually pay for.

Installing Coze Loop with Docker Compose

The README's first deployment method assumes Docker Engine is already installed and running. Clone the repository first.

bash
git clone https://github.com/coze-dev/coze-loop.git
cd coze-loop

Before starting anything, edit release/deployment/docker-compose/conf/model_config.yaml. The README says to modify the api_key and model fields, using Volcengine Ark as its example: api_key is the Ark API key, and model is the Endpoint ID of the Ark model access point. The README links separate documentation for users inside and outside mainland China, so the key you need depends on which endpoint you are using. Nothing starts correctly without this step.

bash
# Run in the coze-loop/ directory
make compose-up

The Makefile confirms the Compose directory is ./release/deployment/docker-compose, which is where the file you just edited lives. The README notes this runs in development mode by default. Once the containers are up, open http://localhost:8082 in a browser to reach the open-source edition. If the page does not load, the model configuration and the Compose logs are the two places to look first.

The Helm path and the ingress step people skip

The second deployment method targets a Kubernetes cluster that already has the Nginx Ingress add-on enabled, plus kubectl and Helm installed. The README suggests Minikube if you want to try it locally.

bash
helm pull oci://docker.io/cozedev/coze-loop --version 1.0.0-helm
tar -zxvf coze-loop-1.0.0-helm.tgz && cd coze-loop && rm -f ../coze-loop-1.0.0-helm.tgz

After extraction, the model configuration moves to release/deployment/helm-chart/umbrella/conf/model_config.yaml, with the same api_key and model fields and the same Volcengine Ark example. Then comes the step that is easy to underestimate: the README instructs you to configure templates/ingress.yaml according to your cluster, manually modifying parameters such as ingressClassName and setting class, instance, host and IP allocation. This is not a value the chart can guess for you, and it is the most likely reason a Helm install appears to succeed while nothing is reachable. The README's chart version is 1.0.0-helm, which is a different numbering line from the application releases, so do not expect the two to match.

Where Coze Loop will disappoint you

The deployment story has a hard dependency the README does not soften: you must supply a model API key before the platform is useful. There is no documented offline mode or local model path, so an air-gapped environment is not addressed. The README also does not document rollback, which matters because the Compose path is described as development mode by default and there is no upgrade or downgrade procedure in the README or the wiki link it points to.

The open-source edition is explicitly a subset of the commercial product. The README frames it as "core foundational feature modules", so if you are evaluating Coze Loop against the commercial Coze offering, assume feature parity is not the goal. Finally, the recent release history is uneven: v1.4.1 landed on 2025-10-21, then v1.5.0 and v1.5.1 arrived on 2026-01-19 and 2026-01-20. The last push to the repository was on 2026-09-10, so the codebase is moving, but the tagged releases do not track it closely. Pin a tag rather than tracking main.

Coze Loop compared with Giskard and Agenta

The two alternatives most often mentioned alongside Coze Loop are Giskard and Agenta, and the difference is architectural rather than cosmetic. Giskard is described in the related searches as an LLM scanning tool: its centre of gravity is testing and scanning a model or pipeline for problems, run as a library or scan job against your own code. Agenta is positioned as a self-hosted prompt and evaluation platform, which overlaps most directly with Coze Loop's Playground and evaluation modules.

The distinction that matters is where the boundary sits. Coze Loop is a platform you deploy and then point your agent at, with trace reporting handled by an SDK and the evaluation surface living inside the deployed service. That gives you one place to look, at the cost of operating that place. A library-oriented tool like Giskard embeds in your existing test suite, so there is nothing to run, but there is also no Playground or trace store. If you already have observability and only need evaluation, Coze Loop's breadth is overhead. If you have neither and want both from one Compose file, the breadth is the point.

Licence, maintenance and what upgrades cost

The licence is Apache-2.0, which is permissive and includes an explicit patent grant. The practical implication for most teams is that you can run a modified Coze Loop internally without publishing your changes, but the usual Apache-2.0 conditions still apply to redistribution, including retaining the licence and notice files. That is a summary of the licence identifier, not legal advice; read LICENSE in the repository root before you ship anything derived from it.

On maintenance, the evidence is mixed in a way worth stating plainly. The repository is not archived, and the last push was on 2026-09-10, so commits are landing. But the most recent tagged release is v1.5.1 from 2026-01-20, roughly eight months earlier. The upgrade cost therefore depends on which of those two lines you follow. Running a pinned tag means you are on code that has not been released since January. Running main means you inherit whatever is in the tree, and the README offers no migration or rollback documentation to fall back on. The Makefile does define image build and push targets, including a multi-platform buildx target for linux/amd64 and linux/arm64, so building your own image from a specific commit is a supported-looking path if you need to freeze a version.

Editorial conclusion

Adopt Coze Loop if you already run OpenAI or Volcengine Ark models and want prompt debugging, evaluation sets and trace reporting in one self-hosted service rather than three separate tools. Do not adopt it if you need a managed SaaS with no cluster to operate, or if you cannot give it a model API key at deploy time, because the README's deployment path stops until model_config.yaml is edited. Before committing, verify three things: that your model provider is reachable from the host, that port 8082 is free or remapped, and that the Helm chart's ingress.yaml matches your cluster's ingressClassName. Check the release page for anything after v1.5.1 rather than assuming the main branch is what you will run.

Frequently asked questions

What is an agentic coding loop?

The README does not define this term. Coze Loop is described as a platform for the development and operation of AI agents, covering prompt debugging, evaluation and observability across the agent lifecycle.

How does loop engineering work?

The README does not describe loop engineering as a discipline. Coze Loop's own mechanism is a deployed platform: prompts are authored and versioned, tested in a Playground against a configured model, scored by evaluation experiments, and traced in production through an SDK.

What does "loop" mean in the context of AI?

The README does not define the term in general. Coze Loop's own use of it refers to the full lifecycle of agent development and operation, from prompt writing through evaluation to monitoring.

Official sources

  1. coze-dev/coze-loop on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/coze-dev-coze-loop.svg)](https://hysenlabs.com/projects/coze-dev-coze-loop)