CUGA Agent Review: An Enterprise Agent Harness You Configure Instead of Build
CUGA is an open-source generalist agent harness for the enterprise, supporting complex task execution on web and APIs, OpenAPI/MCP integrations, composable architecture, reasoning modes, and policy-aware features.
At a glance
- What is it?
- CUGA is an open-source generalist agent harness from IBM Research that wires OpenAPI specs, MCP servers and LangChain tools into a policy-governed agent. It is aimed at teams who already have APIs and want orchestration, guardrails and evaluation without writing them from scratch.
- Who is it for?
- Adopt CUGA if you have a catalogue of REST or MCP tools and need planning, tool budgets, human approval gates and policy enforcement without building that layer yourself. Do not adopt it if you need a fully managed service or a stable public API surface: the licence file is not recognised as a standard SPDX identifier by the repository metadata even though pyproject.toml declares Apache-2.0, and the dependency pins move with upstream security advisories.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What CUGA Actually Replaces
The README frames the problem in one sentence: building a domain-specific enterprise agent from scratch means handling agent and tool orchestration, planning logic, safety and alignment policies, plus evaluation for performance and cost tradeoffs. CUGA is the project's answer to that stack. Instead of writing a planner, a tool registry, a variable manager and an approval gate, you configure tools and policies and let the harness run.
The target user is not a solo developer prototyping a chatbot. It is a team that already owns APIs, has an OpenAPI spec or an MCP server for them, and needs an agent that can call those tools under constraints. The topics list confirms the emphasis: enterprise, guardrails, policies, sandbox. The repository ships a Dockerfile, a Helm chart under deployment/, and a UBI variant (Dockerfile.ubi), which points at regulated or on-premises environments rather than hobby deployments.
Where it is weaker is the boundary of that claim. The README says the project is "evolving toward enterprise-grade reliability," which is a hedge, not a guarantee. Treat the enterprise framing as a design goal visible in the architecture, not as a certification.
Planner, Executor and the Code Generation Modes
CUGA combines what the README calls best-of-breed agentic patterns: planner-executor and code-act, with structured planning and variable management intended to reduce hallucination on long tasks. The practical expression of that is the code generation profile, set as cuga_mode under [features] in src/cuga/settings.toml. The repository carries a configurations/modes/ directory, and the README names three profiles: fast, balanced and accurate. That is a cost and latency dial, and it is the first thing worth tuning, because the mode changes how much code the agent writes before acting.
Tool integration runs through three routes: OpenAPI specs, MCP servers, and LangChain tools. MCP servers are registered in src/cuga/backend/tools_env/registry/config/mcp_servers.yaml; Python-side, the SDK entry point is CugaAgent(tools=[...]). The registry is a file, not a database, so onboarding a tool is a config edit plus a restart rather than a service call.
Two mechanisms deserve attention because they are unusual. First, reflection: [advanced_features] reflection_enabled in settings.toml lets the agent review its own step before continuing. Second, tool-call budgets: [advanced_features] max_tool_calls_per_block, _per_run and _per_thread. Budgets are the cheapest defence against a loop that keeps calling the same API, and they are enforced by the harness rather than by the prompt.
Installing CUGA and Running the First Demo
The project targets Python 3.12 in its badge and declares requires-python = ">=3.10, <3.15" in pyproject.toml. Dependencies are managed with uv; the Dockerfile installs uv, copies pyproject.toml and uv.lock, and runs uv sync. The environment is configured through a .env file, and .env.example ships a Groq section that is uncommented by default.
Start by copying the example environment file and filling in a key. The default provider is Groq with MODEL_NAME="openai/gpt-oss-120b" and AGENT_SETTING_CONFIG="settings.groq.toml"; alternative blocks for OpenAI, MiniMax, WatsonX and watsonx Orchestrate are present but commented out.
cp .env.example .env
# then edit .env and set GROQ_API_KEYThe README's quick start uses the cuga CLI. Running the CRM demo brings up the web UI, and the Dockerfile overrides the demo port with DYNACONF_SERVER_PORTS__DEMO=7860 and binds CUGA_HOST=0.0.0.0, which tells you the default port is 7860.
uv run cuga start demo_crm --cuga-workspace /app/cuga_workspaceFor a knowledge-backed agent, the README points at cuga start demo_knowledge. Knowledge is on by default (enable_knowledge=True) and ingestion of PDFs, Office files, HTML and Markdown goes through Docling, with agent-level and session-level scopes. Agent skills follow the same pattern: SKILL.md files under .cuga/skills, exercised with cuga start demo_skills, which defaults to sandbox_mode = "native" or can use opensandbox.
uv run cuga start demo_skillsIf you want to draft tools, MCP servers, LLM settings and policies in a browser and then publish a versioned configuration, the README gives cuga start manager. That is the path to production chat, and it separates authoring from the running agent.
Policies, Human Approval and Where the Guardrails Stop
The policy system is CUGA's most concrete differentiator. There are five types: Intent Guard, Playbook, Tool Approval, Tool Guide and Output Formatter. They are configurable through the Python SDK or a standalone UI in demo mode, and Tool Approval is the human-in-the-loop gate: the agent pauses before a tool runs and waits for a person. For an agent that can move money or send email, that gate is the difference between a demo and something an auditor will look at.
The limits are as important as the feature. Policies constrain what the agent does through the harness; they do not sandbox the model's reasoning, and the README does not claim they do. The dependency block in pyproject.toml also shows how much of CUGA's safety depends on upstream: langchain is pinned to >=1.3.9,<1.3.15 with a comment explaining that 1.3.9 fixed a path traversal and sandbox escape in the file-search middleware, and langchain-openai is pinned at >=1.1.14 for an SSRF guard bypass, and langchain-core at >=1.3.3 for unsafe deserialization. Those comments are unusually candid, and they also mean your security posture tracks someone else's release cadence.
The upper bound on langchain is a deliberate trade-off. The comment states that 1.3.15 changes what happens when the summary model call fails: it retries and then raises, which CUGA turns into RuntimeError("middleware_invocation_failed: ...") and falls back to keeping the last N messages, while 1.3.9 through 1.3.14 store an "Error generating summary: ..." placeholder. Both drop the same older messages; the difference is the placeholder and the metrics. Capping a dependency to preserve observable behaviour is defensible, but it does mean you cannot simply take the newest langchain.
Hybrid Browser Tasks, Supervisor Mode and the Second Service
CUGA is not limited to API calls. Setting [advanced_features] mode = 'hybrid' combines API and browser execution using Playwright plus a browser extension, whose readme lives at src/frontend_workspaces/extension/readme.md. This is the computer-use side of the project, and it is the part most likely to break in practice: browser automation depends on page structure, and the README does not describe a fallback when a selector changes.
Multi-agent work goes through CugaSupervisor, started with cuga start demo_supervisor and configured under [supervisor] in settings.toml. External agents can be registered as entries in the supervisor config, and the README references an A2A and remote agents path. The event-driven layer is the most operationally demanding feature: channels, triggers and standing flows run as a second service beside CUGA, started with python -m cuga.backend.events.service alongside cuga start demo. The README states it uses ports 7860 and 8100, that it covers web chat plus Slack, Discord and Telegram, webhooks and cron/poll/push flows armed from natural language with a human confirming each one, and that CUGA itself is unchanged when the events service is not deployed. Setup and per-connector guides live in a separate events documentation repository, which means the integration detail you need is not in this repository.
The Alternative: Langflow and Hand-Built LangGraph Agents
The README itself names Langflow as the low-code visual companion, and the two are complementary rather than competing: Langflow gives you a canvas for designing workflows, CUGA gives you a running agent with policies and budgets. If your team wants to see the graph and edit it visually, Langflow is the more direct fit, and CUGA's value is in the parts Langflow does not supply, namely the policy types, the tool-call budgets and the reflection switch.
The real alternative is building on LangGraph directly, which CUGA itself depends on. That route gives you full control over the state machine and no dependency on CUGA's release cadence, at the cost of writing the planner, the tool registry, the variable manager and the approval gate yourself. CUGA's honest positioning is that it has already made those choices, including the cuga_mode profiles and the max_tool_calls_* knobs, and you inherit them. If your workflow does not resemble a planner-executor loop over API tools, that inheritance is overhead. A retrieval-only question-answering service, for instance, gains nothing from a tool budget or an Intent Guard.
Editorial conclusion
Adopt CUGA if you have a catalogue of REST or MCP tools and need planning, tool budgets, human approval gates and policy enforcement without building that layer yourself. Do not adopt it if you need a fully managed service or a stable public API surface: the licence file is not recognised as a standard SPDX identifier by the repository metadata even though pyproject.toml declares Apache-2.0, and the dependency pins move with upstream security advisories. Before committing, verify the exact Python range in pyproject.toml (>=3.10, <3.15) against your runtime, and run cuga start manager to publish a versioned config rather than editing settings.toml in place.
Frequently asked questions
How exactly do AI agents work?
In CUGA's case the README describes a planner-executor pattern combined with code-act: the agent plans, then executes by generating code that calls registered tools. Tools arrive through OpenAPI specs, MCP servers or LangChain, and the harness enforces limits such as max_tool_calls_per_run and can pause for human approval before a tool runs.
What are the key components of an AI agent?
The repository layout shows the components CUGA treats as essential: a tool registry with mcp_servers.yaml for MCP servers, a settings.toml holding code generation modes and advanced features, a policy layer with five policy types, and a supervisor for multi-agent setups. Knowledge ingestion through Docling and SKILL.md agent skills are optional layers on top.
Which Python version does CUGA Agent require?
pyproject.toml declares requires-python = ">=3.10, <3.15", while the README badge shows Python 3.12 and the Dockerfile builds from python:3.12-slim-trixie. The watsonx Orchestrate extra is noted as requiring Python 3.11 or newer.
How do I connect an MCP server to CUGA Agent?
MCP servers are registered in src/cuga/backend/tools_env/registry/config/mcp_servers.yaml, and the README also shows the SDK route of passing tools to CugaAgent(tools=[...]). The manager UI offers a third path: draft MCP servers in the web interface and publish a versioned config.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cuga-project-cuga-agent)