Model or dataset
Abilityai/trinity avatar
Abilityai/trinity

Trinity runs every agent in its own Docker container

Self-hosted AI Agents Platform supporting Claude Code, Codex, Gemini agents. Apache 2.0.

618 stars119 forksPythonApache-2.0

At a glance

What is it?
Abilityai/trinity turns fleets of Claude Code, Codex and Gemini agents into isolated Docker containers on hardware you control, adding cron scheduling, agent-to-agent delegation, cost tracking and a tamper-evident audit trail. Its own compose files and environment template show where the operational edges are.
Who is it for?
Trinity fits teams that want agents running on infrastructure they control and who can accept the Docker and PostgreSQL obligations that come with it. Before deploying, verify three things: that every secret is generated per environment instead of inherited from a fallback, that agent containers pick up the log cap configured outside compose, and that the compose file your start script loads is the one carrying the restart policy.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Claude Code authors the agent, the container runs it

One line in the README sets the whole design apart: Claude Code writes the agent, Trinity runs it in production. An agent here is not a chat window attached to somebody's account. Each one is started as its own isolated Docker container, and the layer around those containers supplies what a single-agent setup does not have: real-time observability into each one, fleet-wide health monitoring, cron-based scheduling, agent-to-agent delegation, cost tracking, and a tamper-evident audit trail of what ran.

That split is why the project argues against three alternatives instead of against rival agent frameworks. A SaaS platform is described as letting data leave your security perimeter and creating vendor lock-in. Building the same thing yourself is costed at 6 to 12 months and $500K or more in engineering. A framework is described as having no observability and no fleet management. Trinity's answer is your own hardware, or any cloud you control.

The licensing follows from that pitch. It is Apache 2.0, free for any use including commercial, and the project states it was independently pentested to UnderDefense Grade A, with the authors running it in production themselves alongside their customers. The main branch was last pushed on October 2, 2026.

Workspace became the working surface in v0.9.5

Release v0.9.5, published September 17, 2026, moved the centre of gravity from fleet dashboards to a chat surface the project calls Workspace. The feature list is specific rather than aspirational: chats per topic, a conversation rail, live execution cards, agents that ask you questions, a canvas for every agent, and real-time voice mode.

The same release widened deployment. Prebuilt images and a guided DigitalOcean install arrived alongside first-run setup in the browser, so a fresh instance no longer assumes you already have a shell session open. Two operational details landed with it: subscription usage and headroom are visible instead of hidden, and instance telemetry is opt-in rather than assumed. Credentials are encrypted at rest, which is what lets the deployment table claim a multi-user production setup.

Getting there is described as three steps. Stand up an instance with one command on your own machine, on any server you control, or on a DigitalOcean Droplet in about ten minutes. Connect Claude Code through the abilities plugins, which scaffold, connect and deploy agents over MCP. Then run it scheduled, multi-user and audited inside your own perimeter. The repository description names Claude Code, Codex and Gemini as the supported agent kinds, and AGENTS.md is the router pointed at first, holding exact commands, key facts and verification steps.

SQLite by default, PostgreSQL for production

Trinity ships SQLite as the default so a first evaluation costs nothing to configure, then tells you not to keep it. The important callout in the README names PostgreSQL as the recommended production database and says production instances should run it, opted into with a single DATABASE_URL. The walkthrough lives in docs/POSTGRESQL_SETUP.md, and the request is tracked as issue #300.

The migration path matters more than the default. The suggested route is the separate trinity-ops-public repository and its Trinity Ops Agent, specifically the /migrate-to-postgres skill, described as a gated validate-then-cutover flow. Validate first, cut over second, and the existing SQLite file is left untouched, so a rollback means pointing back at a file you still have rather than restoring a dump taken under pressure.

The test configuration takes the same distinction seriously. The pytest marker list includes requires_postgres, defined as an assertion that can only fail on PostgreSQL, which the schema-parity CI runs with TEST_POSTGRES_URL set under issue #2434. The gap between the two engines is treated as a place where tests are expected to fail on SQLite, not a difference to paper over.

Three secrets in .env, and a rotation that needs the old key

The environment template marks its first block as required for production and security critical, and the three values in it carry different blast radii.

SECRET_KEY signs JWTs and is generated with openssl rand -hex 32. CREDENTIAL_ENCRYPTION_KEY encrypts sensitive data including OAuth tokens and subscription credentials, uses the same generation call, and carries the warning that losing it makes every encrypted credential unrecoverable. INTERNAL_API_SECRET handles scheduler-to-backend communication and is tracked as C-003. If it is unset, it falls back to SECRET_KEY, and the template tells you not to rely on that fallback because it conflates two secrets with different blast radii. Issue #848 widened the reason: with MCP_INLINE_AUTH_ENABLED on, a holder of INTERNAL_API_SECRET can assert any verified email over the /api/internal/mcp-auth endpoints.

Rotation has its own documented path. CREDENTIAL_ENCRYPTION_KEY_SECONDARY is a decrypt-only fallback meant only for the rotation in issue #267: set it to the previous key while the primary holds the new one, run scripts/deploy/rotate-credential-key.py --apply, then delete the line. The procedure is written up in docs/migrations/CREDENTIAL_KEY_ROTATION.md, and the template says to leave the secondary empty in normal operation.

Agent containers miss the compose log cap

docker-compose.yml opens with a comment explaining a fix from issue #1871. Docker's json-file log driver has no maximum size and no maximum file count by default, so each container's log under /var/lib/docker/containers/ keeps growing until the Docker data root reaches 100% and the daemon wedges. The file defines an x-logging anchor on that driver with a size limit of ${CONTAINER_LOG_MAX_SIZE:-10m} and a file count of ${CONTAINER_LOG_MAX_FILE:-3}, then attaches the anchor to its services.

The agent containers are the awkward part. They are created through the Docker SDK rather than through compose, so the shared anchor never reaches them and they carry a separate cap defined by AGENT_LOG_CONFIG in src/backend/services/agent_service/capabilities.py. Anyone sizing disk usage from the compose file alone has to go find that second setting in the backend source, and the comment says so directly.

The pattern repeats elsewhere in the same file. The backend service takes six build arguments for provenance, VERSION, GIT_COMMIT, GIT_COMMIT_SUBJECT, GIT_COMMIT_TIMESTAMP, GIT_BRANCH and BUILD_DATE, each defaulting to unknown when unset so an out-of-band build still returns a well-formed /api/version response. The top-level tree carries seven compose files, including gitea, hosted, sibling and prod.enterprise variants, plus a packer directory, and the first AWS Marketplace AMI ships as the v1.0.0-aws.1 pre-release dated October 2, 2026.

The base compose file only gained a restart policy recently

The backend service in docker-compose.yml runs with restart set to unless-stopped, and the comment above that line explains why the setting is late. Issue #2541 describes the platform half of the same bug already fixed in the agent containers. The prod and hosted compose files carried the policy; the base file, the one scripts/deploy/start.sh uses without the hosted flag, which covers both the quickstart path and the source install, did not. The consequence was that a self-hosted control plane did not survive a host reboot. The value chosen keeps an operator's own docker compose stop meaningful, since unless-stopped still honours an explicit stop.

Provenance rides along in the build arguments. Issue #926 has scripts/deploy/start.sh export VERSION, GIT_COMMIT, GIT_COMMIT_SUBJECT, GIT_COMMIT_TIMESTAMP, GIT_BRANCH and BUILD_DATE from the local repo, using git rev-parse HEAD and similar calls, before compose builds run. Leaving them unset is safe rather than fatal: the Dockerfile then falls back to unknown and /api/version still answers with the right shape.

That is the operational character of this repository in miniature. Comments carry issue numbers and the reasoning behind each fix, defaults are written out explicitly instead of left to Docker, and the small operational details, log caps, restart policies, version strings, are treated as part of the product rather than as deployment folklore.

A 900 second pytest ceiling with the thread method

The pytest configuration is defensive for a Python project, and the comment above the timeout setting explains why. asyncio_mode is auto and pythonpath covers tests, src/backend and src, so backend code imports without an install step. The timeout is 900 seconds, and issue #2080 explains both the number and the method. The thread method is used rather than the default signal method because signal re-enters the interpreter from a signal handler, which on a hung socket read produced a pytest INTERNALERROR instead of a timeout, turning one blocking test into an indefinite stall. The ceiling is generous on purpose: per-tier timeouts in tests/run-full.sh are the real budget, and this value is only the backstop for someone running pytest by hand.

The marker list is where the project's promises are written down. There is smoke for fast tests that need no agent, requires_agent for tests that need a running agent, plus unit, integration and asyncio. Then two markers cover the harder claims: journey, defined as a promise the platform makes to a person, tested end to end against a live stack under issue #2335, and requires_postgres from issue #2434.

Read together with the compose files, that tells you what running Trinity well actually involves: a Docker host that survives reboots, a database chosen deliberately rather than by default, secrets generated per environment, log caps applied to containers that compose never sees, and a test suite split into tiers that each need something different running underneath them.

Editorial conclusion

Trinity fits teams that want agents running on infrastructure they control and who can accept the Docker and PostgreSQL obligations that come with it. Before deploying, verify three things: that every secret is generated per environment instead of inherited from a fallback, that agent containers pick up the log cap configured outside compose, and that the compose file your start script loads is the one carrying the restart policy.

Frequently asked questions

Does Trinity require PostgreSQL, or does SQLite work in production?

SQLite stays the zero-config default for local development and evaluation, but production instances should run PostgreSQL, opted into with a single DATABASE_URL. The setup walkthrough is kept in docs/POSTGRESQL_SETUP.md.

What happens to an existing SQLite database when a Trinity instance moves to PostgreSQL?

The /migrate-to-postgres skill in the Trinity Ops Agent repo runs a gated validate-then-cutover flow and leaves your SQLite file untouched, so rollback is a matter of pointing back at the file you still have.

Which coding agents can Trinity run?

The repository description names Claude Code, Codex and Gemini agents. Claude Code connects through the abilities plugins, which scaffold, connect and deploy agents over MCP.

Can Trinity be used commercially?

Yes. It is Apache 2.0, described as free for any use, commercial included, and deployable anywhere you run it. The project also states it was independently pentested to UnderDefense Grade A.

How do you stop agent container logs from filling the Docker data root?

Agent containers are created through the Docker SDK rather than compose, so the x-logging anchor with the 10m by 3 file defaults does not reach them. They are capped separately by AGENT_LOG_CONFIG in src/backend/services/agent_service/capabilities.py.

What has to be set in .env before Trinity runs in production?

SECRET_KEY, CREDENTIAL_ENCRYPTION_KEY and INTERNAL_API_SECRET, each generated with openssl rand -hex 32. The template warns against leaning on the INTERNAL_API_SECRET fallback to SECRET_KEY, because the two carry different blast radii.

Official sources

  1. Abilityai/trinity on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/abilityai-trinity.svg)](https://hysenlabs.com/projects/abilityai-trinity)