# TrueForge: a runtime layer for LLM agents, shipped as a TypeScript server

> TrueForge is an MIT-licensed agent harness from TrueFoundry that owns the execution loop around a model: MCP tools, skills, sandboxing, approvals and session state. It runs in one process with SQLite, or against Postgres and Redis for a team deployment.

**truefoundry/trueforge** — The open-source agent harness - the runtime layer that turns an LLM into a working agent.

- Repository: https://github.com/truefoundry/trueforge
- Website: https://trueforge.dev
- Stars: 6,030 · Forks: 486
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/truefoundry-trueforge

## The gap TrueForge targets: an agent that runs, not one that demos

A tool-calling loop is maybe fifty lines. What surrounds it is not. Streaming responses, keeping session state across turns, connecting remote tool servers, running generated code somewhere isolated, pausing for a human approval, and trimming context before it overflows. Most teams write that scaffolding themselves, once, badly, and then maintain it forever.

TrueForge takes the position that this scaffolding is the product. The README describes it as "the runtime layer that turns an LLM into a working agent", and the repository layout backs that up: a workspace of packages rather than a single library, plus a Dockerfile, a Helm chart under charts/trueforge, and a docker-compose.yml. It is aimed at engineers who already know which model they want and would rather not rebuild the loop around it.

The intended audience is narrower than "anyone building with LLMs". If your agent is one prompt and one function call, this is a service to deploy where a script would do. The fit is a team that needs a shared, inspectable runtime with a UI on top, and is prepared to operate it.

## Inside the loop: catalogs, MCP servers, skills and a sandbox tool

The architecture diagram in the README shows a single shape: a chat UI and SDKs talk to the TrueForge server over HTTP, the server runs the agent loop, and that loop reaches out to SQLite or Postgres on one side and to your models, MCP servers and a sandbox on the other. Everything the agent can do is configured before the agent exists.

That configuration is catalog-driven. Models, MCP servers, skills and the sandbox are set up once, and agents pick from what has been connected. The README says presets come from shipped YAML catalogs you can customize, which is a meaningful design choice: the connection details live in files you can review and diff, not in per-agent code. Model access is not restricted to one vendor. The README lists OpenAI, Anthropic and Google Gemini among catalog providers, and any OpenAI-compatible endpoint is accepted.

Tools come from remote MCP servers, with header auth or OAuth, including authorization performed inside the chat. Skills are git-backed SKILL.md instruction packs loaded on demand inside the sandbox. The sandbox itself is exposed to the model as a tool, and the README states it is provisioned only when needed, with Daytona as the provider today and more planned. One detail worth pausing on: secrets stay in the harness rather than being handed to the sandbox, which is the right default but also means the sandbox cannot authenticate to anything on its own.

Context management is treated as a first-class concern rather than an afterthought. The README names subagents, deferred tool loading, Code Mode, large-result offloading and compaction. Human checkpoints appear as tool approval, ask-user-questions, and Generative UI in the chat. Together these are the parts teams usually discover they need only after an agent has already misbehaved in production.

## Installing TrueForge locally and getting to a first agent turn

The fastest path is the published npm package, which needs Node.js 22.14 or newer according to the README badge and the engines field in package.json.

```bash
npx @truefoundry/trueforge@latest
```

That starts the server and the bundled chat UI. The README's quickstart guidance points at trueforge.dev/quickstart for the other supported methods: local, Docker Compose, Kubernetes and Railway. Local mode is one process with SQLite and no extra infrastructure, which is why it is the sensible first stop.

If you would rather run the full stack the way the repository does, the compose file builds from Dockerfile.dev and serves the API and UI on a single port. Note the port mapping: the container listens on 8790 and compose maps it to 8791 on the host so it does not collide with a local `pnpm dev` on 8790.

```bash
docker compose up
```

After the server is up, connect a model, then an MCP server or a skill, and create an agent from what you connected. The README's own ordering is explicit about this: configure models, MCP servers, skills and a sandbox once, and agents pick from what you connected. The first thing to look at is not the chat window but the catalog files, because that is where the agent's capabilities are actually defined.

For a production image, the root Dockerfile installs a published version from npm rather than building the monorepo. It requires an APP_VERSION build argument and fails the build if it is missing.

```bash
docker build --build-arg APP_VERSION=0.1.0 -t trueforge:0.1.0 .
```

The comment in that Dockerfile is worth reading before you rely on it: the image installs the published package, so what you deploy matches the npm release rather than floating workspace source.

## Local mode is a laptop mode, and the README says so plainly

The most important limitation is stated by the project itself, in a block quote under the architecture table. Local mode is for your machine only. There is no login by default, and data lives in a local SQLite file. The README asks that it be kept on localhost and says the maintainers cannot take responsibility for data loss or unauthorized access if it is used beyond that.

That is an unusually direct warning, and it should be read as a boundary rather than a disclaimer. An agent harness executes tools and code on behalf of a model. Running one without authentication on a reachable interface is a different risk class from running a read-only dashboard the same way.

The second limitation is operational. Hosted mode requires Postgres and Redis, and the compose file explains why Redis is there: it is shared by all server replicas and carries executor peering, described as cross-replica turn cancel via request-reply. So the moment you want more than one replica, you are running two stateful dependencies, not one. The compose file also keeps peering enabled even with a single replica, specifically so that path stays exercised.

Third, the sandbox provider is a single option. The README names Daytona and says more providers are planned. If your environment cannot use Daytona, the sandbox-as-a-tool feature is not available to you today, and the README does not document a fallback.

Finally, the release list shows 0.2.0 candidates: @truefoundry/trueforge@0.2.0-rc.3, @truefoundry/trueforge-ui@0.3.0-rc.3 and charts/trueforge@0.2.0-rc.0, all dated 2026-09-10. The last push to the repository was on 2026-09-10 as well. Nothing here suggests abandonment, but release candidates are release candidates, and the README does not document a rollback procedure for a version you have already deployed.

## How TrueForge differs from wiring up the provider SDK yourself

The obvious alternative is not another harness. It is the vendor SDK you already have, plus your own glue: the OpenAI or Anthropic client library, a database table for messages, a tool dispatcher, and a container you shell out to for code execution. That approach is completely reasonable and it is what most teams start with.

The difference is where the work sits. With a provider SDK, the loop is code in your repository, and every capability you add (streaming resumption, tool approval, context compaction, a second model provider) is a change to that code. With TrueForge, the loop is a server you run, and those capabilities are configuration plus an HTTP surface. The README frames the same trade-off from the other direction: it exposes the loop three ways, as a chat UI, an HTTP API with a TypeScript SDK, and an embeddable UI SDK.

That framing also tells you what you give up. You inherit the harness's opinions about how sessions are stored, how tools are approved, and how context is compacted. If your product needs a loop shape the harness does not have, you are working around a server rather than editing a function. The catalog model has the same character: it makes configuration reviewable, but it also means a capability has to exist as a catalog entry before an agent can use it.

There is a middle path worth naming. If you only want the UI, @truefoundry/trueforge-ui is published separately and described as embeddable, so the chat surface can be adopted without taking the whole runtime.

## Licence, upgrades and what maintenance actually costs

TrueForge is MIT licensed, per the repository metadata and the badge in the README. That is permissive: you can use it commercially, modify it, and ship it inside a closed product, provided the licence and copyright notice are preserved. This is not legal advice, and the LICENSE file in the repository is the text that governs.

The upgrade surface is larger than a single package. The workspace publishes @truefoundry/trueforge, @truefoundry/trueforge-sdk, @truefoundry/trueforge-ui and @truefoundry/trueforge-core, and the releases listed for 2026-09-10 show the server and UI versioned independently (0.2.0-rc.3 against 0.3.0-rc.3). The chart has its own version line. If you deploy the Helm chart, you are tracking three version numbers, and the repository uses changesets with a dedicated script, changeset:sdk-regen, which implies the SDK is regenerated as part of the release process rather than hand-edited.

The production Dockerfile adds a constraint that is easy to miss. It installs a pinned npm version and the comment states there is no workspace fallback, so a version that is not on the registry fails the build rather than silently building from source. That is the correct failure mode, but it means your image build depends on registry availability at build time.

Running from source is a separate path. The workspace uses pnpm 11.16.0 and Node 22.14.0 or newer, and the build script chains core, sdk, ui, frontend and the server in that order. The repository also carries docker-compose.dev.yml and a dev:infra script for bringing up local infrastructure, so contributing does not require assembling the dependencies by hand.

## Conclusion

Adopt TrueForge if you want the agent loop, session storage and a chat UI already assembled, and you are willing to run a Node 22.14 or newer service. Do not adopt it for a single-script prototype, and do not expose local mode to a network: the README states there is no login by default and that data lives in a local SQLite file. Before committing, verify that your model provider appears in the catalog or speaks the OpenAI-compatible protocol, that Daytona is acceptable as the sandbox provider, and that the 0.2.0 release candidates are a version you can pin.

## FAQ

### What is TrueForge?

It is an open-source agent harness: the runtime layer that runs the agent execution loop, covering model calls, MCP tools, skills, sandboxing, approvals, context management and session state. It exposes that loop as a chat UI, an HTTP API with a TypeScript SDK, and an embeddable UI SDK.

### How do I install TrueForge?

The README's quickstart is to run npx @truefoundry/trueforge@latest, which requires Node.js 22.14 or newer. The documentation also covers Docker Compose, Kubernetes and Railway, and the Dockerfile builds a production image from a published npm version via an APP_VERSION build argument.

### Can TrueForge run in production?

Hosted mode is the production path: Postgres for storage plus Redis, run through Docker Compose, Helm or Railway. The README states that local mode is for your machine only, has no login by default, stores data in a local SQLite file, and should be kept on localhost.

### Which model providers does TrueForge support?

The README lists OpenAI, Anthropic and Google Gemini among catalog providers, and states that any OpenAI-compatible endpoint is also accepted. Models are configured in catalogs, and agents pick from what you have connected.

### What does TrueForge use for sandboxed code execution?

The sandbox is exposed as a tool and is provisioned only when needed. The README names Daytona as the provider today and says more providers are planned, with secrets kept in the harness rather than passed into the sandbox.

## Sources

- [License: MIT](https://github.com/truefoundry/trueforge/blob/main/LICENSE)
- [Project website](https://trueforge.dev)
- [README](https://github.com/truefoundry/trueforge/blob/main/README.md)
- [Releases](https://github.com/truefoundry/trueforge/releases)
- [truefoundry/trueforge on GitHub](https://github.com/truefoundry/trueforge)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/truefoundry-trueforge
