forge-guardrails: a reliability layer for self-hosted LLM tool-calling
A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows
At a glance
- What is it?
- forge is a Python framework that sits inside one agentic loop and makes local-model tool calls hold together. It is not an orchestrator, and the eval numbers in its README are the author's own, not independently reproduced.
- Who is it for?
- Adopt forge if you already run a local model server and your problem is malformed tool calls, missing steps or context growth inside a single agent loop, not multi-agent coordination. Skip it if you need a DAG planner or a coding harness, since the README puts both out of scope.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 31 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem forge-guardrails actually solves
Local models call tools badly. An 8B model that answers questions well will still emit a tool call with the wrong argument name, skip a step it was told to take, or return prose where JSON was expected. That failure is not a reasoning failure, it is a parsing and validation failure, and it is the one forge targets. The README frames the project as "a reliability layer for self-hosted LLM tool-calling": you hand it a set of tools, the model picks whichever it wants in whatever order, and forge's guardrails (rescue parsing, retry nudges, response validation) clean up what comes back.
The intended user is a developer running a model on their own hardware, usually through llama-server, Ollama, Llamafile or vLLM, who wants tool-calling to stop being flaky without moving to a hosted API. The README also lists Anthropic as a backend, so the same layer can wrap a hosted model. The scope is deliberately narrow: forge sits inside one agentic loop. Multi-agent graphs, DAG planners and cross-agent coordination are named as out of scope, and so is being a coding harness. If you want a planner that decomposes a task across five agents, this is the wrong layer.
How the guardrails, context manager and backends fit together
Three pieces do the work. A backend client talks to the model server. A ContextManager decides what stays in the prompt. A WorkflowRunner drives the loop, injecting the system prompt, executing the tool the model asked for, feeding the result back, and applying guardrails at each turn.
The guardrails are the part that distinguishes forge from a thin HTTP wrapper. Rescue parsing tries to recover a usable tool call from a response that is not clean JSON. Retry nudges re-prompt the model when it produced nothing actionable. Response validation checks the call against the tool's Pydantic parameter model before anything is executed. Workflow structure is opt-in: required_steps, prerequisites and terminal_tool constrain the loop when you need determinism, but the README states that the guardrails apply with zero required steps too. That is the design bet, and it is a reasonable one, because it means you can add forge to an existing free-form loop without first rewriting the loop as a state machine.
Context is handled by strategies rather than a fixed window. The Quick Start uses TieredCompact with keep_recent=2 and a budget_tokens of 8192, which implies older turns are compacted in tiers once the budget is approached. The README does not spell out the tier boundaries, so treat the exact compaction behaviour as something to read in the source or docs before you rely on it for long conversations. Backend support covers generic OpenAI-compatible endpoints, Ollama, llama-server, Llamafile, vLLM and Anthropic. SlotWorker is a separate piece: priority-queued access to a shared inference slot with auto-preemption, aimed at multi-agent setups where specialist workflows share one GPU.
Installing forge-guardrails and running a first tool call
There are two install paths and they are not interchangeable. The Python package is what you want for WorkflowRunner, the guardrails middleware or a Python-managed proxy. It requires Python 3.12 or newer and a running LLM backend. The package name on PyPI is forge-guardrails, not forge.
pip install forge-guardrails # core only
pip install "forge-guardrails[anthropic]" # + Anthropic clientThe README is explicit that the Python package does not install a global forge-proxy command. That command belongs to the standalone installer, which bundles forge, a private Python runtime and the Anthropic SDK so the host needs no Python or pip. Run the Python package's proxy with python -m forge.proxy instead.
curl -fsSL https://raw.githubusercontent.com/antoinezambelli/forge/main/install.sh | shThat is the Linux and macOS standalone path. Windows PowerShell uses irm on install.ps1 piped to iex. After refreshing the terminal, forge-proxy init creates a profile and forge-proxy check validates it. Note what the standalone distribution does not do: it does not install a backend executable, a model, a GPU stack, credentials or client configuration.
Start a backend before running anything. The README recommends llama-server and notes that the top 10 eval configs all run on it.
llama-server -m path/to/Ministral-3-8B-Instruct-2512-Q8_0.gguf --jinja -ngl 999 --port 8080Then a minimal workflow defines one tool with a Pydantic parameter model, marks it as terminal_tool, and runs it through WorkflowRunner with a LlamafileClient and a TieredCompact context manager at budget_tokens=8192. The example in the README asks for the weather in Paris and expects the model to call get_weather with city set. The callable itself is ordinary Python, so the tool result is whatever your function returns.
Where forge-guardrails stops being the right tool
The README's own scope statement is the honest limitation: forge is not an agent orchestrator and not a coding harness. If your problem is task decomposition across agents, forge will not help, and SlotWorker only addresses GPU contention between workflows that already exist, not how they are planned.
The second limitation is the eval evidence. The headline claim, an 8B local model going from single digits to 84% across a 26-scenario v0.7.0 suite, is the project's own measurement, and the README itself notes that the Anthropic numbers were measured in v0.6.0 and not re-run in v0.7.0 because the cost is non-trivial. The repository does carry eval_results files for v0.6.0 through v0.9.0, so the raw data is inspectable, but there is no independent replication. Treat 84% as a number tied to a specific suite, model and configuration, not a general property of the library.
Third, the package is classified as Development Status 4 - Beta at version 0.9.5. The recent release history shows small compatibility and hotfix releases, including a proxy command ownership hotfix in v0.9.3, which is normal for a young project but means the proxy lifecycle and command ownership have changed recently. If you pin an older release, check the changelog before assuming the current install instructions match. Finally, the README does not document rollback behaviour for the standalone installer or for context compaction, so verify those yourself before putting either in a path you cannot easily reverse.
Forge Proxy versus wiring the library in yourself
The real alternative is not another library, it is the integration path. You can point an existing OpenAI-compatible client (the README names opencode, Continue and aider) or Claude Code at the standalone forge-proxy and get guardrails without touching the client. The proxy speaks both the OpenAI chat-completions API and the Anthropic Messages API at /v1/messages, so the client believes it is talking to a smarter model. The Dockerfile shows the packaged entry point: python -m forge.proxy on host 0.0.0.0, port 8081, with a health check at /forge/health.
The other path is to import forge directly and own the loop. WorkflowRunner manages system prompts, tool execution, context compaction and guardrails for you. The third path, documented in examples/foreign_loop.py, is to keep your own orchestration loop and drop in forge's middleware for response validation, malformed-call rescue and required-step enforcement.
The trade-off is control versus surface area. Proxy mode is the smallest change and the README calls it the most popular entry point, but it puts a process between your client and the model, and the standalone installer owns that process's update and uninstall lifecycle. Importing the library gives you the tool definitions and context strategy in your own code, at the cost of writing the runner setup. Middleware gives you the most control and the least help. Pick based on how much of the loop you want to keep.
Maintenance, licence and upgrade cost
The repository is not archived and the last push was on 2026-09-01, so it is being worked on. Releases are frequent and small: v0.9.5 on 2026-08-30, v0.9.4 on 2026-08-26, v0.9.3 on 2026-08-22. Version numbers in the 0.9.x range with hotfix releases suggest the API is still settling, so pin a version in production rather than tracking main.
Licensing is MIT, declared in pyproject.toml, the LICENSE file at the repository root and the Docker image label. MIT permits commercial use and modification with the copyright notice retained. That is the whole of what the repository supports; questions about your own distribution obligations belong with a lawyer, not with this article.
Upgrade cost splits by install path. The Python package upgrades through pip and its dependencies are small: pydantic, httpx and tomli-w, with anthropic optional. The standalone proxy has its own update and uninstall lifecycle, and the README points to docs/PROXY_INSTALLATION.md for exact-version installation, profiles, updates, recovery and uninstall. Because the Python package deliberately does not install forge-proxy, a host that has both the pip package and the standalone binary has two things named similarly with one owner. Keep them separate.
Editorial conclusion
Adopt forge if you already run a local model server and your problem is malformed tool calls, missing steps or context growth inside a single agent loop, not multi-agent coordination. Skip it if you need a DAG planner or a coding harness, since the README puts both out of scope. Verify first that your backend speaks the OpenAI-compatible or Anthropic Messages API, that your Python is 3.12 or newer, and that you are willing to run with required_steps left empty, because forge's guardrails apply with zero required steps too.
Frequently asked questions
What does forge-guardrails require to run?
The Python package requires Python 3.12 or newer and a running LLM backend. Backends listed in the README are generic OpenAI-compatible endpoints, Ollama, llama-server, Llamafile, vLLM and Anthropic.
Is forge-guardrails free to use?
It is published under the MIT licence, which permits commercial use and modification provided the copyright notice is retained. The README does not describe any paid tier.
How do I install forge-guardrails?
For the Python library, run pip install forge-guardrails, or pip install "forge-guardrails[anthropic]" to add the Anthropic client. For the standalone proxy on Linux or macOS, the README gives a curl install script, and Windows PowerShell uses install.ps1.
Does forge-guardrails install a forge-proxy command?
No. The README states the Python package intentionally does not install a global forge-proxy command and that the standalone installer is the sole owner of that command and its update and uninstall lifecycle. From the Python package, run the proxy with python -m forge.proxy.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/antoinezambelli-forge)