Model or dataset
ianarawjo/ChainForge avatar
ianarawjo/ChainForge

ChainForge: a visual environment for comparing prompts across LLMs

An open-source visual programming environment for battle-testing prompts to LLMs.

3,032 stars258 forksTypeScriptMIT

At a glance

What is it?
ChainForge is an open-source, MIT-licensed visual programming environment for sending the cross product of prompt inputs to multiple LLM providers and scoring the responses. This article covers how it installs, how the data flow works, and where it stops being the right tool.
Who is it for?
Adopt ChainForge if you need to see how a prompt template behaves across many input values and several providers before committing to one, and if a local Flask app on port 8000 fits your workflow. Do not adopt it as a production evaluation harness for a CI pipeline: the README describes a desktop-style visual tool, not a headless runner.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem ChainForge solves, and for whom

Ad-hoc chat with one model answers one question at a time. If you want to know whether a prompt template survives a different argument, a different model, or a different temperature, you end up copying text between browser tabs and keeping results in a scratch file. ChainForge replaces that with a data flow graph: you place nodes for prompts, models, and evaluation, wire them together, and press run. The README describes the target as "rapid-fire, quick-and-dirty comparison of prompts, models, and response quality that goes beyond ad-hoc chatting with individual LLMs."

The audience is prompt engineers, applied researchers, and product teams who are choosing between providers or between prompt variants and want evidence rather than a feeling. The repository also ships example flows, including flows generated from OpenAI evals benchmarks, so a new user has something concrete to open before writing their own graph. The project is written in TypeScript on the front end and Flask on the back end, and is licensed MIT.

The cross product is the core mechanism

The README names the central idea directly: "ChainForge takes the cross product of inputs to prompt templates, meaning you can produce every combination of input values." A Prompt Node holds a template with variables such as {game}. A Tabular Data node supplies a column of values. Wiring the column into the template variable produces one prompt per row. Add a second variable and you get the product of both columns, which is how a handful of rows becomes hundreds of queries.

Responses land in a table you can inspect and export. Evaluation happens through scoring functions: the README describes setting up evaluation metrics and visualizing results across prompts, prompt parameters, models, and model settings. For ground-truth work, the README shows importing a dataset, hooking it to a template variable, and comparing each response against an expected answer. Locally installed instances can also run Python code to evaluate responses, which is the escape hatch when a built-in scorer does not fit.

Providers are pluggable. The README lists OpenAI, Anthropic, Google Gemini, DeepSeek, HuggingFace Inference and Endpoints, Together.ai, Ollama, Microsoft Azure OpenAI, Aleph Alpha, and Amazon Bedrock, with custom provider scripts for anything else. Ollama matters here because it means the comparison can include a locally hosted model without sending data to a vendor.

Installing ChainForge and running a first comparison

The README requires Python 3.8 or higher. Installation is a single pip command, and the server starts on port 8000.

bash
pip install chainforge
bash
chainforge serve

Open http://localhost:8000 in Chrome, Firefox, Edge, or Brave. The README notes the web version at chainforge.ai/play has a limited feature set, so a local install is what you want if you plan to load API keys from environment variables, write Python scoring code, or query Ollama models.

Keys can be entered through the Settings icon in the top-right corner, but the README recommends saving them to your local environment instead so you do not retype them each session. The Docker path builds the repository Dockerfile and maps the same port.

bash
docker build -t chainforge .
bash
docker run -p 8000:8000 chainforge

After the container starts, the README says to open http://127.0.0.1:8000. The docker-compose.yml in the repository goes further: it mounts a named volume at /home/chainforge/.local/share/chainforge for persistent storage, sets NODE_ENV and FLASK_ENV to production, and includes a healthcheck that curls http://localhost:8000 every 30 seconds with a 40-second start period. The commented environment block lists OPENAI_API_KEY, ANTHROPIC_API_KEY, COHERE_API_KEY, GOOGLE_API_KEY, DEEPSEEK_API_KEY, and HUGGINGFACE_API_KEY, which tells you the intended way to pass credentials to a container.

For a first real run, open the Example Flows button in the top-right corner and pick a comparison flow. The README's basic example plots response length across different models and arguments for the {game} prompt parameter. Run it, then swap in your own template variable and watch the response table fill.

Where the visual approach costs you

A graph editor is a poor fit for anything you want to run unattended. The README documents a server you open in a browser, example flows you click, and a Share button for links. It does not document a headless command that executes a saved flow and exits with a status code, which is what a CI job would need. If your requirement is a nightly regression suite that fails a build, ChainForge is the wrong shape.

The Share feature has explicit limits that matter if you treat it as storage. The README states you can share up to 10 flows at a time, each under 5MB after compression, and that sharing an eleventh link breaks the oldest one. It tells you to export important flows to cforge files and use Share "to only pass data ephemerally." Anyone who treats a share link as a durable artifact will lose work.

Provider coverage is broad but not uniform. Bedrock support is described as on-demand inference, and HuggingFace is split between Inference and Endpoints. A model that only exists behind a bespoke internal gateway needs a custom provider script, which is Python you now maintain. The README does not document rollback or version pinning for flows, so a graph that depends on a specific model version can change behavior underneath you when the provider updates that model.

ChainForge compared with writing evaluation scripts by hand

The obvious alternative is a Python script: a loop over your prompt variants, a client per provider, and a pandas DataFrame for results. That approach wins on repeatability. It lives in version control, runs in CI, and produces a diffable artifact. It loses on the first hour. Standing up multi-provider auth, rate limiting, retries, and a table view is real work before you learn anything about your prompts.

ChainForge inverts that trade. You get the comparison surface immediately, including the response inspector and table export, and you pay later when you want to automate. The two are not exclusive: a common pattern is to explore in the graph editor until a prompt variant looks promising, then port the winning configuration into a script for the regression suite. The README's mention of prompting a model to generate starter code for evals points at that same handoff.

A second alternative is a hosted prompt playground. Those usually give you one prompt and one model at a time with no cross product and no scoring function, which is exactly the gap the README claims ChainForge fills. The trade you accept with ChainForge is running a local Flask process and managing your own keys.

Maintenance, licensing, and upgrade cost

The repository is not archived, and the last push was on 2026-09-02, which is recent enough that the project is being touched. The most recent tagged release listed is v0.3.6 from 2025-05-11, which added a Media Node and image inputs; v0.3.5 added favorites, an Ollama list, encryption, and persistent settings. The setup.py in the repository declares version 0.3.7.0, ahead of the newest tag, so the version string you see after installing may not match a release name you can search for. That mismatch is worth knowing before you file a bug.

Dependency weight is the real upgrade cost. The core install list is modest: flask, flask_cors, requests, platformdirs, urllib3, openai, cryptography, mistune, and markitdown with pdf, docx, xlsx, xls, and pptx extras. The optional rag extra is heavy, pulling grpcio, numpy, pymupdf, python-docx, tiktoken, nltk, transformers, scikit-learn, sentence-transformers, rank-bm25, whoosh, cohere, chonkie, model2vec, pyarrow, lancedb, pandas, tqdm, and accelerate. A comment in setup.py warns against capping pyarrow or lancedb because older pyarrow ships no wheels for Python 3.13 and pip falls back to a source build that fails. If you install the rag extra on a new Python, expect that class of problem.

The license is MIT, which permits commercial use and modification with the copyright notice retained. That is a permissive starting point, but it says nothing about the terms of the model providers you query through it; your obligations to OpenAI, Anthropic, or Bedrock come from those agreements, not from ChainForge. This is not legal advice, and the LICENSE.md file in the repository is the authoritative text.

Editorial conclusion

Adopt ChainForge if you need to see how a prompt template behaves across many input values and several providers before committing to one, and if a local Flask app on port 8000 fits your workflow. Do not adopt it as a production evaluation harness for a CI pipeline: the README describes a desktop-style visual tool, not a headless runner. Before relying on it, verify that your provider keys load from the environment, that the flows you care about survive an export to a cforge file, and that the Share links you depend on are under the ten-flow limit.

Frequently asked questions

How do I use ChainForge?

Install with pip install chainforge on Python 3.8 or higher, then run chainforge serve and open http://localhost:8000 in Chrome, Firefox, Edge, or Brave. A Docker path also exists: build the repository Dockerfile with docker build -t chainforge . and run it with docker run -p 8000:8000 chainforge.

What is LLM chaining?

LLM chaining usually means feeding one model's output into the next call as input. ChainForge is not that: it is a data flow environment for comparing prompts, models, and settings, and its defining operation is the cross product of inputs to a prompt template.

What is chain of thought in simple terms?

Chain of thought is a prompting technique where the model is asked to reason step by step before answering. ChainForge does not implement it as a node; you would express it in the text of a Prompt Node and then compare the resulting responses across models and settings.

Can you provide an example of chain of thought reasoning?

The README does not give a chain of thought example. It gives comparison examples instead, such as plotting response length across different models and arguments for the {game} prompt parameter, and comparing each model's answer to a math problem against an expected answer using Tabular Data nodes.

Official sources

  1. ianarawjo/ChainForge on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ianarawjo-chainforge.svg)](https://hysenlabs.com/projects/ianarawjo-chainforge)