Open-source project
zonghaoyuan/infiplot avatar
zonghaoyuan/infiplot

infiplot: the next scene is already painted before you choose

InfiPlot is the world's first interactive plot game that AI generates all text and images in real-time. InfiPlot是全球首个在游玩过程中由 AI 实时生成全部图文内容的互动剧情游戏

381 stars56 forksTypeScriptAGPL-3.0

At a glance

What is it?
An AI story game with no preset plot, where four agents write, design characters, dress scenes and paint backgrounds, the engine renders the scenes your options lead to before you pick one, and clicking anywhere on the image routes through a vision model. Running it costs image generations, about $0.00078 each.
Who is it for?
InfiPlot is an interesting demonstration of predictive generation, and the cost arithmetic tells you what that means. Image generation dominates the bill at roughly $0.00078 a scene, and the engine deliberately generates scenes you may never see, so the spend exceeds the scenes you watched.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 87 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four agents, four jobs, and one of them owns the plot

The framing is that InfiPlot is an AI version of a Chinese interactive romance game, the kind where you play through scenes rather than through a menu. There is no preset plot and no preset cast; everything is generated to order.

The examples of what a playthrough can be are deliberately the reader's own: studying magic in the Harry Potter world, being the person everyone at school admires, publishing without stopping and funding never running out, reliving palace intrigue, or going back to youth to make different choices about things you regret.

The mechanism behind that is a multi-agent framework over text, image and audio models, with the work split four ways. A screenwriter, a character designer, a scene dresser and a painter each take one responsibility, and they coordinate so the plot stays coherent, the characters stay consistent and the scenes stay consistent.

The screenwriter carries an extra duty: it is also responsible for the overall plot architecture.

That last clause matters. Consistency across a long playthrough is not something an image model or a dialogue model can enforce alone, so someone has to own the shape of the story, and giving that job to the writing agent is the cheapest place to put it.

A beat tree keeps the image still until the story moves

The scene model is worth stating precisely, because it decides what the player sees.

The whole experience of one playthrough is called a story. A story unfolds as a sequence of scenes. Each scene is one AI-painted background image plus a short beat tree, which is narration, dialogue and the occasional option.

Clicking through the beats of a scene does not change the image. A new picture is painted only when an option takes you somewhere genuinely different: a different space, a different point of view, or a jump in time.

That is the cost and the pacing decision in one sentence. Dialogue inside a location is cheap because the image already exists, and the image is redrawn only when the story leaves the room. A scene that is pure conversation costs one generation regardless of how many lines of dialogue it holds.

It also explains the visual grammar: the background is the invariant of a scene, and the beats are what move within it. A player who stays in one place watches the picture hold still, which is also what makes it obvious when the story has actually moved.

Predictive generation is why the transition feels instant

While you are reading a scene, the engine is already generating the scenes your options lead to. For an option you cannot avoid, it goes one step further and renders the scene after that too.

When you finally choose, the picture is normally already there, so the switch completes instantly with no pause. The page adds that if you still feel some latency, work on it is continuing.

The mechanism is straightforward and the cost is stated later in the document rather than here, which is the right place for it: because the engine speculatively renders scenes you might choose and end up not choosing, real spending is a little higher than the number of scenes you actually saw.

Put together, that means the experience is paid for by predicted branches rather than by watched scenes. It is the same trade every predictive system makes, and it is worth knowing before you point it at a metered image API.

One detail that reduces the risk: stepping through beats is free, so a long conversation in one location adds nothing.

Clicking the background goes through a vision model

Buttons are not the only way to interact. Clicking the background itself, rather than a button, is handled by a vision model.

It reads where you clicked and decides what you meant. If you are exploring the current scene, it inserts a beat, and no new image is generated. If you are continuing forward, it generates a new scene.

The distinction is the whole design: exploration is cheap and immediate, progress costs a generation, and the model decides which one you meant from a coordinate.

The page is candid about the provenance and the ambition. This understanding came from working with Flipbook, and the authors believe it will become the key feature of InfiPlot, because it removes the need to guess which button a player wanted.

The stated direction of travel is that no traditional game interface will be baked into the picture at all. The AI paints the entire world in the style you chose, and the dialogue box and option buttons become a lightweight HTML layer laid over it and tuned to fit the scene, so the interface fits each story rather than staying fixed.

Nine variables, three required providers, speech optional

InfiPlot talks to four kinds of model provider, and the configuration table is where the cost and the choice live.

Text is described as the plot director, with a base URL, an API key and a model name, required, with DeepSeek's deepseek-v4-flash recommended. Image is the scene renderer, also required, with Runware's runware:400@6 running FLUX.2 klein recommended. Vision is the click interpretation layer, required, with Google's gemini-3.5-flash recommended. Speech is the character voice, optional, with MiMo's mimo-v2.5-tts free and StepFun's step-tts-2 as the paid alternative.

So nine variables are mandatory and three more are not. Leaving the speech variables empty runs the game silently, which is a sensible default for a first run: you can check that the plot works before adding a voice layer.

Where they are set depends on the platform. Locally it is .env.local. On Vercel it is Project Settings, then Environment Variables. On Cloudflare Workers you run wrangler secret put with the variable name at the repository root, or set it in the dashboard under the Worker's variables and secrets. For a staging deployment that needs a gate, the suggestion is to put Cloudflare Access in front of the Worker.

Two wire formats survive, and the URLs are normalized for you

The protocol layer is simpler than the provider list suggests, because most providers are reached through an OpenAI-compatible endpoint.

An optional provider variable per model type selects the wire format. openai_compatible is the default and covers text, vision and image, using OpenAI Chat Completions and the images generation endpoint. openai selects OpenAI gpt-image for images, with reference image editing. runware selects the Runware task-array protocol for images.

Text and vision support openai_compatible only. Gemini is reached by pointing the base URL at its OpenAI-compatible endpoint. Claude is recommended to go through a compatible gateway such as LiteLLM, and the reason given is cost and latency: Anthropic's own OpenAI-compatible layer does not support caching.

Image auto-detects from the base URL, treating a runware.ai host as Runware and anything else as OpenAI-compatible, with a documented exception for models on that host which are served over the OpenAI protocol and can be forced with the provider variable. Base URLs also tolerate a missing or extra /v1, or a trailing chat/completions segment, because the engine normalizes them.

That last detail is the one that saves time. Most integration bugs with configurable endpoints are a trailing path segment, and handling it in the engine rather than in the documentation is the right place for it.

Images are the bill, and MOCK_IMAGE exists to avoid it

The cost section is short and specific.

With the recommended combination, the cost of a scene is dominated by image generation, and one FLUX.2 klein image is about $0.00078 at 1792 by 1024, four steps, sub-second. The text model at deepseek-v4-flash is described as very cheap. Stepping through the beats of a scene is free.

The caveat follows directly from predictive generation: the engine renders scenes you might choose and never do, so the real bill runs slightly above the number of scenes you watched.

There is one switch created for exactly this problem. Setting MOCK_IMAGE to true skips image generation entirely and has the renderer return a static placeholder, while the plot, the voice and the options run normally. It is described as the way to debug the speech layer without consuming image credit.

That is a good design for a metered pipeline: it lets you validate three of the four provider integrations at zero cost, and leaves only the one you actually pay for as the last step.

A second optional piece exists for a browser-specific failure: when images load progressively, which happens when Chrome renders them line by line on some networks after a protocol error, you can deploy a very small Cloudflare Worker to forward images server-side. The default needs no configuration at all, since the browser connects to the image provider directly.

One repository, three platforms, and a prebuilt image for the fourth

Deployment is where the project is most accommodating, and the differences between hosts are stated rather than left to be discovered.

For personal use, Vercel is the one-click path, and the repository root is the application itself, so no root directory needs setting. OpenDeploy exists for letting an agent do the deployment for you. Cloudflare needs a Workers Paid Plan, and the reason is specific: the scene pipeline requires longer CPU time.

For your own server, Docker is the path, and you do not clone anything. Two files are fetched:

bash
mkdir -p infiplot && cd infiplot
curl -fsSL https://raw.githubusercontent.com/zonghaoyuan/infiplot/main/docker-compose.yml -o docker-compose.yml
curl -fsSL https://raw.githubusercontent.com/zonghaoyuan/infiplot/main/.env.example -o .env.example
[ -f .env.local ] || cp .env.example .env.local

The compose file is one service on port 3000 with the environment read from a local env file, so starting it is a single command and the game is at localhost:3000. Running the image directly works too, with the env file passed in and a published tag from the GitHub container registry. Both x86 and ARM are supported, including Apple Silicon.

The image itself is built from a Node 22 Alpine base with a non-root user, and the web build sets a standalone output flag so the runtime stage copies only what it needs.

One dependency explains a class of bug: jsonrepair is in the runtime list, which means the model output is repaired when it does not parse. The package is licensed AGPL-3.0-only, and the last push is dated 8 July 2026 with no GitHub releases.

Editorial conclusion

InfiPlot is an interesting demonstration of predictive generation, and the cost arithmetic tells you what that means. Image generation dominates the bill at roughly $0.00078 a scene, and the engine deliberately generates scenes you may never see, so the spend exceeds the scenes you watched. If you want to try it, set the three required providers and leave the speech ones empty for a silent run, or turn on MOCK_IMAGE to debug without spending credit. Before hosting it yourself, pick a platform deliberately: Vercel needs no configuration, Cloudflare needs the paid Workers plan because the scene pipeline wants longer CPU time, and Docker pulls a prebuilt image so you never need the repository itself.

Frequently asked questions

How does infiplot generate a story?

Four agents divide the work: a screenwriter that also owns the overall plot architecture, a character designer, a scene dresser and a painter. A story is a sequence of scenes, and each scene is one AI background plus a short beat tree of narration, dialogue and occasional options.

Why do I need three model providers to play infiplot?

Text for the plot, image for the scene rendering and vision for click interpretation are all required, nine variables in total. Speech is optional, and leaving those variables empty runs the game silently.

How much does one infiplot scene cost?

With the recommended providers, image generation dominates at about $0.00078 per image at 1792 by 1024, the text model is very cheap, and stepping through beats is free. Predictive generation means you also pay for scenes you did not end up choosing.

Can I run infiplot without deploying anything?

Yes, it plays free in a browser at infiplot.com. For self-hosting, fetch the compose file and the env example with two curl commands and run docker compose up -d, which needs no clone and supports x86 and ARM including Apple Silicon.

Official sources

  1. Issues
  2. License: AGPL-3.0
  3. Project website
  4. README
  5. zonghaoyuan/infiplot on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zonghaoyuan-infiplot.svg)](https://hysenlabs.com/projects/zonghaoyuan-infiplot)