Nova Studio's video tab is pure JSON: the host ships no upstream video protocol and plugins declare their own form
自托管的 AI 视频/图像生成工作台 · 自定义模型 · 多模式 · PWA · 实时任务 支持Agent模式,UI设计模式,工作台模式,无限画布,反推提示词,提示词广场,GIF生成。前后端任务机制轻量后端;三端兼容 UI:桌面端、平板端、移动端自适应布局
At a glance
- What is it?
- A self-hosted image and video generation workbench whose frontend is a static Next.js export and whose backend is three npm dependencies over SQLite and a WebSocket. Unusual and defensible: video capability arrives only as plugin packs, egress is allowlisted, and API keys never leave the browser.
- Who is it for?
- Nova Studio fits a small team that already holds API keys for image and text providers and wants a browser workbench it hosts itself rather than another hosted generator. It does not fit you if you need video on day one without writing or installing a plugin, because the host contains no video protocol at all.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 11 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The host ships no upstream video protocol, so video capability is only ever a plugin
Nova Studio is explicit that video generation contains no upstream-specific protocol. The host supplies the surrounding machinery and nothing more: the workbench tab, the task queue, history, media upload, and form rendering. Everything that makes a video provider a video provider lives in the plugin pack instead, and the project enumerates it plainly as who to call, what to send, how to poll, where the result is, and what the form looks like.
A plugin is three JSON files and no executable code: `manifest.json`, `ui.schema.json`, and `provider.json`. That is the whole security model in one sentence. Because a pack cannot execute anything, installing one cannot run code on the server, and the blast radius of a bad plugin is limited to what it declares in its manifest. The reference implementation is `backend/plugins/ccode-h3/`, covering MiniMax H3 with eight models and support for first and last frame, reference image, video and audio, and upscaled tiers.
Installation is an administrative act rather than a UI one. An admin drops the plugin directory into `backend/plugins/` and then restarts the backend, or clicks Reload in Settings. The official collection doubles as a template and is designed to be cloned straight into place:
cd backend/plugins && git clone https://github.com/tianjiangqiji/nova-studio-plugins.git .One practical wrinkle: the plugin protocol documentation under `docs/plugins/` is written in Chinese, while the rest of the project's documentation is English. For a non-Chinese-reading maintainer that is the single biggest obstacle in an otherwise well-structured extension path.
The entire video form is rendered from the plugin's ui.schema.json
Form rendering is the part that makes the plugin approach more than a URL template. The whole left-hand panel of the video workbench is generated from the plugin's `ui.schema.json`, and the schema is where tiers, resolutions, durations, and media slots are declared. The host knows nothing about any specific upstream, which is what allows a pack targeting a first-and-last-frame video model and a pack targeting a text-to-video model to present coherent, different interfaces from the same shell.
The credential model is a deliberate part of the same design. `apiKey` and `baseUrl` are entered under Settings, on a Plugins section, per plugin, and they are stored in the browser rather than in the database. Combined with the fact that all client configuration lives in the browser's localStorage, the server never holds your provider credentials. If the server is lost or replaced, keys are re-entered on each client rather than restored from a database dump.
The Settings page is also deliberately read-only in one direction. It lists what is installed and, importantly, why a pack failed to load, but installing and removing plugins requires server access. So there is no self-service plugin install from the browser, and the failure explanation you see there is the diagnostic channel when a pack will not load. Protocol reference material is split by concern across the manifest, the UI schema, and the provider documents, with a separate page for errors, a cookbook of known-good patterns, and a lifecycle page describing how a task progresses.
Egress is allowlisted and every private or loopback address is refused
Because a plugin decides where to send requests, the host has to constrain where those requests may go. The rule is that any host outside `permissions.hosts` is refused, and additionally every private or loopback address is refused. That second clause is the one doing real work: even if a plugin declares a host permission, a destination that resolves to a loopback or private-range address is still blocked. It closes the obvious path where a plugin pack reaches back into the machine it is installed on.
The complementary rule concerns results. Returned URLs pass through untouched, with no domain rewriting, so the video link you get is exactly what the upstream returned. That matters for two reasons at once. It means a plugin cannot quietly redirect a finished asset to infrastructure you did not intend, and it also means the project does not proxy media through itself, which keeps large video files off the server's disk and bandwidth.
The other half of the trust model is that video is not special-cased anywhere. The same JSON pack structure describes what to call and how to poll, and the same allowlist applies to every plugin, so adding a second video provider does not create a second security surface. For an organisation reviewing packs before installation, the review surface is three JSON files per pack, which is a tractable thing to read.
Image and text models are configured separately, and text means Google or OpenAI Responses
The bring-your-own-model story is deliberately narrow. Image models and text models are configured as separate categories, each with its own API key and base URL. You define the model list and the endpoints yourself, and the backend routes by protocol while passing your parameters through unchanged. That last part is the point: there is no parameter whitelist rewriting your request, so whatever a provider accepts can be sent.
For text, two protocols are named. Google models are called through `generateContent`, and OpenAI models through the Responses protocol. Anything else is out of scope for the text path, so the set of usable text providers is narrower than the set of usable image providers. The reverse-prompt mode depends on this layer: it uploads an image and streams a prompt back using any configured text model, which means the quality of that feature is bounded by whichever text protocol you have configured.
The three runtime dependencies map one to one onto the backend's described architecture. `server.js` runs on Node, SQLite holds task state through `better-sqlite3`, generation APIs are proxied over `undici`, and live task updates are pushed to the browser over `ws`. Those three packages, plus the standard library, are the entire server surface named in the package manifest, which is a small enough dependency list to audit by reading it.
Above that sit seven working modes with named entry components: text-to-image, image-to-image, an agent workspace, the UI design workspace, reverse prompt, GIF generation, and the plugin-driven video workbench. The GIF mode encodes in the browser with `gifenc` after multi-frame generation and grid assembly, so GIFs cost no server-side media work.
The Docker build installs a C++ toolchain, then purges it before shipping
The Dockerfile is three stages on `node:22-slim`, and one detail in the middle stage is worth reading twice. Installing `better-sqlite3` needs a native build toolchain, so the backend-deps stage installs `python3`, `make`, and `g++` before running `npm ci --omit=dev`, and then immediately purges all three with `apt-get purge -y --auto-remove`. The compiler is present exactly as long as it is needed and absent from the image that runs, so the native module ships without carrying a build environment around it.
The frontend stage copies only `frontend/package.json` and its lockfile before the full source, which is the standard cache-friendly layering. The production stage copies the backend, the resolved `node_modules`, and the built frontend from `frontend/out`, which is where a Next.js static export lands, then creates the data directory, exposes port 3000, and starts `node backend/server.js` with `NODE_ENV=production`.
Compose sits on top with one service published on port 3000 and a restart policy of `unless-stopped`. Two environment variables point the task database and image directory into the mounted volume: `NOVA_TASK_DB` and `NOVA_IMAGE_DIR`, both under `backend/data`. Alongside the data mount, compose also binds `blacklist.json`, `prompts.json`, `.env`, and a `./plugins` directory onto their expected paths, which is the supported way to add plugin packs and supply configuration without rebuilding the image. Note that the port mapping is `3000:3000` with no loopback bind, so the published port listens on all interfaces.
UI design mode stores nothing until you tick off what the model proposed
The UI design mode is the most involved of the seven modes, and it is the one with the strictest human checkpoint. You start from a flat mockup, a vision model proposes slices and background candidates, and none of it is stored until you tick items off in a confirmation dialog. That confirmation step is the design decision: an agent proposing a decomposition of your design should not be able to commit it unilaterally.
The slice editor that follows is a full image editor, with zoom and pan, drag-to-create, rubber-band multi-select, eight-way resize, per-corner radius, snapping, a context menu, keyboard shortcuts, and undo and redo to fifty steps. Three view modes let you check the work: original shows the source with outlines, knockout shows the post-cutout result, and slices only shows a checkerboard. Knockout is generated locally and stated as costing nothing.
The mode is wide-screen only, and narrow viewports show a hint to switch rather than a degraded layout. That is a deliberate cutoff rather than an omission, though it does mean the mode is unavailable on a phone. The list of asset operations is described as four in number, covering algorithmic transparency and AI transparency among them, but the documentation here cuts off before enumerating all four, so the complete set is not readable in this copy.
Beyond that mode sit the smaller tools: an agent workspace that goes chat to plan to images with vision descriptions, web search, reasoning, and local browser CDP for opening and reading pages, plus reverse prompt, a prompt gallery, an assets view, and an infinite canvas.
Editorial conclusion
Nova Studio fits a small team that already holds API keys for image and text providers and wants a browser workbench it hosts itself rather than another hosted generator. It does not fit you if you need video on day one without writing or installing a plugin, because the host contains no video protocol at all. Two things to check first. Whether the plugin protocol documentation is usable by you, since those documents are written in Chinese while the rest of the project is not. And what you are exposing, because the Docker compose file publishes port 3000 on all interfaces rather than on loopback; the design keeps keys in browser storage instead of the database, which is the right instinct, but the instance is still reachable. Pin a version rather than tracking latest, since there are no GitHub releases and the compose file defaults to the latest tag.
Frequently asked questions
What is Nova Studio?
Nova Studio is a self-hosted AI video and image generation workbench for individuals and small teams. The frontend is a Next.js 16 and React 19 static export PWA, and the backend is a small Node.js service using server.js, SQLite, and WebSocket that schedules tasks and proxies generation APIs. The current version is v3.4.0.
Does Nova Studio support video generation out of the box?
No. Video generation is fully plugin-based and the host ships no upstream video protocol at all. Capability comes only from plugin packs, which are three JSON files named manifest.json, ui.schema.json, and provider.json, with no executable code.
How do I install a Nova Studio video plugin?
An admin drops the plugin directory into backend/plugins/, then restarts the backend or clicks Reload in Settings. The official collection can be cloned straight into place with `cd backend/plugins && git clone https://github.com/tianjiangqiji/nova-studio-plugins.git .`
Where does Nova Studio store my provider API keys?
In the browser. apiKey and baseUrl are entered per plugin under Settings, Plugins, and stored in the browser rather than in the database. The Settings page is read-only for installs and removals, which require server access.
Which text model protocols does Nova Studio support?
Google models through generateContent, and OpenAI models through the Responses protocol. Image models and text models are configured separately, each with its own API key and base URL, and the backend routes by protocol while passing your parameters through.
What are the runtime dependencies of the Nova Studio backend?
Three: better-sqlite3 for task state in SQLite, undici for proxying generation APIs, and ws for WebSocket task updates. The Docker image runs on node:22-slim, exposes port 3000, and starts node backend/server.js.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tianjiangqiji-nova-image-studio)