# JarvisHub: the canvas is the project state, not a picture of it

> JarvisHub is a canvas-native harness for long-horizon creative agents. Text, references, storyboards, images and revision notes become addressable nodes that the agent reads and writes each turn, with a protocol bridge in between deciding which mutations are allowed. It is young: open-sourced 2026-07-28, last pushed 2026-08-03, no releases, and no tests in the distribution.

**LYL1015/JarvisHub** — JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

- Repository: https://github.com/LYL1015/JarvisHub
- Stars: 611 · Forks: 55
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/lyl1015-jarvishub

## The canvas is the state, not a picture of it

The distinguishing claim is that the canvas holds project state rather than displaying it. Everything the work produces becomes an addressable node or link: text, references, character sheets, scenes, storyboards, images, videos, audio, webpage previews, presentation slides, candidates and revision notes. References, generation dependencies, version lineage and workflow continuation are expressed as typed nodes and edges, so lineage is part of the graph instead of a naming convention.

The agent loop is short and inspectable. On every turn it observes the current canvas, selects a permitted action, invokes a model or tool, and commits the returned artifacts and evidence back to the same workspace. Requests, actions, observations, feedback, repair decisions and canvas updates are all recorded, which is what makes a failure recoverable later instead of lost in a transcript.

The project's own framing is that prompt tools, chatbot agents and node workflows each cover part of creative production, and that none of them keeps the whole state visible and editable at once.

## Three layers split storage, protocol and runtime

The architecture is a table with three rows, and the separation is the design decision. Canvas State stores editable artifacts, spatial layout, dependencies, versions, runtime status, user choices and feedback; its implementation is spread across `apps/web`, `packages/canvas-layout`, `packages/schemas` and PostgreSQL. Protocol Bridge exposes capabilities and execution grants, validates canvas mutations and tool actions, and commits state transitions; that is `apps/hono-api`. Agent Runtime observes the canvas, plans actions, invokes Skills, Memory, Tools and Subagents, and returns commit-ready observations; that is `apps/agents-cli`.

The bridge is what keeps the canvas trustworthy. An agent does not write to state directly, it proposes a mutation that the bridge validates, so the set of permitted actions is a protocol rather than a convention, and an unapproved action fails at a defined point.

The vocabulary on the runtime side is deliberately small: Skills for reusable procedures, Memory for preferences and prior decisions, and Subagents that explore independent subtasks before the parent integrates what proved useful.

## run.sh does six jobs before you see an interface

Requirements are pinned harder than most: Node.js `v24.15.0`, pnpm, Docker and Docker Compose for the default local PostgreSQL, and Python `3.10+` with `python-pptx` and Pillow. With those in place, startup is one command:

```bash
./run.sh
```

On the first run the launcher installs missing workspace dependencies, starts local PostgreSQL and waits for it, validates the vendored PPT Master runtime, selects a compatible Python executable, builds the Agents CLI, and only then starts Web, API, Agents Bridge and Trace Viewer. It stays in the foreground and aggregates service logs, and `Ctrl+C` stops the processes that run started.

The build step is the slow part, which is why there are maintenance flags for the pieces: `./run.sh restart` rebuilds Agents and restarts every service, `--install` reruns pnpm install, `--no-build` skips the Agents build, and `--clean` clears Agents caches before building. `./run.sh db` starts PostgreSQL alone.

## Docker mode adds Redis and needs a mirror when pulls are slow

The container route is a different command and a different service list:

```bash
./scripts/dev.sh docker --build
```

Docker mode starts Web, API, Agents Bridge, Trace Viewer, PostgreSQL and Redis, which is the difference from the local run. The first run pulls images and installs dependencies and can take several minutes; dependencies land in Docker volumes, so later starts are faster.

There is an escape hatch for a slow or blocked Docker Hub, passed as an environment variable on the same command:

```bash
DOCKERHUB_REGISTRY=mirror.gcr.io/library ./scripts/dev.sh docker --build
```

Shutdown uses the compose file directly rather than a run.sh flag:

```bash
docker compose -f apps/hono-api/docker-compose.yml down
```

## Five services, five default ports, and env overrides for all of them

The default endpoints are fixed and worth writing down, because the Trace pair is easy to miss: Web on `http://localhost:5173`, API on `http://localhost:8788`, Agents Bridge on `http://localhost:8799`, Trace API on `http://localhost:5781` and Trace Web on `http://localhost:5782`. Each has a matching environment variable, plus one for the database, so a second copy of the stack can run beside the first:

```bash
WEB_PORT=5174 API_PORT=18788 AGENTS_PORT=18799 TRACE_API_PORT=15781 TRACE_WEB_PORT=15782 POSTGRES_PORT=15432 ./scripts/dev.sh docker --build
```

Trace is the part worth understanding rather than skipping. It is a separate API and web pair, which is what makes the recorded requests, actions and repair decisions viewable while a long job is still running, rather than only readable from a log file after something goes wrong.

The workspace root is a pnpm monorepo pinned to `pnpm@10.8.1`, with `apps/`, `packages/`, `sql/`, `tools/`, `vendor/` and `docs/` beside it, and per-package scripts for web, api, a stable api variant and agents.

## The workspace test script is a sentence saying there are none

The root package.json is worth reading before anything else, because it sets expectations immediately. The `test` script is not a runner. Its whole body is an echo of the sentence stating that tests are not included in this open-source distribution.

Everything else in that file is ordinary: `dev` maps to `./scripts/dev.sh local`, `build` chains build:web, build:api and build:agents, `compose:up` and `compose:down` wrap Docker Compose, and `prepare` runs husky. Dev dependencies amount to husky alone, since the real dependencies live in the workspace packages.

Two smaller signals in the same file. The package is marked private, so it was never meant to be published to a registry. And the repository carries both `pnpm-lock.yaml` and `package-lock.json` even though the declared package manager is pnpm, which is the kind of leftover that makes a fresh install reproduce less reliably than it looks like it should.

## The paper's three tasks are long jobs, not asset demos

The published paper demonstrates three long-horizon tasks, and the shape of each one is the argument. Narrative media generation runs a story or script through character and location references, shot planning, storyboards, candidates and cross-shot revision, ending in character sheets, scene designs, shot lists, image sequences, video clips and animatics. Interactive web development turns design goals and interaction requirements into layouts, frontend code, rendered previews and iterative visual revisions. Presentation deck generation selects and organises content, synthesises diagrams, and holds narrative and visual consistency across slides.

None of the three is a single generation. Each is a chain in which a later step depends on what an earlier step produced and rejected, which is exactly the state that prompt tools hide and node workflows make you specify by hand.

The stated goal is not to replace generation models. It is an open, inspectable runtime for building agents that keep context, orchestrate tools, use feedback and recover from failure, with image, video, audio, code, browser, file, document, presentation and MCP-backed tools all invocable from the same canvas.

## Open sourced on 2026.07.28, with no release since

The updates section has a single line: the repository was open-sourced and the paper uploaded on 2026.07.28. The paper is on arXiv and mirrored on Hugging Face papers, the site is jarvishub.site, a demo video is linked, and a Chinese README ships alongside the English one.

What is absent is as informative. The repository has no GitHub releases, no tags to install and no changelog, so there is no upgrade path and nothing to pin. The last push to main is dated 2026-08-03, which places the public history at roughly a month and a half as of this writing, and the licences are Apache-2.0 with a LICENSE file at the root.

Given the license and the absence of releases, the sensible reading is that this is the code as it stood at open-sourcing, still being reshaped. Read the architecture tables rather than the release history, because there is no release history.

## Conclusion

Use JarvisHub if your creative work spans many turns and keeps failing because the agent forgets what it already made, because the canvas keeps artifacts, dependencies, versions and feedback in one place it can inspect. Hold off if you only need one-shot generation, or if you cannot run five local services plus PostgreSQL and Redis, since that is the actual deployment. Before building on it, read the `test` script and know the open-source distribution ships without tests, and treat the 2026-07-28 open-sourcing date as the start of its public life rather than a mature release.

## FAQ

### What is JarvisHub?

A canvas-native harness for long-horizon creative agents, licensed Apache-2.0 and written in TypeScript. The canvas holds artifacts, dependencies, versions, status and feedback as addressable nodes and links that the agent observes and commits to on every turn.

### What does JarvisHub run.sh do on the first launch?

It installs missing workspace dependencies, starts and waits for local PostgreSQL, validates the vendored PPT Master runtime, picks a compatible Python executable, builds the Agents CLI, then starts Web, API, Agents Bridge and Trace Viewer in the foreground.

### Which local ports does JarvisHub use?

Web on 5173, API on 8788, Agents Bridge on 8799, Trace API on 5781 and Trace Web on 5782 by default. Each is overridable through WEB_PORT, API_PORT, AGENTS_PORT, TRACE_API_PORT and TRACE_WEB_PORT, with POSTGRES_PORT for the database.

### Does the open-source JarvisHub distribution include tests?

No. The workspace package.json sets the test script to echo that tests are not included in this open-source distribution. The repository also has no GitHub releases, and its last push is dated 2026-08-03.

### What can a JarvisHub agent invoke on the canvas?

Image, video, audio, code, browser, file, document, presentation and MCP-backed tools, plus Skills for reusable procedures, Memory for preferences and prior decisions, and Subagents that explore independent subtasks for the parent to integrate.

## Sources

- [Issues](https://github.com/LYL1015/JarvisHub/issues)
- [License: Apache-2.0](https://github.com/LYL1015/JarvisHub/blob/main/LICENSE)
- [LYL1015/JarvisHub on GitHub](https://github.com/LYL1015/JarvisHub)
- [README](https://github.com/LYL1015/JarvisHub/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lyl1015-jarvishub
