Model or dataset
Orkas-AI/Orkas-VideoStudio avatar
Orkas-AI/Orkas-VideoStudio

OrkasVideoStudio: a plan.json your coding agent can edit, not a black-box video generator

Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file.

497 stars23 forksTypeScriptMIT

At a glance

What is it?
OrkasVideoStudio turns Claude Code, Codex or Cursor into a video studio by making the timeline an editable, re-renderable plan.json. It is early: the npm packages are still being published, so source installs are the path today.
Who is it for?
Adopt OrkasVideoStudio if you already drive a coding agent from a shell or over MCP and you want the timeline to stay a readable, diffable file rather than a hidden project format. Do not adopt it if you need a finished npm install with a support contract, if you have no ffmpeg and ffprobe on PATH, or if you expect the generative line to work without your own provider keys.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OrkasVideoStudio solves, and for whom

Most text-to-video tools hide the timeline behind a prompt box. You get an mp4, and when one caption is wrong you regenerate the whole thing. OrkasVideoStudio takes the opposite position. The README describes it as "not a black-box video agent": a video is expressed as a readable, diffable, re-renderable plan (`plan.json`) that your agent and you can edit, and changing one line re-renders only that piece.

The audience is narrow and specific. It is for engineers who already work inside Claude Code, Codex, Cursor, or any agent that can run a shell or speak MCP, and who want video production expressed as files in a repository rather than as sessions in a web app. The project ships three things: knowledge in the form of skills (what makes a good video, which production line to take), deterministic capabilities (render, edit, transcribe, generate) as thin wrappers over hyperframes, ffmpeg and whisper.cpp, and the editable intermediate representation that ties them together.

If you do not already have a coding agent in your workflow, the value proposition mostly evaporates. The CLI is real and usable on its own, but the design assumes an agent is the brain and the toolkit is the hands.

Four production lines and the router that locks one

The architecture is three capability axes plus an automatic pipeline that weaves them. Compose turns a script into designed HTML motion graphics and then an mp4: explainers, kinetic typography, lower-thirds, data viz, title cards, transitions. Edit works on footage you supply: cut, join, trim-silence, de-filler, mix, burn-in subtitles, dub, localize, plus highlight selection over long recordings. Both of those lines are documented as needing no paid keys. Generate is the one that does, and it is opt-in: talking-head or cinematic footage via your own provider keys (the README names OpenAI, Gemini and Doubao), with no managed backend.

Auto is not a fourth capability so much as a routing layer. When a deliverable crosses axes, the agent writes a single cross-modal `plan.json`. `stage-plan` builds the EDL, `stage-assemble` walks it and delegates each segment back to compose, generate or edit, and a delivery guard (`ovs plan promise-check`) verifies the finished cut keeps its promise before anything ships. The README gives the example of catching a silent slideshow that was supposed to have real motion.

That guard is the most interesting design decision here. It is a deterministic check, not a model judging its own output, which means it can fail loudly and reproducibly. It also means it only checks what someone wrote a rule for. A cut that keeps its motion promise but has bad pacing will pass.

Installing the ovs CLI from source and running doctor

The README is explicit that the npm packages are being published and that until then you install from source, with the note that "the `ovs` CLI works exactly the same." Prerequisites are Node 22 or newer and `ffmpeg` plus `ffprobe` on your PATH, which the edit, transcribe and local media QA paths need. The repository's package.json confirms the engine constraint of `node >=22` and pins `[email protected]`.

Clone, install the workspace, build, then verify the environment:

bash
https://github.com/Orkas-AI/Orkas-VideoStudio.git
cd Orkas-VideoStudio
pnpm install && pnpm build
node packages/cli/dist/index.js doctor

The `doctor` command checks ffmpeg, ffprobe and node, and the README says it guides any install it finds missing. If you want the shorter command name, the README suggests aliasing it to the built entry point. Once published, the npm path is `npm i -g @orkas/video-studio`, which provides the `ovs` command, followed by `ovs doctor`.

Then connect it to your agent. Three mechanisms exist. Native skills for Claude Code or Codex materialize a `SKILL.md` knowledge pack so the agent discovers it by progressive disclosure:

bash
ovs skills --install --target claude
ovs skills --install --target codex

The first writes to `~/.claude/skills` and the second to `~/.agents/skills`; adding `--scope repo` targets `./.claude/skills` instead. MCP registration mirrors the CLI:

bash
claude mcp add ovs -- npx -y @orkas/video-studio-mcp
codex  mcp add ovs -- npx -y @orkas/video-studio-mcp

If your agent has neither a native skill loader nor MCP, the CLI is self-describing. `ovs skills` lists the skills, `ovs skill video-router` prints a skill's full instructions into context, and `ovs --help` shows the command surface. In a session the agent is expected to read `video-router` first, because that skill locks the production line, then the relevant stage skills, then author the composition or plan and run the deterministic operations.

Where OrkasVideoStudio is the wrong tool

The honest limitation is maturity, and the README states it in bold rather than burying it: the npm packages are being published. Anything you build today depends on a source checkout that you build yourself. There are no retrieved releases, so there is no version to pin against and no upgrade notes to read before you move forward. A team that needs a stable artifact to install in CI is not served by this yet.

The second limitation is the dependency surface. ffmpeg and ffprobe must be on PATH, and the compose QA gate is backed by a pinned HyperFrames `0.7.60` package dependency, with `npx` described as only a compatibility fallback. That is a real constraint on air-gapped or tightly controlled build images: you are not installing one binary, you are installing a Node workspace plus external media tooling plus a pinned rendering dependency.

The third is scope. Generate needs your own provider keys and is opt-in, so a user who wants AI footage without managing provider accounts gets nothing from that line. And the delivery guard is only as good as its rules. It verifies a stated promise, such as real motion rather than a silent slideshow. It does not judge whether the video is any good. If your problem is taste rather than correctness, a deterministic guard will not solve it.

How this differs from HyperFrames and from editing suites

HyperFrames is the closest reference point, and it is a dependency rather than a rival: OrkasVideoStudio wraps it for rendering and pins version `0.7.60` for the compose QA gate. The difference is the layer. HyperFrames is the rendering engine; OrkasVideoStudio adds the routing knowledge, the `plan.json` intermediate representation, the edit and transcribe wrappers over ffmpeg and whisper.cpp, and the delivery guard. Choosing between them is choosing whether you want to author the rendering layer yourself.

The sharper contrast is with conventional editing suites, where the project file is a binary or an opaque format and the human is the only one who can meaningfully edit it. Here the artifact is JSON that an agent writes and a person reviews in a diff. That is the actual bet: reviewability over convenience. The cost is that you need an agent that can read and write that file coherently, and you need to trust the routing skill to pick the right line. A suite with a GUI does not ask you to trust a router.

Licence, maintenance and what an upgrade costs

The licence is MIT, declared both in the repository metadata and in the root package.json, and the top-level LICENSE file is present. MIT is permissive, so the practical implication is that you can vendor, modify and redistribute the toolkit, including inside a commercial product, provided you keep the copyright notice and licence text. That is a summary of what the licence identifier means, not legal advice; if you are redistributing, read the LICENSE file yourself.

The maintenance signal is thin in a specific way. The repository is not archived, and the last push was on 2026-09-10, so the code is recent. But there are no retrieved releases and the README says the npm packages are being published, which means the distribution story is incomplete rather than the development story. Upgrading means pulling `main` and rebuilding, and the risk sits in the pinned pieces: the HyperFrames `0.7.60` dependency and the ffmpeg and ffprobe versions on your machine. The repository does carry a CHANGELOG.md and a PLAN.md, so there is a place to look before you move, and the root package.json exposes `pnpm verify`, which chains build, typecheck, test, benchmark and two video test scripts. Running that after a pull is the concrete way to find out whether an upgrade broke your setup.

Editorial conclusion

Adopt OrkasVideoStudio if you already drive a coding agent from a shell or over MCP and you want the timeline to stay a readable, diffable file rather than a hidden project format. Do not adopt it if you need a finished npm install with a support contract, if you have no ffmpeg and ffprobe on PATH, or if you expect the generative line to work without your own provider keys. Before committing, run node packages/cli/dist/index.js doctor on the machine that will render, and read video-router to confirm the routing rules match the deliverable you actually need.

Frequently asked questions

What are the system requirements for OrkasVideoStudio?

The README lists Node 22 or newer plus ffmpeg and ffprobe on your PATH, which the edit, transcribe and local media QA paths need. The root package.json confirms the Node engine constraint of >=22 and pins pnpm 10.28.2.

How do I install the ovs CLI before the npm packages are published?

Clone the repository, run pnpm install and pnpm build, then invoke node packages/cli/dist/index.js doctor to verify ffmpeg, ffprobe and node. The README notes the ovs CLI works exactly the same from a source install.

Does OrkasVideoStudio need paid API keys?

The compose and edit lines are documented as needing no paid keys. Generation is opt-in and uses your own provider keys, with the README naming OpenAI, Gemini and Doubao and stating there is no managed backend.

How does a coding agent pick up OrkasVideoStudio?

Three ways are documented: native skills installed with ovs skills --install --target claude or --target codex, MCP registration that mirrors the CLI, or the self-describing CLI where ovs skills lists the skills and ovs skill video-router prints one into context.

Official sources

  1. Issues
  2. License: MIT
  3. Orkas-AI/Orkas-VideoStudio on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/orkas-ai-orkas-videostudio.svg)](https://hysenlabs.com/projects/orkas-ai-orkas-videostudio)