vidu-s is a README and a figures directory, and the API lives somewhere else entirely
Vidu S: Real-Time Interactive, Editable, and Spatial Video Generation
At a glance
- What is it?
- shengshu-ai/Vidu-S documents Vidu S2 and Vidu S1, ShengShu AI's real-time interactive video models, with two arXiv papers, a web trial, four documentation links per language, and an agent skill you install from a different repository. There is no code here, no licence file, and no releases, so every capability claim is a paper claim rather than something you can run.
- Who is it for?
- vidu-s is worth reading if you are assessing Vidu S2 or Vidu S1 and want the claims, the papers, and the exact wording of the capabilities rather than a product page. It is not a codebase: the repository holds a README and a figures directory, there is no LICENSE file, the licence field is unset, and there are no releases, so nothing here can be installed, pinned, or vendored.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The tree holds two entries and one of them is a folder of images
The repository contents are `README.md` and `figures/`. That is the whole list.
There is no source directory, no package manifest, no tests, no continuous-integration configuration, and no licence file. The primary language field is unset, which follows from the contents: a Markdown document and images do not register as a language. The licence field is unset too, and with no LICENSE file in the tree there is nothing to read, so the terms under which the text and figures can be reused are not stated anywhere in the repository.
There are no GitHub releases either. The last push was on 2026-09-20, and the repository is not archived.
All of this makes Vidu-S a paper repository rather than a software one. That is a legitimate thing for a lab to publish, and it means the README's job is to point at work rather than to let you run it. What it cannot do is give you an artifact to install.
Everything executable is hosted somewhere else. The trial is a web application on the vendor's own domain, and the integration path is a documented API on a platform subdomain. There is no checkpoint download, no container image, no SDK listing, and no mention of weights anywhere in the README, so the entire relationship between this repository and anything you can execute is a URL.
The agent skill is installed from a different repository by asking
The one thing in the README that resembles installation is an instruction to say a sentence to a program.
For agent-assisted API integration, the project points at a skill called `vidu-s-api`, and says that in Claude Code, Codex, OpenClaw, or any agent that supports Skills, you can say this directly.
Install this skill: https://github.com/shengshu-ai/vidu-s-api/tree/main/skills/vidu-s-apiThe agent clones it and installs it into the proper skills directory, you restart the agent if required, and then you ask it to load `vidu-s-api` for Vidu S2 API integration.
Three things are worth noticing about how that works.
The skill is not in this repository. The URL points at `shengshu-ai/vidu-s-api`, a separate project, and the path names the `skills/vidu-s-api` directory on its main branch. So the branch tip, not a tagged release, is what you install, and the version you get is whatever that repository's main branch held at the moment the agent cloned it.
The destination is described only as the proper skills directory. Which directory that is, and whether the agent has permission to write it, is settled by the agent rather than by anything in this README.
And the install is a natural-language instruction rather than a command. Nothing here is pinned, checksummed, or listed as a dependency. For a team that treats agent-installed code as untrusted input, that is the intended workflow rather than an oversight, but it does mean the review step has to happen on whatever the agent fetched rather than on a version you chose.
Four documentation links per language, on two domains and two wikis
The documentation section is the longest part of the README and the most awkward to act on.
It is split into English and Chinese, and each half lists the same four things: a user guide, API documentation, a Vidu S2-Avatar quickstart, and a Vidu S2-Editing quickstart. The API links sit on two different hosts. The English pages are under `platform.vidu.com` and the Chinese pages under `platform.vidu.cn`, with matching paths such as the realtime quick start for the avatar and a separate quick start for editing.
The user guide is not on either platform host. Both language versions are Feishu wiki pages, and the English and Chinese versions are different wiki documents with different identifiers.
So there are four distinct documentation systems in play: an international platform domain, a Chinese platform domain, and two Feishu wiki spaces. What the README does not say is how to choose. Nothing states whether the two platform domains serve the same accounts, whether one is a mirror of the other, or whether a user outside China should use the international host with the English wiki. The trial link and the platform links are all under the vendor's own domains, and the README gives no other routing information.
For anyone integrating this, that is the first question to resolve and it has to be resolved with the vendor. The quickstart paths are specific enough to be useful once you know which host you are on.
S1 is 540p at 42 FPS and S2 is 720p at 25 to 42 FPS
Two generations of model are documented, two months apart, and the numbers moved in opposite directions.
Vidu S1 is described as a real-time interactive video generation model for voice-controlled digital characters, where the user guides the generated content at any moment through spoken instructions. Its three claimed advances are real-time speech control over video content, infinite-length real-time interactive generation, and custom character images and voice tones. The performance figure is 540p at up to 42 FPS, and it states that it can run on consumer GPUs. The avatar range covers real people, anime-style characters, pets, and other personalised avatars.
Vidu S2 extends the same idea past talking-head digital characters, to high-resolution interactive avatars, live video editing, and immersive spatial video. Its headline figure is 720p at 25 to 42 FPS, which is a higher resolution over a range whose top end matches S1's. It also accepts new reference images at any moment during streaming and handles complex instructions such as dancing.
The timeline lines up with the arXiv identifiers and the Updates section. The S1 preprint is numbered 2607.03118, which is July 2026, and the Updates list records S1 as available that month. The S2 preprint is numbered 2609.11638, September 2026, and the Updates list records S2 as available then.
Two caveats on the figures. Both frame rates are ranges with no lower bound explained and no hardware named for S2, whereas S1's claim is tied to consumer GPUs. And neither number has a latency figure, a batch size, or a resolution-to-framerate relationship attached, so 720p at 25 FPS and 540p at 42 FPS are not directly comparable as an improvement.
Self-Replay Forcing trains on the model's own drift
The most interesting technical claim in the README is also the least explained.
Long-horizon stability comes from something called Self-Replay Forcing, which replays re-noised, self-generated trajectories during training to reduce error accumulation across streaming segments. The idea is legible from the name: instead of training only on clean ground-truth video, the model is trained on its own output, re-noised, so it learns to recover from the drift it produces rather than only to avoid it in the first place.
That matters for real-time streaming specifically. A live video model that runs for minutes accumulates error segment by segment because each segment conditions on what came before, and any deviation feeds into the next frame. Replaying its own trajectories attacks the failure mode at the source rather than smoothing the output afterwards.
What the README does not give is any evidence. There is no ablation table, no segment-length curve, no comparison of drift with and without the technique, and no statement of how long a stream has to be before the effect shows. The other four claimed advances are similarly unquantified.
The efficiency claim sits alongside it: TurboDiffusion and TurboServe are named as the components that combine efficient attention, low-bit GEMM, kernel optimisations, and multi-GPU pipelining for real-time inference on low-cost GPUs. Low-bit GEMM is the concrete one, since it is what makes consumer hardware plausible, but again no speed, memory, or hardware numbers are attached.
S2-Avatar, S2-Editing, and stereo output are three claims
Reading the feature list carefully, Vidu S2 is not one model doing one thing.
It is described as extending real-time video generation to three separate destinations, and each has its own named component and its own quickstart page in the documentation.
S2-Avatar is for controllable character generation. It is the piece that produces the 720p stream at 25 to 42 FPS, accepts new reference images at any moment during streaming, and follows complex instructions such as dancing.
S2-Editing transforms incoming video streams using reference images. Its listed capabilities are style transfer, virtual try-on, character replacement, and background replacement, and it states that source motion is preserved, which is the constraint that makes it an editing tool rather than a regeneration tool.
The third is real-time spatial video, which produces synchronized stereo views for immersive displays and VR headsets. That is a different output format from the other two: two views, side by side, for a display rather than a flat frame for a browser.
The practical consequence is that a trial tells you about one of these and not the others. A web page at the vendor's stream address is most likely the avatar path. Editing has its own quickstart and its own reference-image semantics. The stereo path has no quickstart in the documentation list at all, only the API documentation and the general user guide, which suggests it is the least documented of the three.
Two citations, both preprints, both with the same first four authors
The README closes with two BibTeX entries and an instruction to cite them if you use Vidu S for research.
The first is keyed to the S2 paper and titled Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation. The second is keyed to the S1 paper and titled Vidu S1: A Real-Time Interactive Video Generation Model. Both are listed as arXiv preprints for 2026, with the identifiers matching the two links at the top of the document.
Ten authors are named in each, followed by others. The first four are the same in both entries, which is what you would expect from one group iterating on one model line, and the tail of the two lists diverges from the fifth author onward.
There is no venue in either entry. No conference or journal, no DOI, no volume or pages, because both works are preprints. If you are writing this into a paper, the arXiv identifier is currently the only stable handle, and neither preprint has a publication venue recorded here.
The Updates section above the citations has only two entries, one marking S2 as available to try and one marking S1 as available. So between July and September 2026 the project went from one model to a three-part system, and the documentation was restructured to match, while the repository itself changed only its README and its figures.
Editorial conclusion
vidu-s is worth reading if you are assessing Vidu S2 or Vidu S1 and want the claims, the papers, and the exact wording of the capabilities rather than a product page. It is not a codebase: the repository holds a README and a figures directory, there is no LICENSE file, the licence field is unset, and there are no releases, so nothing here can be installed, pinned, or vendored. Everything you can actually use is hosted, and the choice of host is spread across two documentation domains and two wikis with no guidance on which applies to your account. Three things to keep in mind. The API skill you would hand to an agent is installed from a different repository and lands wherever that agent decides the skills directory is. The frame rates and resolutions are the project's own figures with no hardware, batch size, or latency attached, so they are not something to size a deployment on. And S2-Avatar, S2-Editing, and the stereo output are presented as separate capabilities rather than one model, so trialling one tells you nothing about the others.
Frequently asked questions
What is actually in the shengshu-ai/Vidu-S repository?
Two top-level entries, README.md and a figures directory. There is no source code, no package manifest, no test or CI configuration, no LICENSE file, and no GitHub releases. The last push was on 2026-09-20 and the licence and primary-language fields are both unset.
How do I install the Vidu S API skill for an agent?
You say a sentence to the agent. In Claude Code, Codex, OpenClaw, or any agent that supports Skills, give it the line Install this skill: https://github.com/shengshu-ai/vidu-s-api/tree/main/skills/vidu-s-api. It clones that other repository and installs into the skills directory, you restart the agent if required, then ask it to load vidu-s-api.
What can Vidu S2-Avatar generate in real time?
720p video at 25 to 42 FPS, following complex instructions such as dancing, and it accepts new reference images at any moment during streaming. Inference uses TurboDiffusion and TurboServe, which combine efficient attention, low-bit GEMM, kernel optimisations, and multi-GPU pipelining.
Where is the Vidu S API documentation?
Two places. API documentation is under platform.vidu.com/vidu-stream/doc for English and platform.vidu.cn/vidu-stream/doc for Chinese, each with a Vidu S2-Avatar quickstart and a Vidu S2-Editing quickstart. The user guide is on a Feishu wiki, with a separate wiki page per language, and the README does not say which host to use.
How does Vidu S2 keep long streams from drifting?
With Self-Replay Forcing, which replays re-noised, self-generated trajectories during training to reduce error accumulation across streaming segments. The README names the technique and explains the intent but gives no ablation, measurement, or comparison for it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/shengshu-ai-vidu-s)