# lanshu-create-ai-presenter-video: A Provider-Neutral Agent Skill for AI Presenter Video Production

> lanshu-create-ai-presenter-video is an Agent Skill that guides a coding agent through the full production of a talking-presenter video, from script and portrait to verified delivery, with an evidence-gated 8-stage pipeline and cost guardrails that prevent duplicate charges.

**cclank/lanshu-create-ai-presenter-video** — Provider-neutral Codex Skill for producing verified AI presenter videos from a script and an authorized presenter image.

- Repository: https://github.com/cclank/lanshu-create-ai-presenter-video
- Stars: 2,049 · Forks: 340
- Language: Python
- License: MIT
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/cclank-lanshu-create-ai-presenter-video

## What lanshu-create-ai-presenter-video Does and Who It Is For

lanshu-create-ai-presenter-video is an Agent Skill, meaning its core artifact is a SKILL.md file that a coding agent reads and follows as a workflow definition. The agent takes a topic or script and an authorized presenter image as inputs, then drives through the full production pipeline: script writing (if a topic is provided), narration generation, presenter video generation, lip-sync, caption and keyword motion graphics, editing, rendering, and a quality assurance pass.

The skill is harness-neutral. The README lists Claude Code, Codex, Gemini CLI, Cursor, OpenCode, and GitHub Copilot as compatible harnesses. Any agent that can read files and run shell commands can follow the SKILL.md directly, even if it is not on the explicit supported list.

The intended user is a developer or content creator who uses a coding agent as a daily tool and wants to integrate video production into that workflow without switching to a dedicated web-based service. The skill is designed to be driven by the agent autonomously, with the operator confirming approval steps at defined checkpoints in the pipeline rather than supervising each individual command.

## The Evidence-Gated Eight-Stage Production Pipeline

The production pipeline has eight states: intake, content_locked, audio_locked, visual_plan_locked, presenter_generated, composition_checked, rendered, and verified. Each state requires evidence before the job may advance. The check_state.py script computes the state from the artifacts present in the job directory, rather than from a flag set by the agent.

The key design decision here is that the pipeline cannot advance by declaration. If the agent claims a narration is ready but no decodable audio file exists in the job directory, the state remains audio_locked. This prevents the agent from hallucinating progress through a state it did not actually complete. The README describes the system as evidence-gated rather than declaration-based.

The audio track is the master clock for the rest of the production: presenter motion, caption timing, scene boundaries, and final duration are all derived from the approved narration. This means that if the narration changes, the subsequent stages must be recomputed from audio_locked forward. The approach avoids desync between the audio and the visual elements.

The verified state requires that both the master encode and the share copy pass a full decode check and a loudness measurement. Nothing is considered delivered until those checks pass.

## Installing in Claude Code, Codex, or Other Supported Agents

The repository directory is the skill. The README requires keeping the folder name `lanshu-create-ai-presenter-video`, because the Agent Skills specification uses the folder name to match the skill's declared name.

For Claude Code, clone into the skills directory:

```bash
git clone https://github.com/cclank/lanshu-create-ai-presenter-video.git \
  ~/.claude/skills/lanshu-create-ai-presenter-video
```

For Codex, the skills directory is ~/.codex/skills/. For other Agent Skills clients, the README points to the client's documentation for the correct directory.

For agents without native Agent Skills support, the README provides an alternative: clone the repository anywhere and start the agent session by instructing it to read the SKILL.md directly:

```text
Read /path/to/lanshu-create-ai-presenter-video/SKILL.md and follow its workflow exactly.
```

To update the skill, run `git pull` inside the installed directory. If the skill is installed in multiple harnesses, each clone must be updated separately.

## Running the Production Workflow Manually

Most supported harnesses select the skill automatically from its description. For a guaranteed selection, use the harness-specific invocation: `/lanshu-create-ai-presenter-video` in Claude Code, `$lanshu-create-ai-presenter-video` in Codex. The skill works in the language of the user's request.

The agent runs the bundled Python scripts automatically as it follows the SKILL.md. To set up a job without an agent, point `SKILL_DIR` at the installation directory and run init_job.py:

```bash
SKILL_DIR=~/.claude/skills/lanshu-create-ai-presenter-video

python3 "$SKILL_DIR/scripts/init_job.py" \
  --job-dir ~/Videos/my-presenter-video \
  --presenter-image ~/Pictures/presenter.png \
  --topic "Context engineering in one minute" \
  --duration 60 \
  --aspect 9:16 \
  --rights-confirmed \
  --adult-presenter-confirmed
```

The `--rights-confirmed` and `--adult-presenter-confirmed` flags are explicit acknowledgments required by the preflight step. After init, complete any manual review steps in job.json and run preflight:

```bash
python3 "$SKILL_DIR/scripts/preflight.py" ~/Videos/my-presenter-video/job.json
```

After each production stage, record the output artifacts in job.json and let check_state.py compute the current state:

```bash
python3 "$SKILL_DIR/scripts/check_state.py" ~/Videos/my-presenter-video/job.json --write
```

## Provider Neutrality and Cost Guardrails

Voice synthesis, presenter video generation, lip-sync, and speech recognition are not bundled in the skill. The skill selects a provider for each capability at runtime based on what is available in the agent's environment. Each job records the provider, model, parameters, and task IDs actually used, which gives a traceable audit of what was called and at what cost.

The cost guardrails are a notable design feature. The skill uses a pilot-first approach for presenter generation: before generating the full video at full cost, it generates a short pilot clip to verify quality. Before the first paid API call, the agent presents an explicit billing statement that the operator must acknowledge. Retry limits are enforced to prevent runaway spending. If a provider call fails but returns a task ID, the skill can recover that job ID rather than submitting a duplicate request.

The README identifies tools like HeyGen and similar cloud services as the category this skill operates in. Those tools offer a web-based upload-and-generate interface without requiring a coding agent or a local script environment. lanshu-create-ai-presenter-video trades that simplicity for automation, auditability, and the ability to integrate the video production step into a larger scripted workflow.

## System Requirements and What the Skill Cannot Do Alone

The skill requires Python 3.9 or later, FFmpeg with ffprobe, Bash, jq, awk, and sed. These must be installed and on PATH before the agent can execute the bundled scripts. The optional HyperFrames compositor handles captions, motion graphics, and final rendering deterministically; without it, those steps require a different compositor.

Most critically, the skill requires access to at least one voice synthesis, presenter video generation, and lip-sync capability. These are not included. The operator must provide a cloud CLI, an API key, or a local model for each. The skill's provider-neutral architecture means it can work with multiple different backends, but it cannot produce any video without at least one real generation capability in place.

The skill does not bundle a script-to-video model. It is a workflow orchestrator that calls external generation tools through whatever interface the coding agent can invoke. A team that expects to generate videos without any external API or model subscription will find that the skill has nothing to call.

## Maintenance, Licensing, and Scope

The last push to the repository was on 2026-08-20. The repository has no GitHub releases. A Chinese-language README (README.zh-CN.md) is present alongside the English one. The repository includes a .github/ directory with CI configuration and a tests/ directory.

The skill is MIT licensed. The MIT licence permits use, modification, and redistribution without restriction. The scripts in the scripts/ directory, the SKILL.md workflow definition, and the agents/ and references/ directories are all part of the repository.

The SKILL.md file is the central artifact that the coding agent reads. This file is separate from the README, which is documentation for human readers. Anyone who wants to understand exactly what the agent will do during production should read SKILL.md directly, not the README.

## Conclusion

lanshu-create-ai-presenter-video is the right tool for a developer who already uses a coding agent (Claude Code, Codex, Gemini CLI, Cursor, or similar) and wants to produce a verified presenter video as part of a scripted workflow. It is not a one-click web app: it requires Python 3.9+, FFmpeg, and access to at least one voice synthesis and presenter video generation service. The last push was on 2026-08-20. Before starting a job, run the preflight script to catch local validation errors before any paid API calls are made.

## FAQ

### Does lanshu-create-ai-presenter-video include a video generation model?

No. The README states that voice synthesis, presenter video generation, and lip-sync capabilities must be supplied by the operator as a cloud CLI, an API, or a local model. The skill selects from available providers at runtime but cannot generate video without at least one real generation backend.

### Which coding agents does lanshu-create-ai-presenter-video work with?

The README names Claude Code, Codex, Gemini CLI, Cursor, OpenCode, and GitHub Copilot as compatible harnesses. Any agent that can read files and run shell commands can follow the SKILL.md directly.

### How do I update lanshu-create-ai-presenter-video after installation?

The README instructs running git pull inside the installed skills directory. If the skill is installed in multiple harness directories, each clone must be updated separately.

## Sources

- [cclank/lanshu-create-ai-presenter-video on GitHub](https://github.com/cclank/lanshu-create-ai-presenter-video)
- [Issues](https://github.com/cclank/lanshu-create-ai-presenter-video/issues)
- [License: MIT](https://github.com/cclank/lanshu-create-ai-presenter-video/blob/main/LICENSE)
- [README](https://github.com/cclank/lanshu-create-ai-presenter-video/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/cclank-lanshu-create-ai-presenter-video
