# Generative Media Skills: Agent Skills for Image, Video and Audio Generation

> SamurAIGPT/Generative-Media-Skills packages image, video and audio generation as reusable skills for Claude Code, Cursor, Gemini CLI and OpenCode, with all model calls delegated to the muapi-cli and the muapi.ai API. The design is clean; the dependency on one hosted provider is the whole story.

**SamurAIGPT/Generative-Media-Skills** — Multi-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video, and audio generation powered by muapi.ai.

- Repository: https://github.com/SamurAIGPT/Generative-Media-Skills
- Website: https://muapi.ai?utm_source=github&utm_medium=about&utm_campaign=generative-media-skills
- Stars: 4,792 · Forks: 539
- Language: Shell
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/samuraigpt-generative-media-skills

## What Generative Media Skills actually packages

The repository is not a model and not a server. It is a set of instruction-and-script packages that an AI coding agent loads as skills, so the agent can call image, video and audio generation without the user writing API glue. The README describes the target clients as Claude Code, Cursor, Gemini CLI and OpenCode, and the top-level entries confirm that with .codex-plugin/, .opencode/ and .github/ directories sitting alongside core/ and library/.

The audience is narrow and specific. If you already work inside an agentic coding tool and you keep hitting the point where the task needs a rendered image or a short clip, this repository is aimed at you. The README frames the value as an "Expert Knowledge Layer": skills that bake in cinematography, atomic design and branding logic rather than exposing raw model parameters. That is the real product. The generation itself is commodity; the prompt structure is the differentiator.

## The Core/Library split and how a request flows

The repository is organised as two layers, which the README calls a Core/Library split. The core primitives live under /core and are thin wrappers around muapi-cli: core/media/ handles file upload, core/edit/ handles prompt-based image editing, and core/platform/ handles setup, auth and result polling. The library under /library holds higher-level skills such as Cinema Director at /library/motion/cinema-director/ and Nano-Banana at /library/visual/nano-banana/.

Data flow follows that split. A library skill turns a creative request into a technical directive, then delegates to a core primitive, which shells out to muapi-cli, which talks to muapi.ai. Generated files land in media_outputs/ at the repository root. The README states that local files are auto-uploaded to the CDN for processing, so a face reference or an audio track on disk enters the pipeline without a manual upload step. Two details matter for agent use: outputs are structured JSON with semantic exit codes, and the --jq flag filters results inline. Semantic exit codes mean an agent can branch on failure type without parsing strings.

## Installing it and generating a first image

There are no releases in the repository, so installation is from source. The README points at muapi-cli on npm as the underlying engine and gives muapi mcp serve as the MCP entry point. Clone the repository first, then confirm the CLI is reachable.

```bash
git clone https://github.com/SamurAIGPT/Generative-Media-Skills.git
cd Generative-Media-Skills
```

The README does not print an npm install line for muapi-cli, so check the package page for the current install command before assuming it is present. Once it is, the MCP route is the shortest path to a working setup: muapi mcp serve exposes the 19 tools to Claude Desktop, Cursor or any MCP-compatible agent, and you should see those tools listed in your client after the server starts.

```bash
muapi mcp serve
```

For a direct call outside an agent, the README documents the --view flag, which downloads the generated media and opens it in your system viewer. Use it on a first run so you can confirm the output landed rather than trusting a JSON response alone. Generated assets accumulate in media_outputs/, so check that directory if a call reports success but you cannot find the file.

## The provider dependency is the design, not a footnote

Every core primitive delegates to muapi-cli, and muapi-cli talks to muapi.ai. The README is explicit that there is no curl and no JSON parsing in the skills themselves. That is a real convenience and a real constraint at the same time.

If muapi.ai changes its model roster, its pricing or its availability, the skills change with it. There is no adapter layer described in the README for swapping in a different backend, and no bring-your-own-key path is documented. The repository is MIT licensed, so the code itself is permissive, but the licence does not travel to the hosted service. A team that needs an air-gapped pipeline, a fixed per-image cost, or a contractual guarantee about where prompts are stored will find this the wrong tool. The README also does not document rollback or version pinning for the underlying CLI, which matters if you are wiring this into anything that runs unattended.

## How it compares with Open-Generative-AI and raw SDKs

The repository's own related-projects list names Open-Generative-AI as a "free self-hosted AI media studio" and a "GUI alternative to these skills for the same model set". The difference in approach is the interface and the control boundary. Open-Generative-AI gives you a graphical studio you run yourself; these skills give you no interface at all, only instructions an agent reads and scripts it executes. If you want to click through model options and compare outputs visually, the skills package is the wrong shape. If you want an agent to make the call mid-task without a human in the loop, the GUI is the wrong shape.

The other comparison point is writing the API calls yourself. Several sibling repositories in the same list, such as flux-3-video-api and midjourney-api, are Python SDKs for specific models. Those give you a typed surface and full control over request construction. The skills package trades that control for prompt knowledge: Cinema Director encodes film-direction vocabulary so the agent does not have to invent it. Pick the SDK when you know exactly which parameters you need; pick the skills when the hard part is knowing what to ask for.

## Maintenance, licence and upgrade cost

The last push to the default branch was on 2026-09-08, which is recent, and the repository is not archived. There are no tagged releases, so there is no versioned upgrade path: you track main. For a skills repository that is mostly markdown and shell wrappers, tracking main is low risk, but it does mean a change to a core primitive can reach you without a version number to pin against.

The licence is MIT. That covers the repository contents. It does not cover muapi.ai usage, which is a separate commercial relationship governed by the provider's terms, and the README does not restate those terms. Treat the two as independent decisions: you can fork and modify the skills freely under MIT, and still be bound by whatever the API account requires. Nothing in the repository suggests a self-hosted fallback for the generation step itself.

## Conclusion

Adopt it if your agent already runs through muapi-cli and you want prompt-level cinematography and design skills rather than raw model calls. Do not adopt it if you need a self-hosted, provider-independent pipeline: every primitive delegates to muapi.ai, and the README documents no offline or bring-your-own-key path. Before committing, run muapi mcp serve and confirm the 19 tools appear in your client, then check core/platform/ for how auth and polling are wired. The repository's own licence file is the authority on reuse terms; the README only states MIT.

## FAQ

### What does generative media mean in the context of Generative Media Skills?

In this repository it means image, video and audio assets produced by AI models through the muapi.ai API, invoked from an AI coding agent rather than from a standalone app. The README lists Midjourney v7, Flux Kontext, Seedance 2.0, Kling 3.0 and Veo3 among the accessible models.

### Can you give me some examples of media skills in Generative Media Skills?

The README names Cinema Director under /library/motion/cinema-director/ for film direction and cinematography, and Nano-Banana under /library/visual/nano-banana/ for reasoning-driven image generation. The core layer adds media upload, prompt-based image editing and platform setup as separate primitives.

### What skills are needed for generative AI work with this repository?

You need a working muapi-cli installation and an account with muapi.ai, since every primitive delegates to that CLI. Beyond setup, the repository supplies the domain knowledge itself, so the skill requirement sits with the agent client you run rather than with the user.

### What are the top 5 AI skills in Generative Media Skills?

The README does not publish a ranked list of five skills. It groups the work into core primitives for upload, editing and platform setup, and library skills such as Cinema Director and Nano-Banana, with the MCP server exposing 19 tools in total.

## Sources

- [Issues](https://github.com/SamurAIGPT/Generative-Media-Skills/issues)
- [License: MIT](https://github.com/SamurAIGPT/Generative-Media-Skills/blob/main/LICENSE)
- [Project website](https://muapi.ai?utm_source=github&utm_medium=about&utm_campaign=generative-media-skills)
- [README](https://github.com/SamurAIGPT/Generative-Media-Skills/blob/main/README.md)
- [SamurAIGPT/Generative-Media-Skills on GitHub](https://github.com/SamurAIGPT/Generative-Media-Skills)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/samuraigpt-generative-media-skills
