Model or dataset
Yacey/agnes-ai-generation-skill avatar
Yacey/agnes-ai-generation-skill

agnes-ai-generation-skill wraps the Agnes text, image, and video APIs behind one Agent Skill, with the video defaults already tuned

Agent Skill for Agnes AI text, image, and video generation APIs.

419 stars77 forksPythonMIT

At a glance

What is it?
A single-file Python script packaged as a standard SKILL.md, installable into Codex, Claude Code, Cursor, and similar clients. It pre-validates video frame parameters, translates non-English prompts before sending them, and polls async video jobs, while being explicit that its tool-calling support is request shape rather than a working agent loop.
Who is it for?
Adopt this if you want an agent to drive Agnes generation without hand-writing HTTP calls, and if you work in a Windows shell where the documented PowerShell commands run as written. Two things to weigh first: the tool-calling path is request-shape compatibility, not a multi-turn agent loop, so do not build an autonomous tool cycle on it.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 112 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Video requests are pre-validated because Agnes rejects malformed frame counts

The video defaults are not arbitrary. The script sends `--num-frames 121 --frame-rate 24` by default, chosen to avoid the instability that comes from omitting key video parameters, and it validates the request before it leaves the machine. Specifically, it checks that `num_frames` satisfies `8n + 1` and does not exceed `441`, and it checks frame rate and size alongside that. So a typo like a frame count of 120 is caught locally instead of becoming a provider-side rejection you have to diagnose from an error body.

The command surface is one script with subcommands. Text generation goes through `text`, images through `image`, video through `video`, and status lookup through `video-get`:

powershell
python scripts/agnes_api.py video --prompt "A cinematic shot of a cat walking on the beach at sunset" --poll

Three model names carry the work. Text uses `agnes-2.0-flash`, images and image editing use `agnes-image-2.1-flash`, and video uses `agnes-video-v2.0`. Images also take a `--size` flag, as in `--size 1024x768`, and image-to-image or editing passes an existing file or URL through `--image`. Seed-based reproducible generation is part of the feature set, which matters if you want the same prompt to return the same output while iterating on something else.

Non-English prompts are rewritten by agnes-2.0-flash before the media call

Agnes video generation is steadier with English prompts, so the script detects non-English characters in an image or video prompt and silently runs a translation pass through `agnes-2.0-flash` first, then sends the translated text to the media API. Nothing about the call site changes:

powershell
python scripts/agnes_api.py image --prompt "一座高信息密度的未来城市集市,拥挤人群,飞行汽车,全息招牌,电影感写实风格"

What survives the rewrite is specific: subject, scene, style, lighting, composition, camera motion, action description, and any negative prompt or constraint. That list is the reason to trust the pass, since camera motion and negative prompts are exactly the elements a naive translation tends to flatten. The translated text is reported back in the output as `translated_prompt`, so you can see what was actually sent rather than guessing. When you want the original wording preserved, or you are iterating on an English prompt and do not want a round trip at all, `--no-translate-prompt` skips the step entirely.

Video is an async job: video_id first, then a polled query

Creating a video returns a job rather than a file, so the skill handles the polling and the format differences between Agnes response versions. When creation returns a `video_id`, the script prefers the newer query interface at `/agnesapi?video_id=...`. Only if an older-shaped response has no `video_id` does it fall back to the compatible `task_id` query path. That preference order is one of the three fixes called out for v0.1.0, and it is the kind of detail that otherwise turns into a silent failure when the provider changes a field.

Once a job completes, the script hunts for the playable link across three places: `video_url`, `url`, or the `remixed_from_video_id` value that Agnes has actually returned in practice, and it normalizes whichever it finds into `urls`. You can inspect a job directly:

powershell
python scripts/agnes_api.py video-get video_123456

Normal output is tidied into fields such as `content`, `urls`, `translated_prompt`, and `next_steps`, while `raw` is preserved alongside so an agent or a person can keep working from the untouched response. Add `--raw` when you want nothing but the original Agnes payload. Multi-image and keyframe generation reuse the same `video` subcommand, passing `--image` more than once and adding `--mode keyframes` for a transition between two supplied frames.

Media is returned as urls and never downloaded unless you ask

The default output contract is deliberately narrow: after an image or video is generated, the skill hands back `urls` and stops. It does not fetch the media to disk unless you explicitly ask it to download, save, or inspect a local file. That is the second v0.1.0 fix, and it removes a class of surprise where an agent quietly fills a working directory with media nobody requested.

The same restraint shows up in how streaming text is handled. A streamed response is aggregated into `content`, and alongside it the script keeps the number of events, whether the stream completed, and a prefix of the raw response, which is enough to tell at a glance whether the streaming endpoint behaved without keeping the whole transcript. The result is that a caller gets something readable and something checkable, rather than one or the other.

Installation follows the same minimal shape. The skill uses the standard `SKILL.md` layout and installs into any client that supports Agent Skills, including Codex, Claude Code, OpenClaw, Cursor, and Windsurf:

powershell
npx skills add Yacey/agnes-ai-generation-skill

Adding `--all` installs it to every supported agent at once. After that, asking the installed agent to generate an image or video with Agnes triggers the skill on its own, and the key is read from `AGNES_API_KEY`, `AGNES_API_TOKEN`, or `APIHUB_AGNES_API_KEY`, with the documented setup written in PowerShell form for Windows shells.

Tool calling here is a compatible request shape, not an agent tool loop

This is the most important caveat in the project and it is stated plainly rather than buried. Agnes's Responses API multi-turn function calling is not currently suitable as an automatic tool loop for agents like Codex or Claude Code. The skill's script uses the chat completions path instead. Its OpenAI-compatible tool-calling support should be read as the ability to send a well-formed tool-call request structure, not as stable multi-turn tool execution.

The distinction shows up in testing. The smoke test does exercise the tool-calling request structure, and it notes that Agnes sometimes accepts such a request without returning any `tool_calls` at all. By default that case only prints a warning rather than failing the run, because a missing `tool_calls` response is not necessarily an error. If you would rather treat it as one, `--strict-tools` makes it fail:

powershell
python scripts/agnes_api.py smoke-test --strict-tools

So if you are planning to let an agent loop on tool calls here, verify that behaviour against your own account first rather than assuming the request structure implies the loop works. The distinction between a well-formed request and a functioning multi-turn cycle is the whole gap, and the project is honest that it has not closed it.

smoke-test tiers exist so one run does not create a pile of video jobs

The default smoke test covers basic text generation, streaming text, the tool-calling request structure, and text-to-image. It does not create video tasks, which makes it the cheap check to run after installing or changing a key:

powershell
python scripts/agnes_api.py smoke-test

Video is opt-in and, when enabled, selected one mode at a time so a single run cannot spawn a batch of long jobs. `--video-case text-to-video` picks one capability, and the four accepted values are `text-to-video`, `image-to-video`, `multi-image`, and `keyframes`. Image editing is likewise opt-in through `--include-image-edit`. The maintainers' own guidance is to test each video mode individually and keep the provider error text, which is sensible given that video jobs are asynchronous and billable.

The project's own status notes are worth reading as a maturity signal. Basic text, streaming, the tool-call structure, text-to-image, image-to-image, high-information-density text-to-image, auto-translated Chinese-prompt image and video creation, and text-to-video plus image-to-video completing with an mp4 URL are all recorded as passing against the real API. Multi-image and keyframe video are listed as supported but not fully re-tested end to end in that round, and it is stated plainly that not every multi-image or keyframe task has been confirmed to return a final video URL. One real text-to-video job also returned a server-side `division by zero` error when queried, while later short jobs completed and returned mp4 links. Video support is real, and the verification is partial.

The key handling is PowerShell-shaped and the repo warns about leaking it

The documented setup is Windows-first. A session-scoped key is set with `$env:AGNES_API_KEY="YOUR_API_KEY"`, and a persistent per-user value uses `[Environment]::SetEnvironmentVariable("AGNES_API_KEY", "YOUR_API_KEY", "User")`. Every command example in the project is a `powershell` fence, and the install path is `npx skills add`. On macOS or Linux the underlying mechanism still works, since the script just reads environment variables, but you are translating the shell syntax yourself and the documentation will not do it for you.

The script is forgiving about the variable name, accepting `AGNES_API_KEY`, `AGNES_API_TOKEN`, or `APIHUB_AGNES_API_KEY`, which matters when the same key is shared with an API hub setup. The project is also direct about the obvious mistake: do not put an API key into a Git repository, into a README, into a screenshot, or into a public chat log. The workflow even suggests handing the key to the AI in a session you trust so the agent can configure the variable and call the skill, which is convenient and is exactly the behaviour the warning is about, so keep that session private.

Everything lives in a small repository: SKILL.md, README.md, README_EN.md, LICENSE, an `agents/openai.yaml`, a `references/api.md`, and the single `scripts/agnes_api.py`. One release tag exists, v0.1.0, published on 21 June 2026, which is also the last push to the default branch. Treat the surface area as small and the version history as short.

Editorial conclusion

Adopt this if you want an agent to drive Agnes generation without hand-writing HTTP calls, and if you work in a Windows shell where the documented PowerShell commands run as written. Two things to weigh first: the tool-calling path is request-shape compatibility, not a multi-turn agent loop, so do not build an autonomous tool cycle on it. And video is asynchronous and provider-side, with the maintainers reporting one real `division by zero` from the Agnes side and multi-image and keyframe modes not yet fully verified end to end. Before the first real job, run `python scripts/agnes_api.py smoke-test`, then test one video mode at a time with `--video-case` rather than creating a batch of tasks.

Frequently asked questions

Can agnes-ai-generation-skill generate video as well as images?

Yes. Text-to-video, image-to-video, multi-image video, and keyframe animation all run through the `video` subcommand using `agnes-video-v2.0`. Video is asynchronous, so creation returns a job and `--poll` waits on the result, while `video-get` queries a job id directly.

How does agnes-ai-generation-skill handle non-English prompts?

The script detects non-English characters in image and video prompts and first calls `agnes-2.0-flash` to translate them, preserving subject, scene, style, lighting, composition, camera motion, action, and any negative prompt. The result is reported back as `translated_prompt`, and `--no-translate-prompt` skips the step.

Does agnes-ai-generation-skill download generated media to my machine?

No. By default it returns `urls` and does not download anything unless you explicitly ask it to save, download, or inspect a local file. Returning only URLs was one of the changes in the v0.1.0 release.

Can I use agnes-ai-generation-skill for multi-turn agent tool calling?

Not as a stable agent tool loop. The Agnes Responses API multi-turn function calling is described as unsuitable for that, so the script uses the chat completions path and its tool-calling support is the request structure only. Treat it as compatibility, not as working multi-turn execution.

How do I install agnes-ai-generation-skill and supply the Agnes API key?

Run `npx skills add Yacey/agnes-ai-generation-skill`, or add `--all` to install to every supported Agent Skills client. The key is read from `AGNES_API_KEY`, `AGNES_API_TOKEN`, or `APIHUB_AGNES_API_KEY`, with PowerShell examples given for setting it.

Which parts of agnes-ai-generation-skill are not fully verified end to end?

Multi-image video and keyframe video are supported but were not fully re-tested end to end in the tested round, and not every such task has been confirmed to return a final video URL. One real text-to-video job also returned a server-side `division by zero` error when queried.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. Yacey/agnes-ai-generation-skill on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yacey-agnes-ai-generation-skill.svg)](https://hysenlabs.com/projects/yacey-agnes-ai-generation-skill)