Library / SDK
ziguishian/xhs-visual-director-skill avatar
ziguishian/xhs-visual-director-skill

XHS Visual Director: a Codex Skill that plans Xiaohongshu image posts before it generates them

这是一个用于规划小红书图文的 Agent Skill。它不是普通文案助手,而是一个“视觉导演”:先判断内容任务,再选择适合的视觉风格,最后输出完整图文结构、逐页视觉方案、图像生成提示词、发布文案和自检清单。

1,354 stars130 forksUnknownMIT

At a glance

What is it?
ziguishian/xhs-visual-director-skill turns a topic or draft into a 6 to 8 page Xiaohongshu carousel plan, a single confirmation image, then the full set. It is a plan-first skill, and the ten-question intake is the part that will decide whether you keep it.
Who is it for?
Adopt it if you already publish Xiaohongshu carousels and want the style decision written down before any image is generated, because the ten-question intake and the style judgment report are the parts that survive a bad generation. Do not adopt it if you need unattended batch output: the workflow stops for your answers and then stops again for a single confirmation image, so it is a slow loop by design.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 78 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem is picking a visual style, not writing the copy

Most Xiaohongshu writing assistants start from the text. This one starts from the content type. The README frames the skill as a visual director rather than a copywriting helper, and the output order reflects that: a topic judgment, then a style judgment report, then three style options with one recommended, and only after that a page structure. The stated audience is Xiaohongshu image post creators, AI and Agent content creators, personal brand operators, product managers, designers and founders, plus anyone who wants to rework drafts, screenshots, product photos, slides or web pages into a carousel.

The problems it claims to address are specific and recognizable: not knowing which visual style fits a topic, carousels that read like slide decks, AI images with blue-purple gradients and unreadable text, pages that spread content evenly with no rhythm, covers that fail to earn a tap, and inner pages with nothing worth saving. The last item on that list is the interesting one. The README describes turning personal taste into a reusable visual workflow, which is a different promise from generating one good post.

How the skill actually runs: intake, style report, one sample, then the set

The mechanism is a fixed pipeline rather than a single prompt. According to the README, a full project triggers ten Socratic questions before anything is designed. The skill asks them to pin down the communication goal, target reader, core argument, available material, style direction, prohibited elements, and the comment and conversion goal. Only after you answer does it produce a summary of your answers plus the assumptions it is generating under. Complex commercial, product, course or personal brand projects get the same ten questions, and the README states that if key decisions are still missing after your answers, it asks another three to five.

The default deliverable is image files, not prompt templates. The documented flow is: ten questions, then the multi-page content and visual plan, then one cover or key-page sample image for visual confirmation, then your sign-off on style, composition, colour and information density, then the full six to eight final pages, then the local path and aspect-ratio check for each image. The README also notes a fallback for image models that render Chinese poorly: the skill prefers to generate a clean visual base with a clear text-safe area and then suggests overlaying real Chinese text afterwards. It still delivers images rather than stopping at a template.

The style layer is a built-in library of 24 styles, per the repository badge. The sample table maps themes to style pairs, for example dark tech magazine plus black-white-grey with fluorescent green for AI and Agent topics, and a phone screenshot rework plus Notion-style card look for real case studies and tool tutorials. The canvas badge states a 3:4 Xiaohongshu ratio, and the README describes the default cover direction as premium, clear and readable on a phone, avoiding a cheap AI template feel.

Installing the skill and running one real project

There is no package manager step. The repository states that the installable skill body lives in the skill/ directory, specifically skill/SKILL.md and skill/agents/openai.yaml. The root docs/, templates/ and examples/ directories are described as maintenance and extension material rather than part of the skill.

text
xhs-visual-director-skill/skill/SKILL.md
xhs-visual-director-skill/skill/agents/openai.yaml

If your Codex Skill setup requires SKILL.md to sit directly inside the skill directory, the README says to use xhs-visual-director-skill/skill/ as that directory. Clone the repository and point your setup at that path.

bash
git clone https://github.com/ziguishian/xhs-visual-director-skill.git

Then give the agent a topic in the shape the README uses. The example input is a theme plus a goal plus a style constraint, and the skill is expected to answer with questions before it answers with a plan.

text
主题:为什么普通人现在必须学习 Vibe Coding?
目标:做成 8 页小红书图文,风格要高级、有科技感,但不要像廉价 AI 模板。

What you should see next is not a finished carousel. The README documents the first reply as the ten clarifying questions. Answer them, and the next output is the answer summary and generation assumptions, followed by the style judgment report. That report has a documented shape: a primary style, supporting styles, an explicitly not-recommended style, and the reasoning. The example in the README picks dark tech magazine as primary, adds an architecture breakdown and a black-white-grey with fluorescent green accent as supporting, and rules out a liquid glass diffused aurora look because the topic needs cognitive impact and credibility rather than dreaminess.

After that you get the page structure and a single confirmation image. Only when you approve it does the skill generate the full set, and it finishes with a title, body copy, tags, a pinned comment and a self-check against templates/visual_review_checklist.md.

The confirmation gate is a feature and a bottleneck

The workflow stops twice: once for ten answers, once for one image. If you want a batch of posts generated overnight, this is the wrong tool. The README is explicit that the skill does not immediately produce a full plan from a bare topic, and the sample prompt "帮我做一篇'小红书图文视觉导演'的小红书" is given precisely as the case where it asks first instead of generating.

The second stop is the one worth thinking about. A single confirmation image is cheap, but it also means your approval decision is made on one page. If the cover is strong and page five is where the density problem lives, the gate does not catch it. The README's answer to that is the review checklist applied at the end, after the full set exists, which is a later and more expensive place to discover a composition problem.

There is a second limitation in the image layer itself. The README acknowledges that image models can be unstable at rendering Chinese text, and the fallback is to generate a base image with a text-safe zone and overlay the Chinese yourself. That is honest, but it moves real work back to you: the skill plans the typography and the safe area, and you still own the final text placement. Anyone expecting a finished, text-correct Chinese carousel straight out of the model should read that paragraph before installing.

How it differs from a general-purpose image prompt workflow

A generic prompt library, or a general image generation assistant, treats each page as an independent request. You describe a picture, you get a picture. The difference here is that the style choice is made once, argued for, and then applied as a master template across six to eight pages, with a per-page visual plan and a prompt record for each page. The README calls this a unified visual master and lists per-page detailed visual planning and prompt records as separate deliverables.

The second difference is the negative decision. The style judgment report is documented to include a style it does not recommend, with a reason. A prompt library has no mechanism for telling you that a look is wrong for your topic, because it has no model of your topic. That is the actual product here: a content-to-style mapping with a rejection path.

The third difference is the material handling. The README gives explicit instructions for reference images: say whether you are borrowing colour, composition, type hierarchy, material, or cover impact, and say which elements you are not borrowing. It warns against saying only "follow this style" and asks for reusable visual features instead. That is a prompt discipline a general assistant will not enforce on you.

Extending the style library, and what maintenance costs

The repository is not archived, and its last push was on 2026-07-03. That is roughly two and a half months before the date of this article, so the codebase is recent but there is no release history to read: no releases were retrieved. Treat the 24 built-in styles as a starting library rather than a maintained catalogue.

Adding a style is documented as a template exercise. Each new style should supply a name, suitable content, unsuitable content, visual character, palette, fonts, composition, common elements, an image prompt template, negative prompts and example title types. The README says to update docs/style_system.md and templates/image_prompt_template.md, and points at templates/style_extension_template.md as the form to fill in, with examples/style_reference_notes.md added when needed.

The example files carry a similar expectation. The README states that examples should demonstrate reusable decisions rather than prose: an input topic, a content type judgment, a recommended style combination, page structure, at least five copyable image prompts, publishing copy and self-check results. The repository ships examples/example_output_plan.md, examples/example_output_yiwu_plan.md, examples/example_image_prompt.md, examples/example_input_topic.md and examples/style_reference_notes.md. If you fork this and change the style system, those example files are the maintenance surface, not the skill logic.

The licence is MIT. That permits commercial use, modification and redistribution provided the copyright notice and permission notice are included; it also means no warranty is provided. This is a description of the licence text, not legal advice, and if you plan to redistribute a modified style library inside a paid product you should read the LICENSE file in the repository yourself.

Editorial conclusion

Adopt it if you already publish Xiaohongshu carousels and want the style decision written down before any image is generated, because the ten-question intake and the style judgment report are the parts that survive a bad generation. Do not adopt it if you need unattended batch output: the workflow stops for your answers and then stops again for a single confirmation image, so it is a slow loop by design. Before you commit, open skill/SKILL.md and confirm your Codex setup accepts the skill/ directory as the skill root, then check templates/visual_review_checklist.md and decide whether its readability criteria match the text you actually put on covers. The repository's last push was on 2026-07-03, so treat the style library as a snapshot you extend yourself rather than one that grows on its own.

Frequently asked questions

What is XHS Visual Director and who is it for?

It is a Codex Skill for planning Xiaohongshu image posts, described in the README as a visual director rather than a copywriting assistant. It is aimed at Xiaohongshu carousel creators, AI and Agent content creators, personal brand operators, product managers, designers and founders.

How do I install XHS Visual Director as a Codex Skill?

The README states the installable skill body is in the skill/ directory, with skill/SKILL.md and skill/agents/openai.yaml. If your setup requires SKILL.md directly inside the skill directory, use xhs-visual-director-skill/skill/ as that directory; docs/, templates/ and examples/ are maintenance material.

Does XHS Visual Director generate the final images or only prompts?

The README says the default deliverable is image files, not prompt templates. The documented flow generates one cover or key-page confirmation image first, and after you approve style, composition, colour and information density it generates the full six to eight final pages with local paths and aspect-ratio checks.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. ziguishian/xhs-visual-director-skill on GitHub
Community notes

Community notes