Library / SDK
ziguishian/xhs-visual-director-skill avatar
ziguishian/xhs-visual-director-skill

XHS Visual Director: An Agent Skill for Xiaohongshu Image Post Planning

这是一个用于规划小红书图文的 Agent Skill。它不是普通文案助手,而是一个“视觉导演”:先判断内容任务,再选择适合的视觉风格,最后输出完整图文结构、逐页视觉方案、图像生成提示词、发布文案和自检清单。

1,378 stars127 forksUnknownMIT

At a glance

What is it?
XHS Visual Director is a Codex agent skill that automates the visual planning workflow for Xiaohongshu (RedNote) image posts. It applies a 10-question Socratic method to define content goals and visual style, then produces a multi-page layout plan, per-page image generation prompts, and publication copy.
Who is it for?
XHS Visual Director is useful for Xiaohongshu content creators, AI-focused writers, and personal brand managers who want a structured visual planning workflow rather than improvising layout decisions. It is not suited for creators who need batch content generation or automated posting: the skill handles planning and prompts, not scheduling or platform integration.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 91 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What XHS Visual Director Does and Who It Is For

XHS Visual Director is an agent skill for planning and generating image posts on Xiaohongshu, a Chinese social media platform where multi-page image carousels are a primary content format. The skill is designed for content creators who find that AI-generated images look generic or like low-cost templates, and who want a structured workflow that produces images with a clear visual identity.

The README lists the target audience as Xiaohongshu image creators, AI and Vibe Coding content creators, personal brand managers, product managers, designers, and entrepreneurs who want to reformat existing materials (drafts, screenshots, product images, slides, or page designs) into high-quality image posts.

The skill's central claim is that it acts as a visual director rather than a text assistant: it makes style decisions before generating anything, explains why a particular style fits the content, and rejects styles that would undermine the post's visual authority. The README names avoiding cheap AI science-fiction aesthetics, PPT-style layouts, and unreadable text as explicit design goals.

The Socratic Questioning Flow: 10 Questions Before Any Image

Before producing any layout or image, the skill asks 10 clarifying questions. These cover: the communication goal, the target audience, the core message, available materials, preferred visual direction, prohibited elements, and conversion or engagement targets.

For complex commercial, product, or personal brand projects, the same 10-question flow applies. If the answers still leave key decisions unresolved, the skill asks up to 5 more follow-up questions. The README explicitly states that the skill does not generate images immediately from a topic alone.

After the user answers the 10 questions, the skill outputs a summary of the answers and a set of generation assumptions. It then produces a style assessment report that names the primary visual style, one or two supporting styles, styles that are explicitly rejected, and the reasoning for each choice. This reasoning step is what separates the skill's workflow from a prompt template: the style decision is explained, not just applied.

An example from the README shows a topic about Vibe Coding: the style assessment names a dark technology magazine aesthetic as primary, rejects a liquid glass aurora style, and explains that the topic requires cognitive impact and professional authority, which a dreamlike aesthetic would undermine.

Installing and Invoking the Skill

The skill lives in the `skill/` directory of the repository. The key files are:

text
xhs-visual-director-skill/skill/SKILL.md
xhs-visual-director-skill/skill/agents/openai.yaml

If your Codex skill management requires a directory that contains `SKILL.md` directly, use `xhs-visual-director-skill/skill/` as the skill directory. The project's `docs/`, `templates/`, and `examples/` directories are maintenance and extension materials, not part of the installable skill.

Once installed, invoke the skill by giving the agent a topic and a content goal. The README shows this example input:

text
主题:为什么普通人现在必须学习 Vibe Coding?
目标:做成 8 页小红书图文,风格要高级、有科技感,但不要像廉价 AI 模板。

The skill begins with the 10 Socratic questions. After the user answers them, the skill outputs the full workflow: style report, three style options with a recommendation, a 6-8 page structure plan, a unified visual template, one sample image for confirmation, and then the complete final image set after approval.

Visual Style System and Image Generation Mechanism

The style system is documented in `docs/style_system.md`. Each style entry specifies: a name, suitable content types, unsuitable content types, visual character, color palette, typography, composition approach, common visual elements, image prompt templates, negative prompts, and example title types.

The image generation flow is two-stage. The skill first generates one sample or cover image and waits for user confirmation of style, composition, color, and information density. After approval, it generates the complete 6-8 page set and outputs each image's local path and aspect ratio check results.

The README documents a limitation specific to AI image generation: if the image model produces unstable or unreadable Chinese text, the skill will generate a high-quality visual base with a clear text safe zone and recommend adding real Chinese text in post-processing. The skill still delivers image files rather than stopping at prompt templates.

The README's checklist for acceptable output requires: every page serving a single primary communication goal, each page carrying progression and collectible value, and every prompt specifying aspect ratio, layout, text zones, typography, color, main visual, negative space, and prohibited elements.

Limitations: Platform Coupling and Content Scope

The skill is designed specifically for Xiaohongshu's 3:4 portrait format and the platform's content culture. The README describes the default output orientation as 3:4 cover direction: high-end, legible, mobile-readable, and avoiding generic AI template aesthetics. Teams creating content for other platforms with different aspect ratios or style expectations would need to adapt the style system and prompts in `templates/image_prompt_template.md`.

The skill handles planning and image generation prompts. It does not integrate with Xiaohongshu's API or any scheduling tool. The README does not mention automated posting, analytics, or feedback loops from actual post performance. A creator must manually upload the generated images and copy to the platform.

The last push to the repository was on 2026-07-03. The repository has no GitHub releases.

Extending the Style System with New Visual Styles

The README documents the extension process explicitly. Adding a new style requires completing all of the following fields: style name, suitable content, unsuitable content, visual character, color palette, typography, composition, common elements, image prompt template, negative prompts, and example title types.

After defining the new style, the files that must be updated are `docs/style_system.md`, `templates/image_prompt_template.md`, and optionally `templates/style_extension_template.md` as a fill-in template for the new style definition. If the new style needs reference imagery, a `examples/style_reference_notes.md` entry is also expected.

The extension process is designed for individual maintainers adding their own visual preferences as reusable workflows, not for collaborative multi-contributor development. The README notes that examples are not meant to showcase writing style but to document reusable decisions, so each example should include the input topic, content type judgment, recommended style combination, page structure, at least 5 copyable image prompts, publication copy, and a self-check result. The MIT license means the style definitions can be forked and adapted without restriction.

Editorial conclusion

XHS Visual Director is useful for Xiaohongshu content creators, AI-focused writers, and personal brand managers who want a structured visual planning workflow rather than improvising layout decisions. It is not suited for creators who need batch content generation or automated posting: the skill handles planning and prompts, not scheduling or platform integration. Before using it, read `docs/style_system.md` to understand the built-in style options and check `templates/image_prompt_template.md` before adding new styles.

Frequently asked questions

What content types work best with XHS Visual Director?

The README lists three use-case examples: AI, agent, and technology opinion content (suited to a dark technology magazine aesthetic), commercial and industry content for business or e-commerce (suited to a high-end business proposal aesthetic), and tutorial or workflow documentation with real screenshots (suited to a phone-screenshot or Notion card aesthetic). The skill asks 10 questions to identify which category applies before suggesting styles.

How does XHS Visual Director handle Chinese text in generated images?

The README states that if the image model produces unstable or unreadable Chinese text, the skill will prioritize generating a high-quality visual base with a clear text safe zone, then recommend adding real Chinese text in post-processing. The skill still delivers image files in this case rather than stopping at prompt templates.

Can XHS Visual Director be used for non-Xiaohongshu platforms?

The skill is built around Xiaohongshu's 3:4 portrait format and the platform's content culture. The README does not describe support for other aspect ratios or platforms. Adapting it would require modifying the style definitions in `docs/style_system.md` and the prompt templates in `templates/image_prompt_template.md`.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. ziguishian/xhs-visual-director-skill on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ziguishian-xhs-visual-director-skill.svg)](https://hysenlabs.com/projects/ziguishian-xhs-visual-director-skill)