Model or dataset
sou350121/VLA-Handbook avatar
sou350121/VLA-Handbook

A VLA handbook organised around the gap between reading and running

本项目旨在为致力于进入VLA(Vision-Language-Action)领域的算法工程师提供一份全中文、实战导向的学习/面试手册。 不同于通用的 CV/NLP 面试指南,本项目聚焦于 Robotics 特有的挑战

675 stars42 forksHTMLCC-BY-4.0

At a glance

What is it?
VLA-Handbook is an all-Chinese, engineering-first knowledge base for vision-language-action work that pairs theory write-ups with distilled field notes from Xiaohongshu, Hugging Face blogs, vendor posts, a Discord and 47 GitHub issues, updated by a scheduled pipeline, with RSS feeds and a separate handbook for world representation.
Who is it for?
It fits an algorithm engineer who already knows computer vision or language modelling and now needs the robotics specifics: action representation, Sim2Real gaps, hardware choices and the community knowledge that never reaches a paper appendix. Two things to weigh.
Can I use it commercially?
Yes, with credit. CC-BY-4.0 allows commercial use as long as you credit the authors and indicate what you changed. It is written for creative content, so check how it applies to any code.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Each page carries an entry script and a shape check

The stated pitch is narrow: the gap between understanding a paper and running its code. The claim is that each write-up names the entry script, the key hyperparameters and the shape conventions, with a sanity check for each.

That is a specific and unusual promise for a handbook. Most summaries of a model describe the architecture; describing the tensor shapes and the entry point is describing how to launch it, which is the part a new arrival to the field actually needs.

The second claim is that the robotics work gets written down. Named explicitly: multimodal synchronisation, Sim2Real breakpoints, action space alignment, and tactile and dexterous hand hardware selection. Each of those is a real category of failure that computer vision or language tutorials never touch, because they are not software problems.

The third claim is that the knowledge base is alive rather than a static document set, with a daily pipeline fetching new VLA papers, rating them, and writing deep analyses back into the repository.

The last two claims are about maintenance, which is where a handbook of this kind usually fails, so they are the ones to check first.

Field notes distilled from three sources on three schedules

The community field notes are described as the most distinctive part of the project, and they come from three places with different cadences.

The Chinese notes are distilled from Xiaohongshu, 300-plus entries covering posts one to two hundred, plus a traceable index of 40 entries and a 28-entry dictionary of community jargon, refreshed by automatic incremental update every three days. The English notes gather 165 traceable entries from Hugging Face blog posts, vendor blogs and a LeRobot Discord, updated on Fridays at 11:00. The GitHub notes are 47 high-interaction issues distilled from six core repositories, scanned automatically every week.

Named topics give a sense of what is in there: measured training costs for GR00T N2 with ACT, GPU compatibility matrices covering the RTX 50 series and Jetson, fine-tuning pitfalls with Pi0, memory optimisation for GR00T, root causes of training convergence failure, real-hardware reinforcement learning for a released model, and measured inference latency.

The Xiaohongshu notes also carry non-technical items such as salary ranges and a published national standard, which is a signal about the mix: some of this is engineering lore and some is community chatter, and the handbook does not separate them by weight.

The counts describe different things at different dates

Several totals appear in the same document, and they do not all measure the same thing.

The current self-description gives 525 theory documents spread across 10 topic directories with two to three new deep analyses added daily, 165 English community field notes, more than 300 Chinese community entries, 47 GitHub issue write-ups, 24 biweekly inference reports, and the industry intelligence layer. A dated update note from 11 June 2026 gives a different figure: 10 mainline topic files were brought up to date with roughly 109 new deep analyses from April to June, keeping older judgements while explicitly recording tensions, 71 scattered articles were filed properly, more than 200 dead internal links were repaired, and the library grew from 254 to 380 documents.

So 380 was the library after a cleanup, and 525 is the theory count later on. Neither is wrong, and both are accurate at their dates, but a reader comparing a figure today with a figure quoted from the web months ago will draw the wrong conclusion.

The maintenance record is also visible: a daily pulse document tracking paper flow and acceleration across 15 VLA method families with 30-day trend charts, generated every day, and a changelog at the root.

The last push is dated 1 October 2026 and there are no tagged releases, so the only way to date a claim is the commit history.

Four RSS feeds, one OPML file, times in Beijing time

The distribution layer is unusually complete for a documentation project. Four feeds are published: one for new theory articles under theory/, one for daily signals carrying the lightning and wrench rated papers plus the state-of-the-art table while filtering out the other two ratings, one for AI agent ecosystem items, and one for weekly and biweekly reports.

There is an OPML file for subscribing to all of them at once, with Feedly, Inoreader and NetNewsWire named as readers that accept it, and a subscription guide covering per-reader walkthroughs, the CC BY 4.0 attribution requirement and a FAQ.

The schedule table is given in Beijing time, which tells you both the cadence and the audience. Paper scoring runs daily between 09:15 and 10:00, and social intelligence collection runs daily at 09:30. The community notes follow their own clocks, with the Chinese set every three days and the English set on Fridays.

So the whole system is a daily job, not a weekly one. For a field producing dozens of papers a day, that is the difference between a handbook that helps you this week and one that helped you last quarter.

One adjacent project deserves a mention because it targets a different reader: a separate skill repository that turns an AI coding assistant into a VLA expert, with Claude Code, Cursor, Codex and OpenCode named as supported tools.

The sister handbook covers world models, not action policies

There is a deliberate division of territory between this repository and a sister one, and the boundary is drawn in a way that tells you what the author thinks matters.

The split is: VLA handles the action policy, and the sister repository, Spatial-Intelligence-Handbook, handles world representation. Named there are 3D Gaussian splatting, a geometry foundation model for visual grounding, depth foundation models, and a cross-embodiment comparison, with the intersection of the two projects identified as 3D-aware VLA.

That framing is useful because it says which half a given paper belongs to. A paper that improves spatial reconstruction is a different literature from one that improves how actions are emitted, and a beginner reading both will otherwise assume they are one field.

The main repository also links a live site that updates daily at 12:00 Beijing time, and a pulse document with paper traffic and acceleration per method family plus 30-day trend charts.

A separate direction list points at specific entry pages for people who know what they want: an open training stack reproduction on a Unitree robot, a line-by-line dissection intended for readers who want to follow every shape transform, a zero-shot world action model write-up, tactile and force alignment, knowledge distillation for edge deployment, and a full learning roadmap.

A search step then a mechanical verification step

The industry intelligence layer is a daily scan of robotics and embodied computing companies, covering funding, products, IPOs and partnerships, and it is described as writing into company profiles rather than sitting in a separate archive.

The method is worth noting because it is explicit about having two stages. A search step using a web-search-capable model collects candidate items, and then a mechanical verification step runs before anything is appended to a company profile. A weekly industry judgement map is produced on Fridays, and the raw records live in an archive directory under memory/blog/archives/.

The two-stage design is the right shape for this kind of monitoring, because the failure mode is a plausible but wrong claim about a funding round. Automating the collection is easy and the verification is where the work is, so naming both stages is more informative than claiming automation alone.

What the layer cannot do is tell you which judgements were revised later. The judgement map is a weekly snapshot, and the company profiles are append-only in practice, so a reader looking for the current view needs to read both and reconcile them.

A 30-minute path that starts with action representation

The reading order is prescribed by dependency, with time budgets, and it starts where the real fork in the field is.

The first page, five minutes, builds the global picture: input as vision plus language, then a backbone, then an action head, then robot actions, traced through the evolution from RT-1 to RT-2 to OpenVLA to pi0. The second page, ten minutes, answers the question that page raises, which is how the action head actually emits actions. The answer is three paradigms: discrete tokens, which are fast but coarse; diffusion, which is precise but slow; and flow matching, which is described as both fast and precise.

That third label is the claim to test, and the next page is where the project argues for it, contrasting a straight ODE path against diffusion's curved denoising and citing 5 to 20 inference steps as what makes 50 Hz control possible.

Then the view pulls back: why ACT and diffusion policy remain the baseline, and how three improvement lines, data scaling, perception enhancement and reinforcement learning post-training, intersect.

The concrete claims in the community notes are the ones to check against your own experience, such as ACT working from 50 episodes, and the observation that Sim2Real failures are usually uncalibrated physics parameters rather than a policy problem.

It positions itself against four reading habits

The comparison table is worth reading for what it concedes, because each competing habit is credited with what it is genuinely good at.

WeChat public accounts from established Chinese outlets are credited with readable Chinese surveys, reliable editing and suitability for reading on a phone. GitHub awesome lists and public surveys are credited with curated bookmarks and fast entry to classic papers and open source projects. Following authors on X is credited with real-time reactions and discussion. Xiaohongshu is credited with first-hand failure reports, reproduction parameters and post-mortems from people doing the work, with the observation that the comment sections are more useful than the posts.

Then the weaknesses. WeChat links expire after ninety days. Awesome lists are static. X buries old threads in the algorithm. Xiaohongshu is hard to search and posts sink.

The handbook's answer to each is archival permanence and daily refresh: git history, full-text search, and automated collection that distils the Xiaohongshu experience into the repository rather than leaving it to be searched.

What that argument does not address is authority. A distilled social post and a peer-reviewed result end up in the same repository with the same file convention, and only the reader's scepticism separates them.

Editorial conclusion

It fits an algorithm engineer who already knows computer vision or language modelling and now needs the robotics specifics: action representation, Sim2Real gaps, hardware choices and the community knowledge that never reaches a paper appendix. Two things to weigh. The corpus is Chinese, and the community notes are distilled from sources with very different reliability, including social posts whose claims are the kind that need verifying before you change a training recipe. And the freshness comes from automation, with daily paper scoring and scheduled community scans, so the update schedule is a promise about a pipeline rather than about a person. Read the theory pages and then check a field note against the source before you trust a number.

Frequently asked questions

What does VLA stand for in AI?

Vision-Language-Action. In this handbook the pipeline is described as visual plus language input going through a backbone into an action head that produces robot actions. The lineage it traces runs from RT-1 through RT-2 and OpenVLA to pi0.

How does VLA work?

The action head emits actions in one of three paradigms according to the handbook: discrete tokens, which are fast but coarse; diffusion, which is precise but slow; and flow matching, which the project argues is both fast and precise, with 5 to 20 inference steps cited as what makes 50 Hz control possible.

What is VLA-Handbook?

It is an all-Chinese, engineering-oriented knowledge base for algorithm engineers entering vision-language-action work. Each theory page aims to give the entry script, key hyperparameters and shape sanity checks, and the repository is licensed CC BY 4.0 with no tagged releases and a last push dated 1 October 2026.

Where do the community field notes in VLA-Handbook come from?

Three sources. Chinese notes distilled from Xiaohongshu with more than 300 entries, a 40-item traceable index and a 28-entry jargon dictionary, updated every three days. English notes with 165 traceable entries from Hugging Face blog posts, vendor blogs and a LeRobot Discord, updated on Fridays at 11:00. And 47 high-interaction GitHub issues distilled from six core repositories, scanned weekly.

How often is VLA-Handbook updated?

The schedule is given in Beijing time: paper scoring runs daily between 09:15 and 10:00, social intelligence collection runs daily at 09:30, new theory articles arrive daily at two to three deep analyses, an industry judgement map is produced on Fridays, and an industry radar scans company funding, product and IPO news every day. Four RSS feeds and an OPML file are published for subscription.

Official sources

  1. Issues
  2. License: CC-BY-4.0
  3. README
  4. sou350121/VLA-Handbook on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sou350121-vla-handbook.svg)](https://hysenlabs.com/projects/sou350121-vla-handbook)