# LingBot-VLA 2.0: Cross-Embodiment Vision-Language-Action Robot Foundation Model

> LingBot-VLA 2.0 is an open-source Python VLA foundation model designed to move from large-scale pre-training toward reliable deployment across multiple robot configurations. It is aimed at robotics researchers and engineers who need a single foundation model that covers arms, dexterous hands, and mobile bases without retraining from scratch for each embodiment.

**Robbyant/lingbot-vla-v2** — From Foundation to Application

- Repository: https://github.com/Robbyant/lingbot-vla-v2
- Stars: 988 · Forks: 113
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/robbyant-lingbot-vla-v2

## Cross-Embodiment Robot Control From One VLA Foundation Model

LingBot-VLA 2.0 addresses the cost of training a separate VLA model per robot platform. The project's design goal is a single foundation model that can be post-trained on a specific robot task without requiring a new pre-training run. The README compares it to a base language model that can be fine-tuned for downstream tasks: the pre-trained model carries cross-embodiment priors learned from a large heterogeneous corpus, and post-training specializes those priors to a target task and configuration.

The model targets three capability improvements over LingBot-VLA 1.0, according to the README: wider generalization across tasks and embodiments through a redesigned data pipeline; an expanded action space that covers arms, end-effectors, grippers, dexterous hands, waist, head, and mobile-base signals; and predictive dynamics modeling through dual-query distillation from LingBot-Depth and DINO-Video.

The primary audience is robotics researchers with GPU compute and access to robot demonstration data, plus engineers integrating the model into simulation or physical robot pipelines. The post-training example in the README uses RoboTwin 2.0 50-task datasets as a concrete case.

## Unified Action Space, MoE Expert, and Dual-Query Distillation

LingBot-VLA 2.0 maps heterogeneous robot configurations into a 55-dimensional canonical state/action vector. The dimensions break down as: 14 for arm joint position, 14 for end-effector pose, 2 for gripper position, 12 for hand joint position, 4 for waist position, 2 for head position, 3 for mobility signal, and 4 reserved dimensions. Any embodiment that can map its sensors and actuators into this canonical space can use the same foundation model without changing the action head.

The action expert uses sparse MoE (Mixture of Experts) layers. Fine-grained expert segmentation and shared expert isolation are used so that universal priors (shared across embodiments) and specialized patterns (unique to a task or configuration) coexist under the same active compute budget. The README states that this design improves cross-embodiment scaling.

Dual-query distillation appends current and future perceptual queries to the visual and text tokens. The current-state queries are distilled from LingBot-Depth (providing geometric cues); the future queries are distilled from DINO-Video (providing semantic temporal priors). The README describes this as encouraging causal inference that captures both current scene geometry and future scene evolution.

## Pre-Training Data: 60,000 Hours of Robot and Egocentric Videos

The pre-training corpus combines two streams. The robotic stream covers 50,000 hours of robot trajectories across 20 robot configurations, including single-arm, dual-arm, half-humanoid, humanoid, and egocentric sources. The egocentric stream provides 10,000 hours of manipulation-centric human videos with reconstructed hand trajectories.

Both streams go through filtering. The robotic stream removes video-state misalignment, blurry or occluded videos, multi-view misalignment, and episodes with abnormal velocity, acceleration, or jerk. The egocentric stream retains only manipulation-centric videos, standardizes hand trajectories, and filters unstable camera or hand-motion estimates.

The README describes the pre-training data as a large, heterogeneous corpus but does not link to a download for it. Post-training on custom tasks requires preparing a LeRobot v2.1 or v3.0 dataset directory, defining a feature mapping in a robot config YAML file, and computing normalization statistics over the dataset. The README provides pre-computed normalization statistics for the RoboTwin tasks at assets/norm_stats/robotwin.json.

## Installing LingBot-VLA 2.0 and Running Post-Training

LingBot-VLA 2.0 requires Miniconda or Anaconda, Python 3.12, and PyTorch 2.8.0. The setup script creates a conda environment and installs flash-attention. Before running the script, Conda must be initialized in your shell so that conda activate works.

```bash
git clone https://github.com/Robbyant/lingbot-vla-v2.git
cd lingbot-vla-v2

bash tools/create_train_env.sh
```

If a local flash-attention wheel is available, pass it explicitly to avoid a source compilation:

```bash
bash tools/create_train_env.sh \
  --flash-attn-wheel /path/to/flash_attn-2.8.3+cu12torch2.8cxx11abiTRUE-cp312-cp312-linux_x86_64.whl
```

Download the pre-trained 6B weights from Hugging Face:

```bash
python3 scripts/download_hf_model.py --repo_id robbyant/lingbot-vla-v2-6b --local_dir lingbot-vla
```

To run post-training on the RoboTwin 2.0 50-task example:

```bash
bash train.sh tasks/vla/train_lingbotvla.py ./configs/vla/robotwin/robotwin.yaml \
  --data.train_path assets/training_data/robotwin.txt \
  --data.data_name multi \
  --train.output_dir output/
```

Post-training also requires weights from Qwen3-VL-4B-Instruct, MoGe-2-vitb-normal, LingBot-Depth, and DINO-VIDEO. The README points to Training_Config.md for the full configuration details.

## Hardware Requirements and Known Limitations

The 6B-parameter model requires significant GPU memory for inference and more for training. The setup script installs flash-attention 2.8.3, which requires a CUDA-capable GPU with a compatible compute capability and a specific PyTorch 2.8.0 build. The requirements.txt pins exact versions for torch (2.8.0), torchvision (0.23.0), torchaudio (2.8.0), and many other packages, which means environment conflicts are likely if these exact versions cannot be installed.

Post-training preparation involves three steps: preparing a LeRobot dataset, defining a robot config YAML, and computing normalization statistics. Each step requires understanding the LingBot-VLA data format. The README provides a complete example only for RoboTwin 2.0; other robot configurations require adapting the robot config file from the provided example.

The repository has no GitHub releases. Versioning is handled through the news section in the README. The Distributed Muon optimizer, added on September 22, 2026, is described as improving single-node H20 post-training speed by 37.6% over the standard Muon optimizer, according to the README's News section. The earlier RoboTwin post-training weights are available separately at huggingface.co/robbyant/lingbot-vla-v2-6b-robotwin.

## LingBot-VLA 2.0 Versus OpenVLA

OpenVLA is an open-source 7B-parameter VLA model from Stanford trained on single-arm manipulation tasks from the Open X-Embodiment dataset. OpenVLA takes a simpler architecture approach: a Prismatic VLM backbone with a fine-tuned action head, without an explicit cross-embodiment design or a unified multi-dimensional canonical action space.

The difference in approach is scope. OpenVLA is optimized for single-arm dexterous manipulation with a relatively straightforward architecture that maps directly onto established VLM fine-tuning workflows. LingBot-VLA 2.0 is larger in ambition: the 55-dimensional unified action space, MoE action expert, and dual-query distillation are all motivated by the goal of handling diverse embodiments without per-robot pre-training.

For a team deploying to a single robot type that matches the OpenVLA training distribution, OpenVLA's smaller scope and simpler setup may be more practical. For teams building a unified model that must serve multiple robot configurations or that need the predictive dynamics signals from LingBot-Depth and DINO-Video, LingBot-VLA 2.0's design addresses requirements that OpenVLA does not cover.

## Apache-2.0 License and Project Activity

LingBot-VLA 2.0 is licensed under Apache-2.0, permitting commercial and proprietary use without requiring derivatives to be open source. The LICENSE file is at the top level of the repository.

The last push was on September 22, 2026, making this an actively updated project. The repository has no tagged releases; version tracking is through the News section and Hugging Face model checkpoints. The technical report is at arxiv.org/abs/2607.06403. Pre-trained weights for the base model (lingbot-vla-v2-6b) and the RoboTwin post-training variant (lingbot-vla-v2-6b-robotwin) are on Hugging Face and ModelScope.

## Conclusion

LingBot-VLA 2.0 suits robotics researchers and engineers who need a single pre-trained VLA model that handles diverse embodiments through a unified action representation, and who are prepared to work through a complex post-training pipeline with GPU hardware and large datasets. Teams with a single-arm manipulation focus and no need for cross-embodiment generalization may find a lighter model like OpenVLA sufficient. Verify GPU availability before starting: the 6B-parameter model and the flash-attention requirement make CPU-only use impractical at any realistic inference or training scale.

## FAQ

### What robot embodiments does LingBot-VLA 2.0 support?

The unified 55-dimensional action space covers arms (14 dims for joint position), end-effectors (14 dims for pose), grippers (2 dims), dexterous hands (12 dims for joint position), waist (4 dims), head (2 dims), and mobile base (3 dims). The README states the pre-training corpus includes single-arm, dual-arm, half-humanoid, humanoid, and egocentric sources across 20 robot configurations.

### Why does LingBot-VLA 2.0 use a Mixture of Experts action head?

According to the README, sparse MoE layers inside the action expert allow universal cross-embodiment priors and specialized task or embodiment patterns to coexist under the same active compute budget. Fine-grained expert segmentation and shared expert isolation are the two mechanisms described as enabling this separation.

### What is dual-query distillation in LingBot-VLA 2.0?

Dual-query distillation appends two sets of perceptual queries to the visual and text tokens: current-state queries distilled from LingBot-Depth for geometric cues, and future-state queries distilled from DINO-Video for semantic temporal priors. The README describes this as encouraging causal inference that captures both current scene geometry and anticipated future scene changes.

## Sources

- [Issues](https://github.com/Robbyant/lingbot-vla-v2/issues)
- [License: Apache-2.0](https://github.com/Robbyant/lingbot-vla-v2/blob/main/LICENSE)
- [README](https://github.com/Robbyant/lingbot-vla-v2/blob/main/README.md)
- [Robbyant/lingbot-vla-v2 on GitHub](https://github.com/Robbyant/lingbot-vla-v2)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/robbyant-lingbot-vla-v2
