Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation

ViGAR: Visual Goal-conditioned Action Reasoning

Shukai Gong1* · Xuanran Zhai2* · Yintianrun Zhang1* · Ruopeng Cui2 · Ye Huang1 · Yiyang Fu1 · Dexuan Lyu2
Chaojie Li2 · Xinyi Song2 · Peiwen Lin2 · Chuang Wang2 · Mingyuan Jia3 · Yufan Deng1
Jiaxin Fang3 · Bo Liang1 · Jiaxin Li1 · Yuxiang Gao3† · Hao Liu2† · Daquan Zhou1†

1 Peking University   2 AgiBot   3 CocoMatrix

* Equal contribution   † Corresponding author

Project Page arXiv HF Daily Paper GitHub Python PyTorch License: OpenMDW-1.1

ViGAR teaser: task-diverse and long-horizon data train a subgoal planner and a world-action policy, enabling long-horizon manipulation by subtask decomposition and in-context learning on unseen tasks

📑 Todo List

  • Evaluation code and checkpoints on RoboTwin.
  • Training code and data on RoboTwin.
  • 900-hour real-robot dataset with subtask annotations.

🎬 Demo

⚙️ Getting Started

git clone https://github.com/DAGroup-PKU/ViGAR.git
cd ViGAR
Step Guide
Set up the environments Installation
Download weights and data Download
Evaluate on Simulation Benchmark Evaluation
Train the policies and subgoal planners in ViGAR Training

💪 Repository

vigar/            Policy model, training and serving
subgoal_planner/  Planner runtime and training implementation
simulation/       RoboTwin tasks and simulator setup
data/             Annotation, goal generation and action conversion
evaluation/       Policy/planner services and closed-loop evaluation
scripts/          Installation and launch commands
configs/          Task and inference settings
docs/             Installation, download, evaluation and training guides

The policy uses VeOmni with a bundled Cosmos runtime. Planner processes select their own Cosmos runtime through PYTHONPATH; they run separately from policy and simulator processes.

🙏 Acknowledgements

Built on NVIDIA Cosmos, VeOmni, RoboTwin, Wan2.2 and Qwen3-VL.

✏️ Citation

@article{gong2026rethinking,
  title={Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation},
  author={Gong, Shukai and Zhai, Xuanran and Zhang, Yintianrun and Cui, Ruopeng and Huang, Ye and Fu, Yiyang and Lyu, Dexuan and Li, Chaojie and Song, Xinyi and Lin, Peiwen and others},
  journal={arXiv preprint arXiv:2610.02368},
  year={2026}
}

License

Original project code uses the MIT license. Bundled components retain their upstream licenses. Cosmos-derived model materials use OpenMDW-1.1.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Paper for DAGroup-PKU/ViGAR