
Xiaomi releases Xiaomi-Robotics-0 vision-language-action model
Xiaomi has introduced Xiaomi-Robotics-0, an advanced vision-language-action model designed for high-performance and smooth real-time execution. The open-source project aims to facilitate robotics research and development.
Published by Jin · 2 min read · 10 AUG 2026
- Xiaomi-Robotics-0
- Vision-Language-Action (VLA) Model
- Asynchronous execution for real-time robotic rollouts
Xiaomi has introduced Xiaomi-Robotics-0, an advanced vision-language-action model optimized for high performance and fast, smooth real-time execution. A vision-language-action model — often called a VLA — connects what a robot sees and hears with the physical actions it needs to perform. By combining visual understanding with robot control, these systems help machines interpret their surroundings and execute tasks in the real world.
According to technical reports, Xiaomi-Robotics-0 is first pre-trained on large-scale cross-embodiment robot trajectories and vision-language data. This training gives the model broad, generalizable action-generation capabilities without losing the visual-semantic knowledge stored in its underlying pre-trained vision-language model. During post-training, the developers use specialized techniques for asynchronous execution to address inference latency during real-robot rollouts.
To ensure continuous and seamless real-time operation during deployment, the system carefully aligns the timestamps of consecutive predicted action chunks. Extensive evaluations on simulation benchmarks and challenging bimanual manipulation tasks demonstrate state-of-the-art performance. Furthermore, Xiaomi-Robotics-0 can run smoothly on consumer-grade hardware while maintaining high success rates and throughput.
By making code and model checkpoints publicly available, the project provides developers and researchers with tools to explore embodied artificial intelligence and robotics control. The open-source release supports ongoing advancements in real-time robot execution and manipulation capabilities.
Verified details
In this report, we introduce Xiaomi-Robotics-0, an advanced vision-language-action (VLA) model optimized for high performance and fast and smooth real-time execution.
Source — huggingface.co ↗
Worth a read?
Comments · 0