Robbyant's Open-Sourcing of LingBot-VLA 2.0: A Step Towards Universal Robot Models
Robbyant, an embodied AI business within Ant Group, has recently open-sourced LingBot-VLA 2.0, a vision-language-action model designed to work across various robot types without the need for retraining. This move is significant as it addresses a critical challenge in robotics: the need for models that can adapt to different robot morphologies and movement patterns.
The Challenge of Cross-Morphology Testing
The crux of the issue lies in cross-morphology testing, where models are tested on different robot designs. Many existing models are tied to specific robots and tasks, making deployment costly and complex when customers switch hardware or expand into new environments. Robbyant's LingBot-VLA 2.0 aims to overcome this by being trained on a diverse set of data.
Training Data and Benchmark Results
The model was trained on an impressive 60,000 hours of real-world physical data, including 50,000 hours of cleaned robot interaction data and 10,000 hours of distilled first-person human manipulation data. This extensive training set covered 20 robot morphologies from 17 manufacturers, showcasing the model's versatility. Benchmark results were impressive, with LingBot-VLA 2.0 outperforming π0.5 and GR00T N1.7 on various platforms, including dual-arm manipulation and long-horizon mobile manipulation tests.
Deployment Efficiency and Commercial Applications
Robbyant emphasizes deployment efficiency, offering a version of LingBot-VLA 2.0 tuned for post-training efficiency, achieving latency below 130 milliseconds on an RTX 4090. This is crucial for reducing the cost of putting robotics software into production, as it minimizes the need for task-specific and machine-specific retraining. The model is already being tested in commercial pilot projects, with partners like Leju, Ti5 Robot, GuoDa Drugstore, and Longsheng Technology.
Expanding the Software Stack
Robbyant's strategy goes beyond a single-purpose model. They have also released LingBot-Depth 2.0 and LingBot-Vision, addressing spatial perception and visual perception challenges, respectively. These models, trained on vast datasets, demonstrate Robbyant's commitment to building a comprehensive software stack for embodied AI. The collaboration with Orbbec further reinforces this vision, combining spatial perception software with Orbbec camera hardware.
The Future of Universal Robot Models
The open-sourcing of LingBot-VLA 2.0 is a significant step towards universal robot models, where a single model can adapt to various robot types and tasks. This approach aligns with the industry's push for standardized data ecosystems and reusable software. As Robbyant continues to innovate, the robotics sector moves closer to a future where robots can seamlessly transition between different environments and applications.
In my opinion, Robbyant's open-sourcing strategy is a bold move that could accelerate the development of universal robot models. The company's focus on deployment efficiency and comprehensive software stack development positions them as a leader in the field. As the robotics industry continues to evolve, Robbyant's contributions will undoubtedly play a crucial role in shaping the future of embodied AI.