GO-1 - The first universal embodied base model launched by Zhiyuan Robotics
GO-1 (Genie Operator-1, the first universal embodied base model launched by Zhiyuan Robotics. The model adopts the Vision-Language-Latent-Action (ViLLA) architecture, consisting of VLM (Vision-Language-Latent-Action)...
What is GO-1?
GO-1 (Genie Operator-1, the first general-purpose embodied base model launched by Logic Robotics. The model adopts the Vision-Language-Latent-Action (ViLLA) architecture, consisting of a VLM (Multimodal Large Model) and a MoE (Hybrid Expert). The VLM leverages massive amounts of internet image and text data to endow the model with general scene perception and language understanding capabilities; the Latent Planner in the MoE obtains general action understanding capabilities through a large amount of cross-ontology and human operation video data; and the Action Expert, based on millions of real-device data points, achieves precise action execution.
GO-1's main functions
- Human video learningBy analyzing a large amount of video data of human actions, the model can learn and understand motion knowledge in the real world and quickly adapt to new tasks.
- Fast generalization with small samplesWith very little data or zero samples, GO-1 can quickly generalize to new scenarios and tasks, lowering the application threshold of embodied intelligence.
- One brain, multiple forms, cross-ontology applicationsGO-1 can be flexibly deployed on different types of robot bodies, supports multiple robot forms, and demonstrates extremely high versatility and flexibility.
- Continuous EvolutionIn practical use, GO-1 can continuously learn and optimize its own performance, and through the data feedback system, it can continuously evolve from the problem data encountered in actual execution, becoming smarter the more it is used.
- High-efficiency action executionAction Expert, trained on millions of real device data sets, possesses sophisticated and efficient action execution capabilities.
The calculation principle of GO-1
- VLM (Multimodal Large Model)VLM (Visual Modeling) leverages massive amounts of internet image and text data to endow models with exceptional general scene perception and language understanding capabilities. It can accurately identify and understand information within images, while efficiently fusing it with text data to achieve a comprehensive understanding of complex scenes.
- MoE (Hybrid Expert System)The MoE system further enhances the model's ability to understand and execute actions. Specifically:
- Latent PlannerBy analyzing a large amount of cross-ontology and human operation video data, we have mastered the general motion planning logic.
- Action ExpertIt is trained on millions of real machine data and has precise and efficient action execution capabilities.
GO-1 project address
- Project official website: https://agibot-world.com/blog/go1
- GitHub repositoryhttps://github.com/OpenDriveLab/AgiBot-World
- HuggingFace model libraryhttps://huggingface.co/agibot-world/GO-1
- Technical Papers: https://agibot-world.com/blog/agibot_go1
Application scenarios of GO-1
- Retail servicesIn a retail environment, GO-1 can be deployed as a service robot to provide services such as customer guidance, product inquiry, and checkout assistance.
- Reception and ConsultationIn places such as hotels, restaurants, or office buildings, GO-1 can serve as a reception robot, providing services such as information consultation, reservation confirmation, and directional guidance.
- Production line supportIn manufacturing, GO-1 can assist in completing repetitive tasks on the assembly line, such as parts handling and assembly.
- Household helperIn a home environment, GO-1 can serve as a household helper, assisting with daily chores such as cleaning and tidying.
- Scientific explorationGO-1 can be used in scientific research, such as for sample collection and data analysis in extreme environments.