NavFoM - A large-scale model of a surround-view navigation base launched by Galaxy Universal.
NavFoM (Navigation Foundation Model) is the world's first cross-ontology, full-domain surround-view navigation foundation model, jointly released by Galaxy General and teams from Peking University, the University of Adelaide, Zhejiang University, and others. It features full-scene support...
What is NavFoM?
NavFoM (Navigation Foundation Model) is the world's first cross-entity, all-domain surround-view navigation foundation model, jointly released by Galaxy General and teams from Peking University, the University of Adelaide, and Zhejiang University. It boasts full-scene support capabilities, applicable to both indoor and outdoor environments, and can achieve zero-shot operation in unseen scenarios. NavFoM supports various navigation tasks, such as natural language command-driven target following and autonomous navigation, and can quickly adapt to different entities such as robot dogs, wheeled humanoid robots, drones, and automobiles. Its core technologies include TVI Tokens and the BATS strategy, establishing a new universal paradigm: "video stream + text command → motion trajectory," completing the entire navigation process end-to-end.
Main functions of NavFoM
-
Full-scenario supportNavFoM can support both indoor and outdoor scenes, and can achieve zero-sample operation in unseen environments without the need for additional mapping or data collection, making it highly adaptable to different environments.
-
Multitasking supportThe model supports various sub-navigation tasks, such as target following and autonomous navigation driven by natural language commands, and can complete corresponding navigation actions according to different commands.
-
Cross-body adaptationNavFoM can quickly and cost-effectively adapt to heterogeneous bodies of different sizes, such as robot dogs, wheeled humanoids, legged humanoids, drones, and automobiles, making it widely applicable.
-
Technological innovationNavFoM uses TVI Tokens (Temporal-Viewpoint-Indexed Tokens) to help the model understand time and direction, and the BATS strategy (Budget-Aware Token Sampling) to keep the model intelligent even with limited computing power. These technological innovations improve the model's performance.
-
Unified ParadigmNavFoM establishes a brand-new universal paradigm: "video stream + text command → action trajectory". It no longer relies on modular splicing, but completes the entire process of "seeing - understanding - acting" end-to-end, simplifying the navigation process.
-
Dataset ConstructionNavFoM has built a massive cross-task dataset containing approximately eight million cross-task and cross-ontology navigation data points, as well as four million open question-answering data points, providing rich data support for model training.
NavFoM's technical principles
-
TVI Tokens (Temporal-Viewpoint-Indexed Tokens)By using time and viewpoint indexes, the model can understand time and direction, thus better handling navigation tasks in dynamic environments.
-
BATS Strategy (Budget-Aware Token Sampling)Under conditions of limited computing power, a budget-aware labeling and sampling strategy can be used to ensure that the model can still run efficiently, thereby improving its feasibility in practical applications.
-
End-to-end general paradigmIt adopts the paradigm of "video stream + text command → action trajectory" to integrate visual input, language command and action output into a unified framework, realizing direct mapping from perception to action.
-
Cross-task datasetA massive cross-task dataset containing approximately eight million navigation data points and four million open question-answering data points was constructed, providing rich multi-scenario and multi-task data support for model training and improving the model's generalization ability.
NavFoM's project address
The relevant address has not yet been announced.
Application scenarios of NavFoM
-
Robot NavigationIn complex environments, such as shopping malls and airports, robots can navigate autonomously and follow targets based on natural language commands, enabling efficient service and guidance functions.
-
autonomous driving: Applied to autonomous driving systems in automobiles, it enhances the vehicle's autonomous decision-making and navigation capabilities in complex road conditions, thereby improving the safety and reliability of autonomous driving.
-
Drone navigationTo provide drones with autonomous navigation capabilities, enabling them to fly autonomously and perform missions in complex terrains and environments, such as logistics delivery and environmental monitoring.
-
humanoid robotIt supports humanoid robots of different forms, such as wheeled and legged humanoids, enabling them to better adapt to various environments and complete complex navigation and interaction tasks.
-
Development Application ModelDevelopers can use NavFoM as a base to further develop application models that meet specific navigation requirements through post-training, thus expanding its application scope in different fields.