AB
AiBoss
project

Ferret-UI 2 - Apple's cross-platform UI understanding multimodal large language model

Ferret-UI 2 is a multimodal, large-scale language model from Apple, used for understanding and interacting with mobile user interfaces. Ferret-UI 2 can recognize and understand UI elements on various mobile device screens, execute complex user commands,...

What is Ferret-UI 2?

Ferret-UI 2 is a multimodal large-scale language model from Apple, used to understand and interact with mobile user interfaces. Ferret-UI 2 can recognize and understand UI elements on various mobile device screens, execute complex user commands, observe user actions on mobile device screens in real time, and be ready to provide assistance and perform tasks. Ferret-UI 2 has undergone significant improvements and updates compared to earlier versions. Based on high-resolution image encoding and advanced data training methods, it improves the accuracy of UI element recognition and interaction capabilities, allowing users to interact with smart devices more naturally and efficiently.

Ferret-UI 2's main features

  • Multi-platform supportFerret-UI 2 can handle user interfaces for multiple platforms, including iPhone, Android, iPad, Webpage, and Apple TV.
  • High-resolution image perceptionBased on adaptive scaling technology, Ferret-UI 2 can achieve more accurate visual element recognition while maintaining the original UI screenshot resolution.
  • Advanced task training data generationBased on GPT-4o and set-of-mark visual cues, Ferret-UI 2 generates training data for complex tasks, improving the model's understanding of the spatial relationships between UI elements.
  • User Center InteractionFerret-UI 2 can understand and execute user-centric interaction tasks, such as confirming submissions and clicking buttons, rather than just mechanical clicks.
  • Cross-platform migration capabilitiesFerret-UI 2 demonstrates powerful cross-platform portability, enabling migration and adaptation between different platforms.

Technical Principles of Ferret-UI 2

  • Multimodal Large Language Model (MLLM)It combines visual perception and language processing capabilities to understand and generate complex interactions with the UI.
  • Adaptive N-grid mechanismThe algorithm determines the optimal grid size, encoding each part of the UI screenshot with minimal resolution distortion and pixel variation.
  • Dynamic high-resolution image codingThe CLIP image encoder is used to extract global and local features, which are then fed into a large language model (LLM).
  • Visual samplerBased on user instructions, identify and select relevant UI areas, and output a perception or interaction description of the UI elements.
  • Set-of-mark (SoM) visual cuesWhen generating training data, use SoM prompts to enhance the model's understanding of spatial relationships between UI elements, especially in multi-turn perception and interactive question answering tasks.
  • End-to-end trainingThe model learns from the original data annotations through an end-to-end training process, generating high-quality training data and optimizing model performance.

Ferret-UI 2 project address

Application scenarios of Ferret-UI 2

  • smartphones and tabletsFerret-UI 2 can understand and execute various user commands on iOS and Android devices, such as navigating applications, sending messages, and setting reminders.
  • Web browsingIn web browsing, it helps users interact with web page elements more effectively, such as clicking buttons, filling out forms, and navigation links.
  • Smart TVFor smart TV platforms such as Apple TV, provide voice control and other interaction methods to enhance the user experience.
  • Multitasking environmentIn scenarios where multiple applications or windows need to be processed simultaneously, it helps users manage and switch between different tasks more efficiently.
  • assistive technologyIt is integrated into assistive technologies to help people with disabilities interact with devices through voice commands or other input methods.