AB
AiBoss
project

Doubao 1.5·UI-TARS - GUI Agent Model launched by ByteDance Doubao

Doubao 1.5·UI-TARS is an agent model launched by ByteDance for graphical user interface (GUI) interaction. Based on human-like capabilities such as perception, reasoning, and action execution, the model enables continuous and smooth interaction with the graphical interface. The model will...

What is Doubao 1.5·UI-TARS?

Doubao 1.5·UI-TARS is an agent model for graphical user interface (GUI) interaction launched by ByteDance's Doubao platform. Based on human-like capabilities such as perception, reasoning, and action execution, the model enables continuous and smooth interaction with the GUI. It integrates visual understanding, logical reasoning, interface element location, and manipulation into a single model, achieving end-to-end task automation without the need for predefined workflows or manual rules. Doubao 1.5·UI-TARS is now available on the Volcano Ark platform.

Main functions of Doubao 1.5·UI-TARS

  • Graphical interface interaction capabilitiesBased on perception, reasoning, and action execution, it enables continuous and smooth interaction with the graphical user interface to complete complex tasks.
  • Visual understanding and positioningIt understands visual information on the screen, supports bounding box and point localization of multiple targets and small targets, performs localization counting, and describes localization content.
  • Logical reasoning and decision makingIt combines visual information and task instructions to perform logical reasoning and generate reasonable operation steps.
  • High execution efficiencyBased on the Ark Bean Bun large model inference service, it boasts the highest throughput across the entire network, with an initial 5 million TPM and extremely low inference latency of 30ms TPOT.
  • Native GUI AgentIt enables end-to-end automated execution of GUI interactive tasks without the need for predefined processes or manual rules.

The technical principles of Doubao 1.5·UI-TARS

  • Visual Large Model (VLM)The model is based on a powerful visual big model that understands and processes visual information in the graphical interface, including images, text, icons, etc.
  • Multimodal fusionIt integrates visual perception, logical reasoning, and action execution capabilities into a single model to achieve the fusion processing of multimodal information.
  • End-to-end learningBased on a large amount of labeled data and reinforcement learning, the model learns an end-to-end mapping from task input to operation output without the need for manually defined rules.

Doubao 1.5·UI-TARS Project Official Website

Application Scenarios of Doubao 1.5·UI-TARS

  • Automated officeAutomatically process tasks such as documents, spreadsheets, and emails to improve efficiency.
  • Software testingSimulate user operations to detect software problems and improve quality.
  • Intelligent Customer ServiceProvides real-time answers to user questions and operational guidance.
  • Robot InteractionIt guides robots to complete complex operations and is used in industry and logistics.