AB
AiBoss
project

UniToken - A unified visual coding framework launched by Fudan University in collaboration with Meituan and other institutions.

UniToken is a novel autoregressive generative model designed for multimodal understanding and generation tasks. By combining discrete and continuous visual representations, it constructs a unified visual encoding framework capable of simultaneously capturing the high-level semantics of images...

What is UniToken?

UniToken is a novel autoregressive generative model designed specifically for multimodal understanding and generation tasks. By combining discrete and continuous visual representations, it constructs a unified visual encoding framework that can simultaneously capture both high-level semantics and low-level details of images. This allows UniToken to seamlessly support visual understanding and image generation tasks, providing multi-dimensional information for different tasks.

UniToken's main functions

  • Illustrated ExplanationUniToken can efficiently handle image and text understanding tasks, such as image captioning and visual question answering (VQA).
  • Image generationUniToken supports high-quality image generation tasks, including generating images from text descriptions, image editing, and story generation.
  • Multimodal dialogueIn multimodal dialogue scenarios, UniToken can generate natural language responses based on input text and image information, supporting more complex interactive tasks, such as interpreting image content or generating new images based on image and text instructions.
  • Complex instruction followingUniToken enhances fine-tuning through instructions, enabling it to better understand and execute complex multimodal instructions, such as generating images with specific layouts given text descriptions and images.
  • Fine-grained vision tasksWith the help of technologies such as AnyRes and ViT end-to-end fine-tuning, UniToken can process high-resolution images, improve the ability to perceive image details, and is suitable for tasks that require high-precision visual processing.
  • Task versatilityUniToken can seamlessly integrate multimodal understanding and generation tasks, supporting a variety of complex tasks such as text and image understanding, image generation, image editing, and story generation, demonstrating powerful general generation capabilities.

UniToken's technical principles

  • Unified visual codingUniToken employs a dual encoder consisting of continuous and discrete encoders, combining the discrete encoding of VQ-GAN with the continuous representation of SigLIP to generate visual encodings that possess both high-level semantics and low-level details, providing complete visual information for multimodal large models.
  • Multi-stage training
    • Visual semantic space alignmentBased on Chameleon, the language model (LLM) is frozen, and only SigLIP ViT and Adapter are trained to align continuous visual encoding with the language space.
    • Multi-task joint trainingJoint training on large-scale image understanding and image generation datasets, and by controlling the data ratio, the performance of the model on both understanding and generation tasks can be improved in a balanced way.
    • Instructions for Enhanced Fine-tuningIntroducing high-quality multimodal dialogue and refined image generation data further enhances the model's ability to follow complex instructions.
  • Fine-grained visual enhancementUniToken supports technologies such as AnyRes and ViT end-to-end fine-tuning to improve the fine-grained perception of high-resolution images while avoiding model crashes and adapting to a wide range of task scenarios.

UniToken's project address

UniToken Application Scenarios

  • Content creation and designUniToken can generate high-quality images based on text descriptions, helping designers quickly create creative sketches or concept diagrams, saving design time and effort.
  • Intelligent customer service and virtual assistantIn multimodal dialogue scenarios, UniToken can understand the text and image information input by the user and generate natural language responses.
  • Education and LearningUniTokens can be used in education to help students better understand and learn complex concepts. For example, by generating images related to scientific experiments, historical events, or literary works, UniTokens can enhance students' visual memory and comprehension.
  • Medical and HealthIn the medical field, UniTokens can be used to generate medical images or interpret medical images.
  • Autonomous driving and traffic managementUniToken can be used for visual question answering (VQA) tasks in autonomous driving scenarios. For example, vehicles can upload road images in real time and use UniToken to generate natural language descriptions of road conditions, traffic signs, and other information to help the autonomous driving system make more accurate decisions.