AB
AiBoss
project

Emu3 - A unified input and generative multimodal model developed by Beijing Academy of Artificial Intelligence.

Emu3 is a native multimodal world model developed by the Beijing Academy of Artificial Intelligence (BAAI). It employs BAAI's self-developed multimodal autoregressive technology, jointly trained on images, videos, and text, enabling the model to possess native multimodal capabilities...

What is Emu3?

Emu3, developed by the Beijing Academy of Artificial Intelligence (BAAI), is a native multimodal world model. Employing BAAI's proprietary multimodal autoregressive technology, it is jointly trained on images, videos, and text, enabling native multimodal capabilities and unified input and output for these media. Emu3 converts various content into discrete symbols and predicts the next symbol based on a single Transformer model, simplifying the model architecture. In image generation, Emu3 can create high-quality images that meet requirements with just a text description, outperforming the dedicated image generation model SDXL. Regarding image and language understanding, Emu3 accurately describes real-world scenes and provides appropriate textual responses without relying on CLIP or pre-trained language models. Emu3 can also extend existing video content naturally, expanding video scenes.

Emu3's main functions

  • Image generationEmu3 can generate high-quality images based on text descriptions, supporting different resolutions and styles.
  • Video generationEmu3 can generate videos by predicting the next symbol in a video sequence, without relying on complex video diffusion techniques.
  • Video predictionEmu3 can naturally continue existing video content, predict what will happen next, and simulate environments, people, and animals in the physical world.
  • Illustrated ExplanationEmu3 can understand the physical world and provide coherent textual responses without relying on CLIP or pre-trained language models.

Emu3's technical principles

  • Next token predictionThe core of Emu3 is next token prediction, which is an autoregressive method. The model is trained to predict the next element in a sequence, whether it is text, image or video.
  • Multimodal sequence unificationEmu3 unifies image, text, and video data into a discrete token space, enabling a single Transformer model to handle multiple types of data.
  • Single Transformer ModelEmu3 uses a single Transformer model trained from scratch to handle all types of data, simplifying the model architecture and improving efficiency.
  • Autoregressive generationIn the generation task, Emu3 predicts tokens in the sequence one by one through autoregression, thereby generating images or videos.
  • Illustrated ExplanationIn image-text comprehension tasks, Emu3 can encode images into tokens and then generate text describing the content of the images.

Emu3's project address

Application scenarios of Emu3

  • Content creationEmu3 automatically generates images and videos based on text descriptions, helping artists and designers quickly realize their creative ideas.
  • Advertising and MarketingGenerate compelling advertising materials based on Emu3 to enhance brand promotion effectiveness.
  • educateEmu3 visualizes complex concepts, enhancing the student learning experience.
  • Entertainment industryEmu3 assists in game and movie production, creating realistic virtual environments.
  • Design and architectureEmu3 is used to generate design prototypes and architectural renderings, improving design efficiency.
  • e-commerceEmu3 helps online retailers generate product display images, enhancing the shopping experience.