AB
AiBoss
project

Ola - A multimodal language model jointly developed by Tsinghua University, Tencent, and others.

Ola is a multimodal language model developed collaboratively by Tsinghua University, Tencent Hunyuan Research Team, and the National University of Singapore's S-Lab. Through a progressive modality alignment strategy, it gradually expands the modalities supported by the language model, from images and text...

What is Ola?

Ola is a multimodal language model developed collaboratively by Tsinghua University, Tencent Hunyuan Research Team, and the National University of Singapore's S-Lab. Through a progressive modality alignment strategy, it gradually expands the modalities supported by the language model, starting with images and text, and then introducing speech and video data to achieve understanding of multiple modalities. Ola's architecture supports multimodal input, including text, images, video, and audio, and can process these inputs simultaneously. Ola employs a sentence-by-sentence decoding scheme for streaming speech generation, enhancing the interactive experience.

Ola's main functions

  • Multimodal understandingIt supports four modalities of input: text, image, video, and audio, and can process these inputs simultaneously, performing excellently in comprehension tasks.
  • Real-time streaming decodingIt supports user-friendly real-time streaming decoding, which can be used for text and speech generation, providing a smooth interactive experience.
  • Progressive modal alignmentBy progressively expanding the modalities supported by the language model, starting with images and text, and then introducing speech and video data, we can achieve understanding of multiple modalities.
  • High performanceIt demonstrates superior performance in multimodal benchmarks, outperforming existing open-source full-modal LLMs, and is comparable to specialized unimodal models on certain tasks.

Ola's technical principles

  • Progressive modal alignment strategyOla's training process begins with the most basic modalities (images and text), gradually introducing speech data (connecting language and audio knowledge) and video data (connecting all modalities). This progressive learning approach allows the model to gradually expand its modal understanding capabilities, keeping the scale of cross-modal aligned data relatively small, and reducing the difficulty and cost of developing a full-modal model from existing vision-language models.
  • Multimodal input and real-time streaming decodingOla supports full-modal input, including text, images, video, and audio, and can process these inputs simultaneously. Ola has designed a sentence-by-sentence decoding scheme for streaming speech generation, supporting a user-friendly real-time interactive experience.
  • Efficient utilization of cross-modal dataTo better capture the relationships between modalities, Ola's training data includes traditional visual and audio data, as well as cross-modal video-audio data. This data builds bridges between visual and audio information in the video, helping the model learn the intrinsic connections between modalities.
  • High-performance architecture designOla's architecture supports efficient multimodal processing, including visual encoders, audio encoders, text decoders, and speech decoders. Through techniques such as Local-Global Attention Pooling, the model can better fuse features from different modalities.

Ola's project address

Ola application scenarios

  • Intelligent voice interactionOla functions as a smart voice assistant, supporting speech recognition and generation in multiple languages. Users can interact with Ola using voice commands to obtain information, solve problems, or complete tasks.
  • Education and LearningOla can be used as an English practice tool to help users practice spoken English and correct pronunciation and grammar errors. It can also provide encyclopedic Q&A, covering multiple learning scenarios from K-12 to the workplace.
  • Travel and NavigationOla can act as a travel guide, providing users with introductions to the history and culture of scenic spots, as well as recommending travel guides and restaurants.
  • Emotional companionshipOla offers emotional support services to help users relieve stress and provide psychological support.
  • Life servicesOla can recommend nearby restaurants, provide schedules, and offer navigation services.