AB
AiBoss
project

Lyra - SmartMore, in collaboration with several universities, launched an enhanced multimodal interaction capability.

Lyra is a high-performance, multimodal, large-scale language model (MLLM) developed by the Chinese University of Hong Kong, SmartMore, and the Hong Kong University of Science and Technology. It focuses on improving the interactive capabilities of speech, vision, and language modalities. Lyra is based on open-source large-scale models and multimodal...

What is Lyra?

Lyra, developed by the Chinese University of Hong Kong, SmartMore, and the Hong Kong University of Science and Technology, is a high-performance multimodal large-scale language model (MLLM) focused on enhancing the interactive capabilities of speech, vision, and language modalities. Based on open-source large-scale models, the multimodal LoRA module, and potential multimodal regularizers, Lyra reduces training costs and data requirements. It constructs large-scale multimodal datasets, including long speech samples, to handle complex long speech inputs, achieving powerful full-modal cognitive capabilities. Lyra achieves state-of-the-art performance in various modal understanding and reasoning tasks while being more efficient in terms of computational resources and training data usage.

Lyra's main functions

  • Multimodal understanding and reasoningLyra can understand and process data in multiple modalities, including images, videos, audio, and text, and perform complex understanding and reasoning tasks.
  • Voice center capabilitiesThe model is particularly enhanced in understanding speech, including the recognition and processing of long speech, and performs excellently in voice interaction.
  • High-efficiency processingLyra is more efficient in training and inference, using less data and computational resources, making it suitable for real-time and long-context multimodal applications.
  • Streaming generationIt supports the simultaneous generation of text and voice output, providing real-time responses during dialogues and interactions.
  • Cross-modal interactionBased on potential multimodal regularizers and extractors, information interaction between different modalities is enhanced, thereby improving model performance.

Lyra's technical principles

  • Multimodal LoRA (Low-Rank Adaptation)Based on LoRA technology, the model adapts to multimodal inputs, retaining its original visual capabilities while developing its capabilities in the speech modality, thus reducing the need for training data.
  • Potential cross-modal regularizerBased on the Dynamic Time Warping (DTW) algorithm, the voice token is aligned with the corresponding text token, so that the input of the voice modality is semantically consistent with the text modality.
  • Potential multimodal extractorBased on evaluating the relevance of different modal tokens to text queries, the system dynamically selects and retains the tokens most relevant to the task, thereby improving the efficiency of training and inference.
  • Long speech capability integrationWe constructed a dedicated long speech SFT dataset and used compression techniques to process long speech tokens, enabling the model to handle audio inputs up to several hours long.
  • Streaming text-to-speech generationIt integrates a streaming generation mechanism, enabling models to output corresponding speech while generating text, thus achieving a seamless multimodal interactive experience.
  • Dataset ConstructionTo train and optimize Lyra, researchers built a high-quality dataset containing more than 1.5 million multimodal samples and more than 12,000 long speech samples, covering a wide range of scenarios and domains.

Lyra's project address

Lyra's application scenarios

  • Smart AssistantAs a smart assistant, it understands and responds to users' voice commands, providing services such as information retrieval, schedule management, and reminder settings.
  • Customer ServiceIn the field of customer service, customer inquiries, complaints, and technical support are handled based on voice and text interaction.
  • Education and trainingAs an educational aid, it provides audio explanations, course content comprehension and Q&A, as well as pronunciation and listening training in language learning.
  • Health and Medical CareIn the medical field, it helps patients consult about health issues via voice, or serves as an auxiliary tool for doctors to understand and summarize patients' medical records.
  • Content moderationAnalyze image, video, and text content to conduct content moderation, identify, and filter inappropriate content.