AB
AiBoss
project

Doubao Large Model 1.5 - The latest large model released by ByteDance.

Doubao Big Model 1.5 is the latest version of the big model released by ByteDance. It adopts a large-scale sparse MoE architecture, equivalent to the performance of a Dense model with 7 times the activation parameters, and achieves a high overall score across multiple evaluation criteria including knowledge, code, reasoning, and Chinese language performance...

What is the 1.5 version of the large bean bun model?

Doubao Big Model 1.5 is the latest version of the big model released by ByteDance. It adopts a large-scale sparse MoE architecture, equivalent to the performance of a Dense model with 7 times the activation parameters. Its overall score outperforms models such as GPT-4o and Claude 3.5 Sonnet across multiple benchmarks including knowledge, code, reasoning, and Chinese. Doubao Big Model 1.5 also introduces the Doubao Real-Time Voice Model (Doubao-1.5-realtime-voice-pro) and the Doubao Visual Understanding Model (Doubao-1.5-vision-pro), featuring low-latency, interruptible voice dialogue capabilities and enhanced visual reasoning and document recognition abilities. No data generated by any other models was used during model training.

Main functions of the large bean bun model 1.5

  • Comprehensive capabilities significantly enhancedIt performs globally leading on multiple authoritative benchmarks, including knowledge (such as MMLU_PRO, GPQA), code (such as McEval, FullStackBench), reasoning (such as DROP), and Chinese (such as CMMLU, C-Eval), with an overall score superior to industry-leading models such as GPT-4o, Claude 3.5 Sonnet, etc.
  • Efficient model structure and low costEmploying a large-scale sparse MoE architecture, it achieves performance equivalent to a Dense model with 7 times the activation parameters, far exceeding industry-standard efficiency. The self-developed server cluster solution supports low-cost chips, significantly reducing hardware costs.
  • Multimodal capabilities have been comprehensively improved.
    • Doubao Visual Understanding Model (Doubao-1.5-vision-pro)The system features comprehensive upgrades in multimodal data synthesis, dynamic resolution, multimodal alignment, and hybrid training, significantly enhancing its capabilities in visual reasoning, text document recognition, and fine-grained information understanding.
    • Doubao Real-Time Voice Model (Doubao-1.5-realtime-voice-pro)It adopts the Speech2Speech end-to-end framework, supports end-to-end voice dialogue, and has the characteristics of low latency and interruptibility at any time. It has been fully launched on Doubao App.
  • Deep thinking abilityBased on the Doubao 1.5 base model, through breakthroughs in RL algorithms and engineering optimization, a deep thinking model, Doubao-1.5-Pro-AS1-Preview, was developed, which has achieved leading performance in evaluations such as AIME.
  • Data independenceNo data generated by any other model was used during model training, and a completely autonomous data production system was built to ensure the independence and reliability of data sources.

Technical principles of the large bean bun model 1.5

  • Large-scale sparse MoE architectureDoubao's large model 1.5 adopts a large-scale sparse MoE (Mixture of Experts) architecture, which is pre-trained with smaller activation parameters. It is equivalent to the performance of a Dense model with 7 times the activation parameters, far exceeding the industry's conventional leverage efficiency of 3 times.
  • Multimodal fusion technologyThe model has been significantly upgraded in terms of multimodal capabilities, supporting input and output of multiple modalities such as text, image, and speech.
  • High-efficiency data processing and trainingThe Doubao Big Model 1.5 did not use any data generated by other models during training. It utilizes a self-built data production system, combined with an annotation team and model self-play technology, to ensure the independence and reliability of data sources. The model significantly reduced hardware costs through a self-developed server cluster solution and optimization techniques.
  • Reinforcement Learning and Optimization FrameworkThe Doubao Big Model team proposed the HybridFlow framework, a flexible and efficient reinforcement learning (RL) training framework that combines the advantages of single and multi-controller approaches, significantly improving training throughput.
  • Model optimization and inference accelerationDoubao Big Model 1.5 optimizes the inference efficiency of the model through techniques such as fine quantization and PD separation.

How to use the large bean bun model 1.5

  • Doubao APPThe Doubao large model 1.5 has been launched in a gray-scale release, and users can experience it in the Doubao APP.
  • Volcano Engine APIDevelopers can directly call the API through the Volcano Engine, supporting applications in multiple scenarios.
  • Price advantageThe price of the existing model remains unchanged; more features, no price increase.

Project address for Doubao Large Model 1.5

Application scenarios of the large bean bun model 1.5

  • Sentiment Analysis and FeedbackBy analyzing the sentiment of voice and text, we can better understand users' emotions and provide more targeted services.
  • Intelligent homework tutoringIt helps students solve problems in subjects such as mathematics and science, providing problem-solving strategies and steps.
  • Text generationSupports long text generation, suitable for news reporting, copywriting, story creation, etc.
  • Video generationDoubao's video generation model can generate high-quality videos based on text or images, and supports the creation of dynamic posters and short videos.
  • Visual understandingDoubao's visual understanding model can identify objects and scenes in images and perform logical reasoning, making it suitable for tasks such as question analysis and chart analysis in the education field.
  • Multilingual learningIt supports multilingual speech recognition and generation, and can be used for language learning and teaching.