News
Clear to hear, understandable to see! Doubao Speech Recognition Model 2.0 is here!
Volcano Engine has released Doubao Speech Recognition Model 2.0. Based on the Seed hybrid expert architecture, the model achieves deep contextual reasoning through PPO reinforcement learning, improving keyword recall by 20%. It adds multimodal visual recognition capabilities, accurately distinguishing easily confused words (such as "滑鸡" and "滑稽") by combining image content, and supports accurate recognition of 13 languages including Japanese, Korean, and German.