AB
AiBoss
project

Step-Audio 2 mini - Step-Star's open-source end-to-end large-scale speech model

Step-Audio 2 mini is an open-source end-to-end speech model released by Step-Audio. Breaking away from traditional speech model structures, it adopts a true end-to-end multimodal architecture, directly converting raw audio input into speech response output with lower latency...

What is Step-Audio 2 mini?

Step-Audio 2 mini is an open-source, end-to-end speech model released by Step-Audio. Breaking away from traditional speech model structures, it adopts a true end-to-end multimodal architecture, directly converting raw audio input into speech response output with lower latency and the ability to understand paralinguistic information and non-human voice signals. The model incorporates chain-like reasoning and reinforcement learning for joint optimization, enabling refined understanding and responses to emotions and intonation. It supports external tools such as web search, effectively addresses the hallucination problem, and enhances its scalability across multiple scenarios.

In terms of performance, Step-Audio 2 mini achieved state-of-the-art (SOTA) results on multiple international benchmark datasets. For example, on the general multimodal audio understanding benchmark MMAU, it ranked first among open-source end-to-end speech models with a score of 73.2; on the URO Bench, which measures spoken dialogue ability, it achieved the highest score among open-source end-to-end speech models in both the basic and professional tracks; in the Chinese-English translation task, it significantly outperformed GPT-4o Audio and other open-source speech models; and in the speech recognition task, it achieved first place in multilingual and multi-dialect tasks, leading other open-source models by more than 15%.

Main functions of Step-Audio 2 mini

  • Audio understandingIt can accurately understand various audio content, including natural sounds, music, and speech, and can also capture paralinguistic information such as emotions and intonation, thus realizing the perception of "subtext".
  • Speech recognitionIt performs exceptionally well in multilingual and multi-dialect speech recognition, boasting high accuracy and the ability to quickly convert speech into text, making it suitable for various language environments.
  • Voice translationIt supports voice-to-voice translation, enabling mutual translation between Chinese, English, and other languages, helping users overcome language barriers to communicate.
  • Emotion and Paralinguistic AnalysisIt can analyze the emotions and paralinguistic features in speech, such as anger, happiness, sadness, and other emotions, as well as non-verbal signals such as laughter and sighs, making the interaction more natural.
  • voice dialogueIt possesses excellent conversational skills, can conduct fluent voice communication, understand complex questions and provide appropriate answers, and can be used in scenarios such as intelligent customer service and voice assistants.
  • Tool callIt supports online search and other operations, and can obtain the latest information in real time, providing users with more comprehensive and accurate answers.
  • Content creationIt can assist in generating audio content, such as podcasts and audiobooks, providing creators with inspiration and materials.

The technical principles of Step-Audio 2 mini

  • True end-to-end multimodal architectureBreaking through the traditional three-level structure of speech models, it directly converts the original audio input into speech response output, simplifies the architecture, reduces latency, and can effectively understand paralinguistic information and non-human voice signals.
  • CoT inference combined with reinforcement learningThis is the first time that chain-thinking reasoning and reinforcement learning have been jointly optimized in an end-to-end speech model, enabling sophisticated understanding, reasoning, and natural responses to paralinguistic and non-speech signals such as emotion, tone, and music.
  • Audio knowledge enhancementIt supports external tools such as web retrieval to help the model solve the illusion problem, improve its scalability in multiple scenarios, and enable the model to obtain the latest information and make accurate answers.

Step-Audio 2 mini project address

  • GitHub repositoryhttps://github.com/stepfun-ai/Step-Audio2
  • Hugging Face Model Libraryhttps://huggingface.co/stepfun-ai/Step-Audio-2-mini
  • Experience addresshttps://realtime-console.stepfun.com

Application scenarios of Step-Audio 2 mini

  • Intelligent voice assistantIt provides users with convenient voice interaction services, such as smart home control and smart office assistant, allowing them to complete various operations through voice commands.
  • Intelligent Customer ServiceIn the customer service field, it enables users to quickly and accurately understand their problems and provide solutions, thereby improving service efficiency and user experience.
  • Voice translationIt enables real-time voice-to-voice translation, helping users overcome language barriers and is suitable for scenarios such as international communication and business meetings.
  • Audio content creationIt assists creators in generating audio content, such as podcasts and audiobooks, by providing creative inspiration and content generation support.
  • EducationUsed for language learning, online education, etc., it provides a personalized learning experience through voice interaction to help students improve their language skills.
  • HealthcareIt is used in fields such as medical consultation and rehabilitation therapy to provide patients with health advice and psychological support through voice dialogue.