News
StepAudio 2.5 ASR, a new generation of automatic speech recognition model, has been launched by StepAudio.
StepAudio 2.5, a next-generation automatic speech recognition model, has been launched. This model pioneers the introduction of large language model inference acceleration technology into the field of speech recognition. Based on the ASR+MTP-5 architecture, it achieves a 400% increase in inference speed, a 60% reduction in latency, a peak speed of 500 tokens/s, and an 80% cost reduction. It achieves state-of-the-art (SOTA) performance on multiple mainstream Chinese and English benchmarks, reuses a 32K context window, and can transcribe 30 minutes of audio in a single pass.