News
Alibaba open-sources Fun-ASR-Realtime, a large-scale real-time speech recognition model.
Alibaba officially launched its upgraded real-time speech recognition model, Fun-ASR-Realtime, with first-word latency controlled at the level of hundreds of milliseconds and recognition accuracy approaching that of offline models. It supports 16 dialects and 30 languages. The model possesses contextual understanding capabilities and can self-correct; in tests recognizing 16 dialects, its character accuracy averaged 88.62%, outperforming Volcano and Tencent's related products.