Luna-TTS - A large-scale text-to-speech model launched by Yusheng Yueban
Luna-TTS is a large-scale text-to-speech model independently developed by VUI Labs, a startup company affiliated with Shanghai Jiao Tong University. The model abandons the mainstream autoregressive approach and adopts a Masked Diffusion architecture, based on Qwen3-0.6...
What is Luna-TTS?
Luna-TTS is a large-scale text-to-speech model independently developed by VUI Labs, a startup company affiliated with Shanghai Jiao Tong University. The model abandons the mainstream autoregressive approach, employing a Masked Diffusion architecture, trained on Qwen3-0.6B, and coupled with its self-developed Luna-Codec, pre-trained on 1 million hours of speech data from Chinese, English, Japanese, and Korean. The model topped the Hugging Face TTS Arena V2 blind listening test leaderboard globally and ranked third on the authoritative Artificial Analysis leaderboard, surpassing Google.
Main functions of Luna-TTS
- Multilingual speech synthesisIt supports high-quality speech synthesis in four languages: Chinese, English, Japanese, and Korean. The voice is natural and the rhythm is smooth, making it suitable for long text scenarios such as audiobooks and podcasts.
- Emotion and Paralinguistic ControlIt can precisely control emotions such as neutral, sad, and surprised, as well as paralinguistic details such as breathing, pauses, laughter, and sighs.
- Real-time streaming outputLuna-TTS Realtime supports block-level streaming generation with a first block latency of only 41.6ms and an end-to-end RTF as low as 0.024.
- Tone consistency and cloningSupports timbre cloning and preservation, ensuring stable timbre without drift during long text generation.
The technical principle of Luna-TTS
-
Masked Diffusion ArchitectureAbandoning the mainstream autoregressive approach, the speech RVQ token grid is treated as a whole, and the masked position is predicted in parallel through multiple rounds of iteration. The generation order is dynamically determined by the model, avoiding the accumulation of errors along the prefix.
-
Self-developed Luna-CodecThe 24kHz waveform is encoded into a discrete token grid of 25Hz with 8 codebooks at a bit rate of only 2.2kbps. A causal codec is used to achieve streaming and frame-synchronized output.
-
Qwen3 Main Model: Based on Qwen3-0.6B, we continue training. The standard version uses bidirectional attention, while the Realtime version uses block-level causal attention and KV Cache to achieve streaming generation.
-
Large-scale training and reinforcement learningThe program underwent approximately 1 million hours of pre-training and 100,000 hours of high-quality annealing in four languages: Chinese, English, Japanese, and Korean. It then transferred GRPO reinforcement learning to the diffusion denoising trajectory to optimize content accuracy and speaker consistency.
Luna-TTS project address
- Project official website:https://vuilabs-ai.github.io/luna-tts/
- arXiv technical paper:https://arxiv.org/pdf/2608.11593
Comparison of Luna-TTS with similar products
| Comparison Dimensions | Luna-TTS | ElevenLabs Eleven v3 |
|---|---|---|
| Development Team | Usei Yueban VUI Labs (China) | ElevenLabs (USA) |
| technical route | Masked Diffusion | Autoregressive generation |
| Model parameters | 0.6B | Not disclosed |
| Language support | China, UK, Japan, and South Korea | Multilingual (32+) |
| Emotional control | Fine-grained control (neutral/sad/surprised, etc.) | Support for mood and style regulation |
| Paralinguistic details | Breathing, pauses, laughter, sighs | Support for some secondary languages |
| Real-time version | Luna-TTS Realtime(RTF 0.024) | Turbo v2.5 (Low Latency) |
| API pricing | $80.0 / 1M chars | $100.0 / 1M chars |
| TTS Arena V2 | 1st place(Rating 1574) | 10th place (Rating 1176) |
| Artificial Analysis | 3rd place(Elo 1220) | 10th place (Elo 1176) |
Application scenarios of Luna-TTS
- Intelligent customer service and telemarketingIt replaces traditional voice robots, providing a natural, emotionally expressive conversational experience, supports real-time streaming responses, and reduces user hang-up rates.
- Audio content and podcastsGenerates high-quality long-text audio for audiobooks, news broadcasts, and knowledge podcasts, maintaining consistent tone and narrative rhythm, and supporting switching between multiple emotional segments.
- Online Education and TrainingIt provides clear and standard multilingual pronunciation for course instruction and language learning, and supports mood modulation to match teaching content (such as encouraging, serious, and gentle).
- In-vehicle and smart hardwareLeveraging the low latency of Luna-TTS Realtime, it provides an instant and natural voice interaction interface for in-vehicle navigation, smart homes, and headphone assistants.
- Real-time simultaneous interpretation and meeting assistantBy combining speech understanding capabilities, it enables low-latency, highly natural speech translation output in cross-language conference scenarios, thereby improving communication efficiency.