AB
AiBoss
project

Indic Parler-TTS - An open-source multilingual TTS model focused on synthesizing Hindi and English.

Indic Parler-TTS is a multilingual text-to-speech (TTS) model developed in collaboration between Hugging Face and AI4Bharat, specifically designed for speech synthesis in Indian languages and English. Indic Parler-TTS is Parl...

What is Indic Parler-TTS?

Indic Parler-TTS is a multilingual text-to-speech (TTS) model developed in collaboration between Hugging Face and AI4Bharat, specifically designed for speech synthesis in Indian languages and English. An extended version of Parler-TTS Mini, Indic Parler-TTS supports 20 Indian languages and English, offering 69 unique speech characteristics and generating natural, clear, and expressive speech output. Based on descriptive text input, the model flexibly adjusts speech characteristics such as pitch, rate, emotion, and background noise to adapt to various application scenarios. Indic Parler-TTS performs exceptionally well in multiple Indian languages and demonstrates strong adaptability in low-resource languages.

Main functions of Indic Parler-TTS

  • Multilingual support:
    • It supports 20 Indian languages and English, including Hindi, Tamil, Bengali, Telugu, Marathi, and others.
    • Limited support is provided for languages that are not officially supported, such as Kashmiri and Punjabi.
  • Rich emotional and vocal characteristics:
    • It supports a variety of emotional expressions, such as anger, happiness, sadness, and surprise.
    • It supports adjusting the tone, speed, background noise, reverb, and overall sound quality of the voice.
  • Flexible input methods:
    • Users control speech characteristics using descriptive text (captions), such as specifying the speaker's gender, accent, emotion, and recording environment.
    • The model automatically identifies the language of the input text and switches to the corresponding language for speech synthesis.
  • High-quality voice output: Performs well in multiple languages, especially Indian languages.
  • speech diversityIt offers 69 unique voices, with a recommended voice for each language to ensure natural and clear pronunciation.
  • Customization capabilitiesUsers can precisely control the background noise, reverberation, expressiveness, pitch, speech rate, and speech quality of the voice based on descriptive text.

Indic Parler-TTS Technical Principles

  • Deep learning-based TTS architectureA deep learning-based text-to-speech model, employing an Encoder-Decoder architecture, converts text input into speech waveforms to achieve high-quality speech synthesis.
  • Multilingual pre-training and fine-tuningIt is pre-trained on a large-scale multilingual dataset and then fine-tuned on specific Indian and English datasets. This pre-training + fine-tuning approach enables it to adapt to multiple languages and dialects.
  • Descriptive text control: Introduce descriptive text (caption) input, and control speech based on the characteristics of natural language description.
  • Double segmenter mechanismThe model uses two tokenizers: one for processing text input (prompt) and the other for processing descriptive text (description).

Indic Parler-TTS project address

Application scenarios of Indic Parler-TTS

  • voice assistantIt provides multilingual voice interaction for smart devices, making them easier for users to operate.
  • audiobooksIt converts text into speech to meet the reading needs of different users.
  • News BroadcastGenerate multilingual audio content to expand the reach of information dissemination.
  • Customer service systemIt supports automatic voice responses in multiple languages, improving service efficiency.
  • Content creationIt provides efficient speech synthesis for film, television, advertising, and other fields, enriching creative forms.