AB
AiBoss
project

EVI 3 - A speech and language model from Hume AI

EVI 3 is a brand-new speech and language model from Hume AI. The model can process both text and speech tags simultaneously, enabling natural and expressive voice interaction. It supports high personalization, generating any voice and... based on user prompts.

What is EVI 3?

EVI 3 is a brand-new speech and language model from Hume AI. The model can process both text and speech tags simultaneously, enabling natural and expressive voice interaction. It supports high personalization, generating any voice and personality based on user prompts, and adjusting emotion and speaking style in real time. In comparative tests with models such as OpenAI's GPT-4o, EVI 3 outperforms in emotion understanding, expressiveness, naturalness, and response speed. EVI 3 also boasts low-latency response capabilities, generating voice responses within 300 milliseconds.

Main functions of EVI 3

  • Multimodal interactionEVI 3 supports simultaneous processing of text and voice input, generating natural and expressive voice and language responses, achieving seamless integration of voice and text.
  • Highly personalizedUsers can create any voice and personality based on prompts, and EVI 3 will generate the corresponding voice and style in real time based on the prompts, supporting more than 100,000 custom voices.
  • Emotion and style adjustmentEVI 3 supports real-time adjustment of emotions and speaking style based on user commands, supporting a variety of emotions from "excitement" to "sadness," as well as unique speaking styles such as "pirate" or "whisper."
  • Real-time interactionEVI 3 supports generating voice and language responses within dialogue latency.

EVI 3 technical principles

  • Autoregressive modelBased on a single autoregressive model, it simultaneously processes text (T) and speech (V) tokens. The model can process text and speech inputs uniformly to generate natural and fluent speech output.
  • System promptThe system prompts include text and voice tags, provide language commands, shape the assistant's speaking style, and generate different voices and styles based on different prompts.
  • reinforcement learningBased on reinforcement learning methods, it identifies and optimizes preferred features of any human voice to achieve highly personalized voice generation.
  • StreamingEVI 3 uses streaming technology to generate voice responses within dialogue latency, ensuring smooth real-time interaction.

EVI 3 project address

Application scenarios of EVI 3

  • Intelligent Customer ServiceProvides customers with natural and fluent voice interaction to quickly answer questions.
  • voice assistantIt can be integrated into the device to provide personalized voice services.
  • Educational guidanceSimulated dialogues to aid language learning and improve social skills.
  • Emotional supportRespond to emotions and provide psychological comfort.
  • Content creationGenerates voice content with specific emotions and styles for use in audiobooks, etc.