AB
AiBoss
project

StepAudio 2.5 TTS - A context-aware speech generation model developed by StepAudio.

StepAudio 2.5 TTS is a Contextual TTS (Context-Aware Speech Synthesis Model) launched by StepAudio, which for the first time introduces contextual understanding capabilities into the entire speech generation process.

What is StepAudio 2.5 TTS?

StepAudio 2.5 TTS is a Contextual TTS (Context-Aware Speech Synthesis Model) launched by StepAudio, which for the first time introduces contextual understanding capabilities into the entire speech generation process. The model uses a dual-level control system, which sets the tone and rhythm of each paragraph through Global Context and finely controls emotional pauses sentence by sentence through Inline Context. Combined with Zero-shot voice replication, it only requires 3 seconds of reference audio to replace traditional labels with natural language descriptions, allowing AI to upgrade from "reading text" to "performing text".

Main functions of StepAudio 2.5 TTS

  • Global context controlIt supports describing the emotional tone, character state, and scene atmosphere of the entire audio in natural language (such as "restrained sadness, no sobbing, slight trembling"), making the expression more consistent and coherent.
  • Context control in the textUse parentheses in text () Insert instructions within sentences to finely control details such as emotion, tone, rhythm, pauses, breathing, and stress changes. The content in parentheses is only an instruction and will not be read aloud.
  • Zero-shot tone reproductionThe target timbre can be cloned with just 3 seconds of reference audio, and the replicated timbre fully inherits the global and contextual control capabilities, without being limited by a fixed sound library.
  • Non-streaming speech synthesis:pass POST /v1/audio/speech The interface synthesizes a complete audio file in one go, prioritizing sound quality, and is suitable for scenarios where latency is not a major concern.
  • Streaming speech synthesis:pass WebSocket /v1/realtime/audio It enables low-latency streaming, making it suitable for dialogue and real-time playback scenarios.
  • Reissue Preview:pass /v1/audio/voices/preview The interface allows for quick previews of the synthesized reference audio, charging only for the synthesis process and not creating official sound assets.
  • Full timbre context controlBoth the replicated and original timbres can be flexibly adjusted in terms of emotion, style, and expression through natural language commands to achieve a performance effect of "same sound, different feeling".

How to use StepAudio 2.5 TTS

  • Obtain accessVisit the Stepfun Open Platform (https://platform.stepfun.com/docs/zh/guides/models/stepaudio-2.5-tts) to register an account and obtain the API Key from the console.
  • Select access method:
    • Online experienceTo try it out, visit the Experience Center at https://www.stepfun.com/studio/audio or the Demo page at https://stepaudiollm.github.io/step-audio-2.5-tts/.
    • API callsChoose between a non-streaming (quality-first) or a streaming (low-latency) interface depending on the scenario.
  • Write context instructions:
    • set up instruction(Global context)Describe the overall tone of the passage in natural language, such as "the voice was extremely tense, the speech was fast and intermittent, with a noticeable sense of suppression."
    • edit input Text (Context)Insert parentheses in sentences that require precise control. () Mark the mood and pauses, such as, "(lowering his voice) Hey... look at my phone. (short inhale)"
  • Call API
    • non-flow:Towards https://api.stepfun.com/v1/audio/speech Send a POST request, carrying the parameters model, voice, input, and instruction.
    • Flow cytometry: Connecting via WebSocket wss://api.stepfun.com/v1/realtime/audioSend first tts.create Establish a session, and then through tts.text.delta Push text stream with parenthetical instructions
  • Sound reproduction (optional)To clone a sound, prepare a reference audio file of at least 3 seconds for the target sound, and then call... /v1/audio/voices/preview Listen to the audio sample and, once confirmed, create the official sound asset.

Key information and usage requirements of StepAudio 2.5 TTS

  • Model Basics
    • The model type is Contextual TTS (Context-Aware Speech Synthesis), which achieves voice performance based on natural language understanding and supports dual control of global context (overall tone) and textual context (intra-sentence details).
    • The maximum input per session is 1000 characters, and the maximum instruction (global contextual natural language guidance) is 200 characters.
  • Pricing Standards
    • Context-based text-to-speech: 5.8 yuan/10,000 characters
    • Voice replication/generation: 9.9 yuan/voice (the listening interface only charges for synthesis; payment is made immediately upon successful replication).
  • Access method
    • Non-streaming speech synthesis: POST /v1/audio/speech, synthesizes a complete audio file in one go.
    • Streaming speech synthesis: WebSocket /v1/realtime/audio, low-latency streaming response suitable for dialogue scenarios.
    • Remastered preview: POST /v1/audio/voices/preview, quick preview without creating official sound assets.
  • Usage restrictions
    • Parentheses are used to control the context in the text. () The enclosed instruction, the content within the parentheses, is treated as an instruction only and will not be read aloud directly.
    • Zero-shot tone replication requires only 3 seconds of reference audio, and the replicated tone fully inherits the ability to control the nuances of the sound.
    • The Step Plan and Step Plan open platforms are now fully operational, allowing users to directly call the API or experience the service online.

The core advantages of StepAudio 2.5 TTS

  • Natural Language Alternation Tagging SystemIt abandons traditional fixed labels such as "sadness/anger" and supports precise tone description using complex natural language such as "restrained sadness, no crying voice, slight trembling", which greatly reduces the threshold for regulation.
  • Dual-level contextual precision controlGlobal Context controls the overall emotional tone and character state, while Inline Context... () The parentheses are used to fine-tune the rhythm, pauses, and breathing of each sentence, achieving a three-dimensional sound directing from macro to micro perspectives.
  • Zero-shot fully controllable replicaIt only requires 3 seconds of reference audio to clone any timbre, and the replicated timbre fully inherits the context control capability, breaking through the limitations of fixed sound libraries. The same voice can express a variety of emotional styles.
  • Performance-level vocal qualityThe system features a comprehensive upgrade in rhythmic dimensions such as pauses, emphasis, and intonation transitions, as well as an upgrade in underlying vocal quality. It bids farewell to the "plastic" and "AI-like" feel of traditional TTS, achieving a lifelike performance effect where "every word is expressive."
  • Low barrier to entry and high flexibilityNo professional audio knowledge is required; you can control complex emotional expressions simply by "speaking out your needs." It also supports both non-streaming (high-quality audio) and streaming (low-latency) modes, adapting to everything from content creation to real-time dialogue.

StepAudio 2.5 TTS Competitive Product Comparison

Dimension StepAudio 2.5 TTS ElevenLabs Fish Audio
Pricing Standards 5.8 yuan per 10,000 characters (approximately $0.08 per thousand characters) Flash: ~$0.06/thousand characters; Multilingual v2: ~$0.12-0.18/thousand characters (approximately 0.87-1.3 yuan/thousand characters)

~$15/million characters (approximately $0.015/thousand characters, 0.11 yuan/thousand characters)

Free quota Please check the official website for specific policies. 10,000 characters/month (Free plan)

500 characters/time, 7 minutes S1 generation per month

Sound reproduction Zero-shot, 3-second audio, 9.9 yuan/sound, supports full context control. Instant Clone (pay-to-use) + Professional Voice Clone (high fidelity, starting with the Creator plan)

Supports voice cloning, available from Plus plan onwards.

Context control Dual-mode control: Global Context + Inline Context (inline parentheses instructions) Based on SSML tags and speed/style control, the v3 model supports sentiment expression.

Adjusting basic parameters (speed, emotion, etc.)
Latency performance Supports non-streaming (quality-first) and WebSocket streaming (low latency). Flash v2.5: ~75ms; Turbo v2.5: ~250-300ms

Standard generation speed (Free), Enhanced generation speed (Plus+)

Language support Primarily optimized for Chinese, but supports multiple languages. 29+ languages, deep multilingual optimization

Multilingual support
Input restrictions A maximum of 1000 characters can be entered at a time, with a maximum instruction size of 200 characters. Maximum 10,000 characters per transaction (API)

Free: 500 characters/time; Plus: 15,000 characters/time; Pro: 30,000 characters/time

Core advantages Natural language description with label substitution, performance-level emotional control, and precise dual-level context regulation. The voice's naturalness is industry-leading (9.5/10), with rich emotional expression and a well-developed ecosystem.

Lowest price, available open-source models, high cost-performance ratio

Applicable Scenarios Film and television dubbing, audiobooks, game characters, and Chinese content creation. Audiobooks, podcasts, international multilingual content, real-time conversational AI Large-scale programmatic generation, budget-sensitive projects, developers

Application scenarios of StepAudio 2.5 TTS

  • Film and animation dubbingBy setting the emotional tone of the character in the overall context and finely adjusting the tone and pauses in the text, professional-level character dubbing is achieved, making the character's voice more layered and realistic.
  • Audiobook and podcast productionBy leveraging dual-level context control capabilities, unique voice personalities can be given to different characters, creating immersive multi-person audio content and lowering the barrier to professional audio production.
  • Game voice generation: Build complete voice character profiles for game characters, achieving comprehensive customization from voiceprints to personality, and enabling NPCs to have vivid expressions that match the scene atmosphere.
  • Intelligent voice assistantLeveraging the low latency of streaming speech synthesis, it empowers intelligent customer service and AI assistants with natural conversational capabilities, supporting real-time context adjustment to match user emotions.
  • Advertising and Marketing Content: Quickly clone brand-specific sounds using Zero-shot sound replication, and combine this with contextual control to generate marketing audio materials with a consistent style and rich emotions.