AB
AiBoss
project

Seeduplex - ByteDance's native full-duplex voice big model

Seeduplex is a native full-duplex speech model developed by ByteDance's Seed team, enabling real-time interaction with simultaneous listening and speaking. The model boasts precise anti-interference capabilities (reducing false interruption rate by 50%) and dynamic stop detection (reducing interruption rate by 40%), even in noisy environments...

What is Seeduplex?

Seeduplex is a native full-duplex voice model developed by ByteDance's Seed team, enabling real-time "listen and speak" interaction. The model boasts precise anti-interference capabilities (reducing false interruption rate by 50%) and dynamic call termination (reducing interruption rate by 40%), delivering natural and smooth performance in noisy environments and complex scenarios involving multiple users. Seeduplex has been fully launched on the Doubao App, providing a high-quality voice call experience to hundreds of millions of users, marking the first large-scale commercial deployment of full-duplex voice technology.

Seeduplex's main functions

  • Full-duplex real-time interactionIt enables "listening and speaking simultaneously," breaking the traditional turn-based limitation of "question and answer," and supporting true real-time two-way voice communication.
  • Precise anti-interferenceIt continuously perceives the global acoustic environment and accurately locates the main user's voice in noisy scenarios such as cars and cafes, reducing the false reply rate and false interruption rate by 50%.
  • Dynamic stoppingIntelligent judgment of dialogue rhythm by combining voice and semantic features: patiently listens when the user is thinking, responds instantly after the user finishes speaking, reduces the proportion of interruptions by 40%, and reduces the pause delay by 250ms.
  • Agile interrupt responseIt can respond to user interruption commands (such as "wait a moment") in real time, reducing the interruption response latency by 300ms and achieving smooth switching.
  • Environmental perception linkageIt automatically analyzes background ambient sounds (such as broadcasts and navigation sounds) and incorporates them into the reasoning context, proactively responding by combining environmental information.
  • Complex Expression ComprehensionIt supports users' fragmented expressions that they can revise as they think (such as repeatedly adjusting their order requirements) and accurately captures their final intent.

How to use Seeduplex

  • Download/update Doubao AppUpdate the Doubao App to the latest version.
  • Enter voice callSelect the "Make a Call" icon in the dialog box to enter the voice call interface and experience it.

Key information and usage requirements for Seeduplex

  • Product Name: Seeduplex (Seed-Full-Duplex)
  • Development TeamByteDance Seed Team
  • Technology typeNative full-duplex voice large model
  • Core breakthroughIt enables real-time interaction of "listening and speaking simultaneously," breaking through the traditional turn-based limitations of "question and answer."
  • Key Indicators:
    • The rates of false interruptions and false responses have been reduced by 50%.
    • The rate of interrupting others decreased by 40%.
    • Stop delay reduced by approximately 250ms
    • Interruption response latency reduced by approximately 300ms
    • User call satisfaction improved by an absolute 8.34%.
  • Online statusIt has been fully launched on the Doubao App, making it the industry's first full-duplex voice model to achieve large-scale deployment.
  • Platform restrictionsOnly supported via Doubao App

Seeduplex's core advantages

  • Native full-duplex architectureThe industry's first "listen and speak" voice model to be deployed on a large scale breaks through the traditional turn-based limitation of "question and answer", and the interaction is as natural as a real person's conversation.
  • Precise anti-interference capabilityBy perceiving the global acoustic environment, it can accurately locate the main user's voice in noisy scenes (in cars, cafes, etc.), reducing the false reply rate and false interruption rate by 50%.
  • Intelligent dynamic stop judgmentThe system combines voice and semantic features to judge the rhythm of the conversation in real time, listens patiently when the user is thinking, and responds instantly after the user finishes speaking (reducing latency by 250ms), thus reducing the rate of interruption by 40%.
  • Ultra-low latency responseInterruption response latency is reduced by 300ms, and it supports interruption at any time, enabling truly smooth real-time two-way communication.

Comparison of Seeduplex's similar competing products

Comparison Dimensions Seeduplex
(ByteDance)
GPT-Realtime
(OpenAI)
Step-Audio
(Stepping onto the stars)
Technical Architecture End-to-end speech large model
Native full-duplex architecture
End-to-end Speech-to-Speech
Streaming Real-Time Transmission
End-to-end unified modeling
Open source full-duplex architecture
Core advantages Precise anti-interference(Incorrect interruption rate ↓50%)
Dynamic stopping(Interception rate decreased by 40%)
Ultra-low latency response
Multimodal fusion(Supports image input)
Emotion recognition (laughter/tone of voice)
Improved tool usage ecosystem
Emotional control(Dynamic switching of emotions within the sentence)
Dialect support(Cantonese, Sichuanese, etc.)
Voice-native Tool Calling
Latency performance Delay in suspension ↓250ms
Interrupt response ↓300ms
Real-time streaming, specific values not disclosed.
Supports SIP telephony protocol access
Low latency, specific optimization values not disclosed.
Anti-interference capability powerful(Accurately pinpointing human voices in noisy environments)
(False response rate reduced by 50%)
Medium (depends on end-to-end generalization ability) Medium difficulty (open-source models require user optimization for specific scenarios).
Openness Closed sourceDoubao App built-in
It is now fully online and no application is required.
API paid(Realtime API)
Support third-party integration development
open source(GitHub/HuggingFace)
Supports local deployment and customization
Scene focus Complex acoustic environments (in car/shopping mall)
High-frequency interactive game (Flying Flower Order)
Multi-person dialogue scenarios
Customer Support Agent
Educational guidance
Multimodal real-time interaction
Smart cockpit voice control
Medical consultation (supports 30 medical terms)
Customer service in dialect areas

Application scenarios of Seeduplex

  • Voice interaction in noisy environmentsIn high-noise environments such as inside a car (where navigation and radio broadcasts are mixed), in a coffee shop, or in a shopping mall, it accurately isolates background interference and locks onto the main user's voice.
  • Multi-person dialogue scenariosWhen users converse with others (such as responding to delivery drivers or friends interrupting), the system can identify commands genuinely addressed to the AI, avoiding accidental triggers. In overlapping conversations with multiple users, it can accurately distinguish which statements are directed at the AI and which are casual remarks from others.
  • Fragmented/Hesitant ExpressionSupports users to think and revise complex expressions as they go, such as repeatedly adjusting their requests when ordering ("I want it iced... no, hot... add two more pumps of syrup").
  • High-frequency interactive gamesIn scenarios requiring rapid responses, such as quick Q&A and word games, it achieves seamless, low-latency responses (reduced by approximately 250ms) and supports smooth competitive dialogues.