Human conversation is governed by subconscious, sub-second timing. Linguistic studies demonstrate that in natural human dialogue, the average pause between conversational turns is roughly 200 to 250 milliseconds. When a voice assistant pauses for 1.5 to 2.0 seconds before responding, the interaction immediately feels awkward, robotic, and unnatural.
The Cascade Latency Trap
Traditional voice assistants relied on a sequential three-stage pipeline (Cascade Architecture):
$$\text{User Audio} \rightarrow \text{ASR (300ms)} \rightarrow \text{Text LLM (600ms)} \rightarrow \text{TTS Synthesizer (500ms)} \rightarrow \text{Audio Out}$$When you add network transport and buffering, total latency inevitably exceeds 1,500 milliseconds, and crucial non-verbal vocal cues—intonation, laughter, hesitation, and emotional pitch—are completely lost during transcription.
[Traditional Cascade Architecture: >1500ms Turn Latency]
Audio In ──► [Whisper ASR: 300ms] ──► [Text LLM: 600ms] ──► [TTS Synthesizer: 500ms] ──► Audio Out
[Native Speech-to-Speech Architecture: <250ms Turn Latency]
Audio In ──► [Continuous Audio Tokenizer (SNAC/EnCodec)]
│
▼ (Streaming Residual Audio Tokens via WebRTC)
[Full-Duplex Multimodal Neural Core]
│
▼ (Instant Cross-Attention Audio Stream)
Audio Out ◄── [Direct Neural Audio Vocoder]
The Conceptual Shift: Native Audio Tokens & Full Duplex
Modern real-time conversational engines eliminate the text-transcription bottleneck entirely through Native Speech-to-Speech Modeling:
- Continuous Audio Tokenization: Audio is compressed into discrete neural audio codecs (such as SNAC or EnCodec) at 50Hz, converting raw waveforms into streaming tokens.
- Full-Duplex WebRTC Streaming: Audio packets stream bi-directionally over low-latency UDP WebRTC connections, allowing the model to listen and speak simultaneously.
- Natural Interruption Handling: Because the neural model receives incoming audio tokens while generating output tokens, it detects user vocal interruptions instantly, pausing output generation in under 80 milliseconds.
By treating sound as a native language token rather than an external transcription, voice AI achieves true conversational presence.