We have VAD false positives causing interruptions without meaningful speech content. Here's what happens:
1. VAD detects activity (breathing, background noise, etc.)
2. Agent stops speaking and waits for transcript
3. Deepgram returns empty/no transcript
4. Agent stays stuck waiting for content that never arrives
Our Goal
Replace VAD-based interruption with word-based interruption:
- Only interrupt when Deepgram sends interim transcript results with actual words
- Cost optimize by not sending all audio to Deepgram - only when VAD thinks there's speech
- Maintain low latency for responsive interruption (current VAD interruption is very fast)
Current Implementation
# We set turn detection mode to interim_transcript
turn_detection_mode = 'interim_transcript'
session_kwargs['turn_detection'] = "stt"
# We create a separate STT stream for interim processing
if turn_detection_mode == 'interim_transcript':
stt_stream = shared_stt.stream()
async def process_stt_events():
async for event in stt_stream:
if event.type == SpeechEventType.INTERIM_TRANSCRIPT:
interim_text = event.alternatives[0].text.strip()
if interim_text:
session.interrupt() # Interrupt on real words
# Push audio frames from user's track
async for audio_event in audio_stream:
stt_stream.push_frame(audio_event.frame)
The Challenge
When we tried using StreamAdapter for VAD gating:
vad_gated_stt = stt.StreamAdapter(stt=shared_stt, vad=vad_instance)
It caused significant delays - interruption became much slower than the original VAD-based system.
---
Questions for LiveKit Expert
1. Is our parallel STT stream approach correct? Should we create a separate STT stream for interim transcripts, or integrate with the session's
existing STT differently?
2. Why does StreamAdapter cause delays? VAD works fast elsewhere - what makes StreamAdapter slow for real-time interruption?
3. What's the proper VAD integration pattern for cost optimization (only send speech audio to STT) without sacrificing interrupt responsiveness?
4. Interim transcript access - what's the right way to get interim results from Deepgram in LiveKit agents? The user mentioned having "a hard time
getting interim results" - are there gotchas?
5. Architecture recommendation - for VAD-gated word-based interruption in LiveKit agents, what's the proper audio flow and component integration?