We have VAD false positives causing interruptions ...
# ask-ai
c
We have VAD false positives causing interruptions without meaningful speech content. Here's what happens: 1. VAD detects activity (breathing, background noise, etc.) 2. Agent stops speaking and waits for transcript 3. Deepgram returns empty/no transcript 4. Agent stays stuck waiting for content that never arrives Our Goal Replace VAD-based interruption with word-based interruption: - Only interrupt when Deepgram sends interim transcript results with actual words - Cost optimize by not sending all audio to Deepgram - only when VAD thinks there's speech - Maintain low latency for responsive interruption (current VAD interruption is very fast) Current Implementation # We set turn detection mode to interim_transcript turn_detection_mode = 'interim_transcript' session_kwargs['turn_detection'] = "stt" # We create a separate STT stream for interim processing if turn_detection_mode == 'interim_transcript': stt_stream = shared_stt.stream() async def process_stt_events(): async for event in stt_stream: if event.type == SpeechEventType.INTERIM_TRANSCRIPT: interim_text = event.alternatives[0].text.strip() if interim_text: session.interrupt() # Interrupt on real words # Push audio frames from user's track async for audio_event in audio_stream: stt_stream.push_frame(audio_event.frame) The Challenge When we tried using StreamAdapter for VAD gating: vad_gated_stt = stt.StreamAdapter(stt=shared_stt, vad=vad_instance) It caused significant delays - interruption became much slower than the original VAD-based system. --- Questions for LiveKit Expert 1. Is our parallel STT stream approach correct? Should we create a separate STT stream for interim transcripts, or integrate with the session's existing STT differently? 2. Why does StreamAdapter cause delays? VAD works fast elsewhere - what makes StreamAdapter slow for real-time interruption? 3. What's the proper VAD integration pattern for cost optimization (only send speech audio to STT) without sacrificing interrupt responsiveness? 4. Interim transcript access - what's the right way to get interim results from Deepgram in LiveKit agents? The user mentioned having "a hard time getting interim results" - are there gotchas? 5. Architecture recommendation - for VAD-gated word-based interruption in LiveKit agents, what's the proper audio flow and component integration?
t
Thanks for reaching out, Clovis. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!
c
In LiveKit agents, what's the correct way to create an STT stream that gives us interim transcript events? # Is this the right approach? stt_stream = my_stt_instance.stream() async for event in stt_stream: if event.type == SpeechEventType.INTERIM_TRANSCRIPT: # Handle interim text