Hi! I'm working on implementing word-based interru...
# ask-ai
c
Hi! I'm working on implementing word-based interruption in a LiveKit agents voice system to replace VAD-only interruption (which causes false positives from breathing/noise). Goal: Only interrupt the agent when Deepgram sends interim transcript results with actual words, while optimizing costs by not sending all audio to Deepgram continuously. Current Setup: - Using turn_detection_mode = 'interim_transcript' - Session configured with session_kwargs['turn_detection'] = "stt" - Deepgram STT with interim_results=True - Creating separate STT stream for interim processing: if turn_detection_mode == 'interim_transcript': stt_stream = shared_stt.stream() async def process_stt_events(): async for event in stt_stream: if event.type == SpeechEventType.INTERIM_TRANSCRIPT: interim_text = event.alternatives[0].text.strip() if interim_text: session.interrupt() # Interrupt on real words # Push audio frames from user's audio track async for audio_event in audio_stream: stt_stream.push_frame(audio_event.frame) What I Tried for Cost Optimization: 1. StreamAdapter Approach: vad_gated_stt = stt.StreamAdapter(stt=shared_stt, vad=vad_instance) stt_stream = vad_gated_stt.stream() 1. Result: Caused significant delays in interruption because StreamAdapter buffers audio until END_OF_SPEECH, so no interim transcripts arrive during speech. 2. Frame-level VAD filtering: if vad_instance.classify(audio_event.frame).speech_probability > threshold: stt_stream.push_frame(audio_event.frame) 2. Result: Failed because VAD doesn't have a classify() method - it works with streams/events, not individual frame classification. Current Status: - Basic interim transcript processing works (interrupts on words) - But sends ALL audio to Deepgram continuously (expensive) - Need proper VAD integration that gates STT without buffering delays Questions: 1. What's the correct LiveKit pattern for VAD-gated STT streaming that maintains real-time interim transcripts? 2. How should VAD integration work with a separate STT stream for interim processing? 3. Is there a way to have VAD act as an instant "gate" (not buffer) that immediately starts/stops sending frames to STT when speech is detected/ends?
t
Thanks for reaching out, Clovis. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To address your goal of implementing word-based interruption while optimizing costs, here are the key insights and recommendations: 1. Correct Pattern for VAD-Gated STT Streaming: • Use the
StreamAdapter
pattern if you want VAD to act as a gate for STT processing. This ensures that STT only processes audio frames when VAD detects speech, avoiding unnecessary processing of silence. • However, note that the
StreamAdapter
buffers audio until VAD detects the end of speech, which means it does not support real-time interim transcripts. 2. VAD Integration with Separate STT Stream: • For real-time interim transcripts, you can run VAD and STT streams in parallel on the same audio input. VAD provides timing signals (e.g.,
START_OF_SPEECH
and
END_OF_SPEECH
), while STT processes audio continuously for interim results. • Example implementation:
Copy code
const sttStream = this.#stt.stream();
const vadStream = this.#vad.stream();

const audioStreamCo = async () => {
  for await (const ev of audioStream) {
    sttStream.pushFrame(ev);  // Send to STT for interim processing
    vadStream.pushFrame(ev);  // Send to VAD for speech detection
  }
};

const vadStreamCo = async () => {
  for await (const ev of vadStream) {
    switch (ev.type) {
      case VADEventType.START_OF_SPEECH:
        this.emit(HumanInputEvent.START_OF_SPEECH, ev);
        break;
      case VADEventType.END_OF_SPEECH:
        this.emit(HumanInputEvent.END_OF_SPEECH, ev);
        break;
    }
  }
};

const sttStreamCo = async () => {
  for await (const ev of sttStream) {
    if (ev.type === SpeechEventType.INTERIM_TRANSCRIPT) {
      this.emit(HumanInputEvent.INTERIM_TRANSCRIPT, ev);
    } else if (ev.type === SpeechEventType.FINAL_TRANSCRIPT) {
      this.emit(HumanInputEvent.FINAL_TRANSCRIPT, ev);
    }
  }
};
3. VAD as an Instant Gate: • VAD can act as an instant gate to start/stop sending frames to STT when speech is detected/ends. This is implemented in the
StreamAdapter
pattern, where VAD controls when STT processes audio. • Key mechanism:
Copy code
const recognize = async () => {
  for await (const ev of this.#vadStream) {
    switch (ev.type) {
      case VADEventType.START_OF_SPEECH:
        this.output.put({ type: SpeechEventType.START_OF_SPEECH });
        break;
      case VADEventType.END_OF_SPEECH:
        this.output.put({ type: SpeechEventType.END_OF_SPEECH });

        try {
          const event = await this.#stt.recognize(ev.frames);
          if (!event.alternatives![0].text) {
            continue;
          }
          this.output.put(event);
          break;
        } catch (error) {
          // error handling
        }
    }
  }
};
Recommendations: • If real-time interim transcripts are critical, use the parallel processing approach with VAD and STT streams running simultaneously. • If cost optimization is the priority and you can work with final transcripts, use the
StreamAdapter
pattern to gate STT processing based on VAD signals. Let me know if you need further clarification or additional details! Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/stt.py | agents/src/pipeline/human_input.ts | Speech-to-text (STT) integrations | LiveKit Docs | Realtime model integrations | LiveKit Docs | agents/src/vad.ts