chilly-nest-87883
07/29/2025, 3:07 AMtall-airline-16756
07/29/2025, 3:07 AMtall-airline-16756
07/29/2025, 3:09 AMStreamAdapter pattern if you want VAD to act as a gate for STT processing. This ensures that STT only processes audio frames when VAD detects speech, avoiding unnecessary processing of silence.
• However, note that the StreamAdapter buffers audio until VAD detects the end of speech, which means it does not support real-time interim transcripts.
2. VAD Integration with Separate STT Stream:
• For real-time interim transcripts, you can run VAD and STT streams in parallel on the same audio input. VAD provides timing signals (e.g., START_OF_SPEECH and END_OF_SPEECH), while STT processes audio continuously for interim results.
• Example implementation: const sttStream = this.#stt.stream();
const vadStream = this.#vad.stream();
const audioStreamCo = async () => {
for await (const ev of audioStream) {
sttStream.pushFrame(ev); // Send to STT for interim processing
vadStream.pushFrame(ev); // Send to VAD for speech detection
}
};
const vadStreamCo = async () => {
for await (const ev of vadStream) {
switch (ev.type) {
case VADEventType.START_OF_SPEECH:
this.emit(HumanInputEvent.START_OF_SPEECH, ev);
break;
case VADEventType.END_OF_SPEECH:
this.emit(HumanInputEvent.END_OF_SPEECH, ev);
break;
}
}
};
const sttStreamCo = async () => {
for await (const ev of sttStream) {
if (ev.type === SpeechEventType.INTERIM_TRANSCRIPT) {
this.emit(HumanInputEvent.INTERIM_TRANSCRIPT, ev);
} else if (ev.type === SpeechEventType.FINAL_TRANSCRIPT) {
this.emit(HumanInputEvent.FINAL_TRANSCRIPT, ev);
}
}
};
3. VAD as an Instant Gate:
• VAD can act as an instant gate to start/stop sending frames to STT when speech is detected/ends. This is implemented in the StreamAdapter pattern, where VAD controls when STT processes audio.
• Key mechanism: const recognize = async () => {
for await (const ev of this.#vadStream) {
switch (ev.type) {
case VADEventType.START_OF_SPEECH:
this.output.put({ type: SpeechEventType.START_OF_SPEECH });
break;
case VADEventType.END_OF_SPEECH:
this.output.put({ type: SpeechEventType.END_OF_SPEECH });
try {
const event = await this.#stt.recognize(ev.frames);
if (!event.alternatives![0].text) {
continue;
}
this.output.put(event);
break;
} catch (error) {
// error handling
}
}
}
};
Recommendations:
• If real-time interim transcripts are critical, use the parallel processing approach with VAD and STT streams running simultaneously.
• If cost optimization is the priority and you can work with final transcripts, use the StreamAdapter pattern to gate STT processing based on VAD signals.
Let me know if you need further clarification or additional details!
Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/stt.py | agents/src/pipeline/human_input.ts | Speech-to-text (STT) integrations | LiveKit Docs | Realtime model integrations | LiveKit Docs | agents/src/vad.ts