Can I use EnglishModel() for turn detection (conve...
# ask-ai
c
Can I use EnglishModel() for turn detection (conversation flow) while simultaneously running a separate STT stream from a user's audio track for custom interruption logic? Are these two audio processing pipelines independent?
t
Thanks for reaching out, Clovis. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Yes, you can use
EnglishModel()
for turn detection while simultaneously running a separate STT stream for custom interruption logic. These two audio processing pipelines are independent. The turn detection model operates in the text domain, so it requires an STT plugin to function. However, this does not prevent you from running additional STT streams for other purposes. The audio processing architecture supports parallel processing, where audio frames are distributed to both VAD and STT models simultaneously. This ensures that each component can process the same audio stream independently without interference. Here’s an example setup:
Copy code
from livekit.plugins.turn_detector.english import EnglishModel
from livekit.plugins import deepgram

session = AgentSession(
    turn_detection=EnglishModel(),  # Uses STT for turn detection
    stt=deepgram.STT(),            # Primary STT stream
    # You can also initialize additional STT instances for custom logic
)
This configuration allows you to implement both turn detection and custom interruption logic effectively. Sources: livekit-agents/livekit/agents/voice/audio_recognition.py | Agents Core Architecture | livekit/livekit_composite | DeepWiki | Turn detection and interruptions | LiveKit Docs
c
I want to add VAD gating to my custom STT stream to reduce Deepgram costs. I have a Silero VAD instance and am processing audio frames like this: async for audio_event in audio_stream: stt_stream.push_frame(audio_event.frame) # Currently sends ALL frames How can I: 1. Check if my VAD instance detects speech for each audio frame? 2. Only push frames to STT when VAD indicates speech is active? 3. Access VAD state/events from a custom audio processing loop? I want: if vad_detects_speech: stt_stream.push_frame(frame) to reduce costs by only sending speech segments to Deepgram.
t
Hey there! Let me walk you through a super cool way to save some money on your speech-to-text processing while keeping things smooth and efficient. We're going to use Voice Activity Detection (VAD) to only send actual speech to your STT service. Trust me, it's easier than it sounds! 1. Check if VAD Detects Speech for Each Frame Think of this like having a smart audio filter. You'll send each audio frame through the VAD, which will basically tell you, "Yep, someone's talking!" or "Nope, it's just silence" using events like
START_OF_SPEECH
and
END_OF_SPEECH
. 2. Only Push Frames to STT When VAD Indicates Speech Here's where the magic happens! You'll keep a little audio buffer and only send frames to STT when there's actual speech. Check out this handy example:
Copy code
from livekit.plugins import silero, deepgram
from livekit.agents.vad import VADEventType

class CustomSTTWithVAD:
    def __init__(self):
        self.custom_vad = silero.VAD.load(
            min_speech_duration=0.2,
            min_silence_duration=0.6,
        )
        self.vad_stream = self.custom_vad.stream()
        self.stt = deepgram.STT()
        self.audio_buffer = []
        self.is_speaking = False

    async def process_audio_frame(self, frame):
        self.vad_stream.push_frame(frame)
        self.audio_buffer.append(frame)
        if len(self.audio_buffer) > 1000:
            self.audio_buffer.pop(0)

    async def process_vad_events(self):
        async for vad_event in self.vad_stream:
            if vad_event.type == VADEventType.START_OF_SPEECH:
                self.is_speaking = True
                if len(self.audio_buffer) > 100:
                    self.audio_buffer = self.audio_buffer[-100:]

            elif vad_event.type == VADEventType.END_OF_SPEECH:
                self.is_speaking = False
                if self.audio_buffer:
                    await self.send_to_deepgram(self.audio_buffer.copy())
                    self.audio_buffer.clear()

    async def send_to_deepgram(self, speech_frames):
        stt_stream = self.stt.stream()
        for frame in speech_frames:
            stt_stream.push_frame(frame)
        stt_stream.end_input()
        async for event in stt_stream:
            if event.type == SpeechEventType.FINAL_TRANSCRIPT:
                print(f"Transcribed: {event.alternatives[0].text}")
3. Access VAD State/Events in a Custom Loop This is where you get to play audio detective! Use the VAD stream to catch those speech events, so you know exactly when to send your audio frames to STT. Alternative: Use StreamAdapter for Automatic Integration Not into writing all that code? No worries! There's a super simple shortcut with the built-in
StreamAdapter
that does all the heavy lifting for you:
Copy code
from livekit import agents
from livekit.plugins import deepgram, silero

vad = silero.VAD.load(
    min_speech_duration=0.1,
    min_silence_duration=0.5,
)

stt_with_vad = agents.stt.StreamAdapter(deepgram.STT(), vad)
stt_stream = stt_with_vad.stream()

async for audio_frame in audio_source:
    stt_stream.push_frame(audio_frame)
And just like that, you've got a smart system that only sends actual speech to Deepgram. Less noise, lower costs, and a smoother experience. Win-win! 🎉 Sources: basics/echo_transcriber_agent.py | Working with plugins | LiveKit Docs | livekit-agents/livekit/agents/stt/stream_adapter.py