chilly-nest-87883
07/28/2025, 5:37 PMtall-airline-16756
07/28/2025, 5:37 PMtall-airline-16756
07/28/2025, 5:38 PMEnglishModel() for turn detection while simultaneously running a separate STT stream for custom interruption logic. These two audio processing pipelines are independent.
The turn detection model operates in the text domain, so it requires an STT plugin to function. However, this does not prevent you from running additional STT streams for other purposes. The audio processing architecture supports parallel processing, where audio frames are distributed to both VAD and STT models simultaneously. This ensures that each component can process the same audio stream independently without interference.
Here’s an example setup:
from livekit.plugins.turn_detector.english import EnglishModel
from livekit.plugins import deepgram
session = AgentSession(
turn_detection=EnglishModel(), # Uses STT for turn detection
stt=deepgram.STT(), # Primary STT stream
# You can also initialize additional STT instances for custom logic
)
This configuration allows you to implement both turn detection and custom interruption logic effectively.
Sources: livekit-agents/livekit/agents/voice/audio_recognition.py | Agents Core Architecture | livekit/livekit_composite | DeepWiki | Turn detection and interruptions | LiveKit Docschilly-nest-87883
07/28/2025, 8:10 PMtall-airline-16756
07/28/2025, 8:12 PMSTART_OF_SPEECH and END_OF_SPEECH.
2. Only Push Frames to STT When VAD Indicates Speech
Here's where the magic happens! You'll keep a little audio buffer and only send frames to STT when there's actual speech. Check out this handy example:
from livekit.plugins import silero, deepgram
from livekit.agents.vad import VADEventType
class CustomSTTWithVAD:
def __init__(self):
self.custom_vad = silero.VAD.load(
min_speech_duration=0.2,
min_silence_duration=0.6,
)
self.vad_stream = self.custom_vad.stream()
self.stt = deepgram.STT()
self.audio_buffer = []
self.is_speaking = False
async def process_audio_frame(self, frame):
self.vad_stream.push_frame(frame)
self.audio_buffer.append(frame)
if len(self.audio_buffer) > 1000:
self.audio_buffer.pop(0)
async def process_vad_events(self):
async for vad_event in self.vad_stream:
if vad_event.type == VADEventType.START_OF_SPEECH:
self.is_speaking = True
if len(self.audio_buffer) > 100:
self.audio_buffer = self.audio_buffer[-100:]
elif vad_event.type == VADEventType.END_OF_SPEECH:
self.is_speaking = False
if self.audio_buffer:
await self.send_to_deepgram(self.audio_buffer.copy())
self.audio_buffer.clear()
async def send_to_deepgram(self, speech_frames):
stt_stream = self.stt.stream()
for frame in speech_frames:
stt_stream.push_frame(frame)
stt_stream.end_input()
async for event in stt_stream:
if event.type == SpeechEventType.FINAL_TRANSCRIPT:
print(f"Transcribed: {event.alternatives[0].text}")
3. Access VAD State/Events in a Custom Loop
This is where you get to play audio detective! Use the VAD stream to catch those speech events, so you know exactly when to send your audio frames to STT.
Alternative: Use StreamAdapter for Automatic Integration
Not into writing all that code? No worries! There's a super simple shortcut with the built-in StreamAdapter that does all the heavy lifting for you:
from livekit import agents
from livekit.plugins import deepgram, silero
vad = silero.VAD.load(
min_speech_duration=0.1,
min_silence_duration=0.5,
)
stt_with_vad = agents.stt.StreamAdapter(deepgram.STT(), vad)
stt_stream = stt_with_vad.stream()
async for audio_frame in audio_source:
stt_stream.push_frame(audio_frame)
And just like that, you've got a smart system that only sends actual speech to Deepgram. Less noise, lower costs, and a smoother experience. Win-win! 🎉
Sources: basics/echo_transcriber_agent.py | Working with plugins | LiveKit Docs | livekit-agents/livekit/agents/stt/stream_adapter.py