chilly-nest-87883
07/28/2025, 11:44 PMtall-airline-16756
07/28/2025, 11:44 PMtall-airline-16756
07/28/2025, 11:45 PMStreamAdapter in LiveKit is designed to enable non-streaming Speech-to-Text (STT) providers to function in streaming contexts by combining them with Voice Activity Detection (VAD). Here's what happens when you initialize it:
• The StreamAdapter takes an STT instance (e.g., my_stt) and a VAD instance (e.g., my_vad) as parameters.
• It sets its capabilities to streaming: true and interimResults: false, making it suitable for streaming use cases.
• It creates a wrapper that processes audio in two stages:
1. Audio Forwarding: Incoming audio frames are forwarded to the VAD stream instead of directly to the STT. The VAD analyzes the audio for speech activity.
2. VAD-Driven Recognition: The adapter listens for VAD events:
◦ On START_OF_SPEECH, it emits a start-of-speech event.
◦ On END_OF_SPEECH, it:
▪︎ Buffers the audio frames detected during speech.
▪︎ Sends these frames to the underlying STT for recognition.
▪︎ Emits the final transcript if text is detected.
This approach allows the StreamAdapter to buffer audio until the VAD signals the end of speech, ensuring that complete speech segments are sent to the STT for better recognition accuracy. It effectively bridges the gap between non-streaming STT models and streaming applications.
Sources: agents/src/pipeline/human_input.ts | livekit-agents/livekit/agents/stt/fallback_adapter.py | Working with plugins | LiveKit Docs | Speech-to-text (STT) integrations | LiveKit Docs