bulky-actor-93512
09/23/2025, 2:49 AMtall-airline-16756
09/23/2025, 2:49 AMtall-airline-16756
09/23/2025, 2:50 AMpip install "livekit-agents[deepgram]~=1.2"
2. Initialize the STT Service: Next, configure the Deepgram STT plugin by specifying parameters like the model (nova-3), language (en-US), and enabling interim results so you can get real-time transcription feedback:
from livekit.plugins import deepgram
stt = deepgram.STT(model="nova-3", language="en-US", interim_results=True)
3. Subscribe to Audio Tracks: Use the LiveKit SDK to subscribe to user audio tracks as they join a room, so you can capture their audio in real-time.
4. Stream Transcriptions to GPT: Now, with the stt.stream() method, you can process the audio streams and send the transcriptions to GPT to generate real-time responses. Here’s a simple example:
stt_stream = stt.stream()
async for event in stt_stream:
if event.type == SpeechEventType.FINAL_TRANSCRIPT:
user_text = event.alternatives[0].text
# Send user_text to GPT
elif event.type == SpeechEventType.INTERIM_TRANSCRIPT:
pass
5. Integrate GPT: You can use GPT models like GPT-3.5 or GPT-4o to process the transcribed text and create thoughtful responses.
6. Optimize Latency: To make everything feel smooth and responsive, enable streaming recognition and use smaller frame sizes (like 20ms) for faster transcription and quicker responses.
If you’d like to dive deeper or see some practical examples, I highly recommend checking out the Deepgram STT Integration Guide and the Live Translated Captioning Example. They’re fantastic resources to support you along the way.
Sources: README.md | Deepgram STT integration guide | LiveKit Docs | Live Conversations with AI using ChatGPT and WebRTC | Live Conversations with AI using ChatGPT and WebRTC