When using gpt real-time with deepgram STT The tra...
# ask-ai
b
When using gpt real-time with deepgram STT The transcript doesn't have the user's terms. What's up with that?
t
Thanks for reaching out, Sebastian. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Thanks so much for reaching out! To help you better, could you share a bit more about your setup? Specifically: • Are you using any custom language models or keyword boosting with Deepgram STT? • Which specific terms are missing from the transcript? • Have you noticed any error messages or unexpected behaviors in the integration?
b
session = AgentSession( stt=deepgram.STT(model="nova-3", language="en-US"), llm=openai.realtime.RealtimeModel( model="gpt-realtime", turn_detection=None, input_audio_transcription=None, voice=config.voice, temperature=config.temperature ), vad=ctx.proc.userdata["vad"], turn_detection=EnglishModel(), ) The transcript is missing the user's terms, so for example it says: {'items': [{'id': 'item_CIxR87455VHbEAJIq1kNH', 'type': 'message', 'role': 'assistant', 'content': ["Hello! I'm your AI assistant running in console mode. How can I help you today?"], 'interrupted': False, 'created_at': 1758634419.8534982}, {'id': 'item_CIxRFbRTYDQXEP8HDvYpI', 'type': 'message', 'role': 'assistant', 'content': ["Hey! I'm doing great, thanks for asking. How about you? Everything going well on your end?"], 'interrupted': False, 'created_at': 1758634427.415773}, {'id': 'item_CIxRNE9ldPO2XZHq3MFkU', 'type': 'message', 'role': 'assistant', 'content': ['Glad to hear that! Anything interesting going on or something you want to chat about today?'], 'interrupted': False, 'created_at': 1758634435.291072}, {'id': 'item_CIxRYcPdaLmB5aUupj9s7', 'type': 'message', 'role': 'assistant', 'content': ["Wow, that’s incredible news! Ousmane Dembélé winning the Ballon d'Or 2025, that’s a huge achievement. He must have had an outstanding season. What do you think made him stand out this year?"], 'interrupted': False, 'created_at': 1758634452.478347}]} so there's no users turn
t
Thank you so much for sharing the code snippet and example output! To help you better, could you please clarify a few things for me? • Are the user's terms being detected at all by Deepgram STT, or do they seem to be missing entirely from the transcription output? • Have you had a chance to test the Deepgram STT integration on its own to see if it’s transcribing the user input correctly? • Is this issue happening consistently with all user inputs, or only with certain phrases or terms?
b
look at the example transcript i gaev, it' smissing the users turn entirely
t
I understand how important it is to have accurate user turns in your transcript, and I’m here to help you with that! The reason you’re seeing missing user turns is that real-time LLMs like OpenAI's GPT Realtime API provide full user transcripts only after a turn ends. However, turn detection needs live, incremental transcription data to know exactly when someone finishes speaking. To get this working smoothly, you’ll want to set up a dedicated STT plugin alongside your real-time LLM integration for turn detection. Here’s a friendly step-by-step guide to fix this: 1. Install and Set Up the Turn Detector Plugin: `pip install "livekit-agents[turn-detector]~=1.2"`Before running the agent, don’t forget to download the model weights:
python agent.py download-files
2. Configure AgentSession: Use this configuration to make sure user turns are captured properly:
Copy code
from livekit.agents import AgentSession
from livekit.plugins import deepgram, openai
from livekit.plugins.turn_detector.multilingual import MultilingualModel

session = AgentSession(
    stt=deepgram.STT(model="nova-3", language="en-US"),
    llm=openai.realtime.RealtimeModel(
        voice="alloy",
        turn_detection=None,  # Disable OpenAI’s internal turn detection
        input_audio_transcription=None,  # Prevent duplication; use Deepgram instead
    ),
    turn_detection=MultilingualModel(),  # Use LiveKit's turn detector
    vad=silero.VAD.load(),  # Recommended for improved VAD performance
)
3. Enable Filler Words: To help turn detection work even better, set
filler_words=True
in your Deepgram STT configuration. 4. Set Language Consistently: Make sure the language in your STT matches what the turn detector expects. For English-only, use
EnglishModel()
and set the STT language to
"en"
or
"en-US"
. 5. Adjust Turn Detection Delays (Optional): If you want, you can fine-tune how sensitive the turn detection is by tweaking
min_endpointing_delay
and
max_endpointing_delay
parameters. By combining a streaming STT provider like Deepgram with LiveKit’s turn detection model, and turning off any conflicting internal turn detection in the real-time LLM, you’ll get accurate user turns captured in your transcripts. If you have any questions or need a hand with any of these steps, please feel free to reach out—I’m here to support you! Sources: LiveKit turn detector plugin | LiveKit Docs | livekit-plugins/livekit-plugins-aws/livekit/plugins/aws/experimental/realti...
b
Awesome. But if you notice in the session above, we already do this.
t
Thank you so much for bringing this to my attention. Since your configuration already has the key components in place, let’s walk through some additional troubleshooting steps together to help resolve the issue of missing user turns: 1. Verify Deepgram STT Settings: Please double-check that
interim_results=True
and
filler_words=True
are explicitly enabled in your Deepgram STT configuration. These settings are really important because they provide the incremental transcription data needed for detecting turns. 2. Check Final Transcript Handling: Make sure the final transcripts you receive from Deepgram aren’t empty and have a good confidence score. If the
current_transcript
is empty or the confidence is low, the system might skip that turn. Looking at debug logs can be a great way to spot if this is happening. 3. Turn Detection Mode: If you’re using
turn_detection_mode="manual"
, please ensure that
session.commit_user_turn()
is being called explicitly to finalize user turns. For automatic mode, it’s helpful to confirm that the turn detector model is properly initialized and paired with the STT plugin. 4. Audio Buffering: If you notice delays in final transcripts, try adding some silence frames to flush out the STT buffer. This often helps trigger endpointing and makes sure the final transcript gets generated. 5. Integration with Real-Time LLMs: To avoid any conflicts with LiveKit’s turn detector, please disable any built-in turn detection in the real-time LLM (like OpenAI GPT Realtime). I see this is already set in your configuration, but it’s always good to double-check for any overlapping settings. If you try these steps and still run into issues, please don’t hesitate to reach out. I’m here to help, and we can definitely explore more debugging options or escalate this for a deeper look if needed. Sources: livekit-agents/livekit/agents/voice/audio_recognition.py | Gemini Live API integration guide | LiveKit Docs | LiveKit turn detector plugin | LiveKit Docs
b
Is this a known limitation? Because everything you've mentioned is fine there's no issues it's not those configurations
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!