I'm using LiveKit Agents Python framework with Age...
# ask-ai
c
I'm using LiveKit Agents Python framework with AgentSession and I want to access interim transcripts from Deepgram STT for custom interruption logic. You mentioned that Deepgram emits INTERIM_TRANSCRIPT and FINAL_TRANSCRIPT events, but how do I actually listen to these events when using the high-level AgentSession API? My current code structure is: session = AgentSession( stt=deepgram.STT(model="nova-3", interim_results=True), llm=openai.LLM(), tts=cartesia.TTS(), vad=silero.VAD.load(), turn_detection=EnglishModel() ) await session.start(agent=agent, room=room) I need to intercept interim transcripts as they arrive (before final transcription) to implement word-based interruption logic. What's the correct way to access the STT stream events or interim transcript data within the AgentSession framework? Should I be using session event handlers, or do I need to access the STT instance directly?
t
Thanks for reaching out, Clovis. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To access interim transcripts from Deepgram STT while using the high-level
AgentSession
API, you can utilize the built-in hooks provided by the framework. Here's how you can achieve this: 1. Enable Interim Results: Ensure that interim results are enabled in your Deepgram STT configuration. This is already set by default with
interim_results=True
in your code.
Copy code
session = AgentSession(
    stt=deepgram.STT(model="nova-3", interim_results=True),
    llm=openai.LLM(),
    tts=cartesia.TTS(),
    vad=silero.VAD.load(),
    turn_detection=EnglishModel()
)
2. Implement Event Handlers: Use the
on_interim_transcript
and
on_final_transcript
hooks to process interim and final transcripts. For example:
Copy code
async def on_interim_transcript(self, ev: stt.SpeechEvent):
    interim_text = ev.alternatives[0].text
    print(f"Interim: {interim_text}")
    # Add your custom interruption logic here

async def on_final_transcript(self, ev: stt.SpeechEvent):
    final_text = ev.alternatives[0].text
    print(f"Final: {final_text}")
3. Access Current Transcript: The
AgentSession
also provides a
current_transcript
property that includes both interim and final transcripts. You can use this for real-time processing. 4. Complete Example: Here’s how you can integrate the hooks into your session:
Copy code
class MyAgent:
    def __init__(self):
        self.session = AgentSession(
            stt=deepgram.STT(model="nova-3", interim_results=True),
            llm=openai.LLM(),
            tts=cartesia.TTS(),
            vad=silero.VAD.load(),
            turn_detection=EnglishModel()
        )
        self.session.on_interim_transcript = self.on_interim_transcript
        self.session.on_final_transcript = self.on_final_transcript

    async def on_interim_transcript(self, ev: stt.SpeechEvent):
        interim_text = ev.alternatives[0].text
        print(f"Interim: {interim_text}")

    async def on_final_transcript(self, ev: stt.SpeechEvent):
        final_text = ev.alternatives[0].text
        print(f"Final: {final_text}")

    async def start(self, agent, room):
        await self.session.start(agent=agent, room=room)
This approach allows you to intercept interim transcripts as they arrive and implement your custom interruption logic effectively. Sources: Deepgram STT integration guide | LiveKit Docs | livekit-plugins/livekit-plugins-baseten/livekit/plugins/baseten/stt.py | Working with plugins | LiveKit Docs
c
I have registered interim transcript handlers using session.on_interim_transcript = my_handler but the handler never fires when users speak. The logs show that Deepgram STT is configured with interim_results=True and smart_format=True, and I can see final transcript events in conversation_item_added, but my interim/final transcript handlers never receive events. What's the correct way to register and receive interim transcript events in LiveKit Agents? Do I need to use a different event registration method or additional configuration?
t
Hey there! 👋 Dealing with some tricky
session.on_interim_transcript
challenges? No worries, I've got your back! Let's break down some common reasons why your interim transcripts might be playing hide and seek: 1. Realtime Model Integration: If you're using a realtime LLM model with
user_transcription
turned on, sometimes the interim transcript events can get a bit shy. It's like the realtime model is stealing the spotlight! Here's a friendly fix: - Try adding a separate STT plugin just for transcription, while letting the realtime model handle responses. 2. Manual Turn Detection Mode: When you're in manual turn detection mode, interim transcripts might feel a bit left out if the user turn is already wrapped up. Just double-check your turn detection settings to make sure everything's playing nice. 3. STT Configuration Issues: Let's make sure your Speech-to-Text setup is ready to rock. Do a quick check to confirm: -
interim_results=True
is switched on in your STT config. - Your STT service is connected and listening for those audio vibes. 4. Audio Input Verification: Think of audio frames like guests at a party - they need to actually show up for things to get interesting! Make sure your audio is being sent through the recognition pipeline. Still scratching your head? No problem! If you can share some details about your LLM setup, turn detection mode, and any logs you've got, we'll solve this together. You've got this! 💪 Sources: Realtime model integrations | LiveKit Docs | livekit-plugins/livekit-plugins-deepgram/livekit/plugins/deepgram/stt.py