chilly-nest-87883
07/27/2025, 9:44 PMtall-airline-16756
07/27/2025, 9:45 PMtall-airline-16756
07/27/2025, 9:46 PMAgentSession API, you can utilize the built-in hooks provided by the framework. Here's how you can achieve this:
1. Enable Interim Results:
Ensure that interim results are enabled in your Deepgram STT configuration. This is already set by default with interim_results=True in your code. session = AgentSession(
stt=deepgram.STT(model="nova-3", interim_results=True),
llm=openai.LLM(),
tts=cartesia.TTS(),
vad=silero.VAD.load(),
turn_detection=EnglishModel()
)
2. Implement Event Handlers:
Use the on_interim_transcript and on_final_transcript hooks to process interim and final transcripts. For example: async def on_interim_transcript(self, ev: stt.SpeechEvent):
interim_text = ev.alternatives[0].text
print(f"Interim: {interim_text}")
# Add your custom interruption logic here
async def on_final_transcript(self, ev: stt.SpeechEvent):
final_text = ev.alternatives[0].text
print(f"Final: {final_text}")
3. Access Current Transcript:
The AgentSession also provides a current_transcript property that includes both interim and final transcripts. You can use this for real-time processing.
4. Complete Example:
Here’s how you can integrate the hooks into your session: class MyAgent:
def __init__(self):
self.session = AgentSession(
stt=deepgram.STT(model="nova-3", interim_results=True),
llm=openai.LLM(),
tts=cartesia.TTS(),
vad=silero.VAD.load(),
turn_detection=EnglishModel()
)
self.session.on_interim_transcript = self.on_interim_transcript
self.session.on_final_transcript = self.on_final_transcript
async def on_interim_transcript(self, ev: stt.SpeechEvent):
interim_text = ev.alternatives[0].text
print(f"Interim: {interim_text}")
async def on_final_transcript(self, ev: stt.SpeechEvent):
final_text = ev.alternatives[0].text
print(f"Final: {final_text}")
async def start(self, agent, room):
await self.session.start(agent=agent, room=room)
This approach allows you to intercept interim transcripts as they arrive and implement your custom interruption logic effectively.
Sources: Deepgram STT integration guide | LiveKit Docs | livekit-plugins/livekit-plugins-baseten/livekit/plugins/baseten/stt.py | Working with plugins | LiveKit Docschilly-nest-87883
07/27/2025, 10:43 PMtall-airline-16756
07/27/2025, 10:44 PMsession.on_interim_transcript challenges? No worries, I've got your back! Let's break down some common reasons why your interim transcripts might be playing hide and seek: 1. Realtime Model Integration: If you're using a realtime LLM model with user_transcription turned on, sometimes the interim transcript events can get a bit shy. It's like the realtime model is stealing the spotlight! Here's a friendly fix: - Try adding a separate STT plugin just for transcription, while letting the realtime model handle responses. 2. Manual Turn Detection Mode: When you're in manual turn detection mode, interim transcripts might feel a bit left out if the user turn is already wrapped up. Just double-check your turn detection settings to make sure everything's playing nice. 3. STT Configuration Issues: Let's make sure your Speech-to-Text setup is ready to rock. Do a quick check to confirm: - interim_results=True is switched on in your STT config. - Your STT service is connected and listening for those audio vibes. 4. Audio Input Verification: Think of audio frames like guests at a party - they need to actually show up for things to get interesting! Make sure your audio is being sent through the recognition pipeline. Still scratching your head? No problem! If you can share some details about your LLM setup, turn detection mode, and any logs you've got, we'll solve this together. You've got this! 💪 Sources: Realtime model integrations | LiveKit Docs | livekit-plugins/livekit-plugins-deepgram/livekit/plugins/deepgram/stt.py