I am using livekit worker agent in python sdk. I a...
# ask-ai
f
I am using livekit worker agent in python sdk. I am noticing a lot of discrepancy in stt for user participant. Can you tell me some tips of parameter tuning for better STT experience. FYI- I am using openai whisper and i am using manual turn detection.
t
Thanks for reaching out, Tushar. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To help you get the best out of the Speech-to-Text (STT) experience using OpenAI Whisper with manual turn detection in the LiveKit Python SDK, here are some friendly tips: • Enable Manual Turn Detection: You can set
turn_detection="manual"
in the
AgentSession
constructor. Then, remember to explicitly commit the turn using
session.commit_user_turn()
once speech ends. Also, the
transcript_timeout
parameter is handy to wait for STT results. • Configure OpenAI Whisper: Using the
whisper-1
model usually gives better accuracy. If you want, you can provide a
prompt
to help guide the transcription, especially if you’re working with domain-specific vocabulary or accents. • Optimize Audio Input: It’s a good idea to disable audio input initially and enable it only during active speech, like with push-to-talk. This helps minimize background noise and improves overall performance. • Ignore VAD Parameters: Since you’re using manual turn detection, VAD or endpointing settings (like
threshold
and
silence_duration_ms
) won’t have an effect because automatic detection is turned off. These settings give you full control over the transcription process, making your STT experience more reliable and responsive. If you have any questions or need more help, feel free to reach out—I’m here for you! Sources: Turn detection and interruptions | LiveKit Docs | livekit-plugins/livekit-plugins-assemblyai/livekit/plugins/assemblyai/stt.p...
f
vad = silero.VAD.load() stt_adapter = stt.StreamAdapter( stt=openai.STT( model="whisper-1", language="en", ), vad=vad ) session = AgentSession( stt=stt_adapter, llm=llm, tts=tts, vad=vad, _turn_detection_="manual" ) can you help me identify problem with this configuration
t
Thank you for sharing your configuration! I noticed that you have both a VAD (Silero) and manual turn detection set up. Since manual turn detection actually disables automatic voice activity detection, I just wanted to check in and see if you’re planning to use the VAD for something specific, or if it might be okay to leave it out. Understanding this will really help us find any possible conflicts or unnecessary parts in your setup. I’m here to help you get everything running smoothly!
f
I honestly dont remember why I have used vad because i set it up a while back. In that case, Would you suggest to remove vad?
t
I'm glad you asked! When you're using manual turn detection, it's actually best to remove the VAD from your configuration. Since manual turn detection turns off automatic voice activity detection, having a VAD in there might not help and could even make things a bit more complicated or cause conflicts. By simplifying your setup and leaving out the VAD, you'll make sure the transcription process follows your clear control through
session.commit_user_turn()
. If you have any questions or need a hand with this change, please feel free to reach out—I’m here to help!