I'm using OpenAI LLM Deepgram STT for streaming a ...
# ask-ai
b
I'm using OpenAI LLM Deepgram STT for streaming a custom Baseten TTS (the plugin was not working great) I do some regular flushes in the TTS Despite the pseudo streaming we have implemented, the call tests still doesn't feel very natural, the turns are off, and taking a little while to respond. How can we make the turns feel natural? If we already have deepgram streaming, should the TTS model be generating tokens immediately so when the user ends we already have a few chars to flush? How to improve this? How to tweak with jumpiness in livekit ?
t
Thanks for reaching out, Andrew. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To improve the naturalness and turn-taking in your LiveKit calls, here are some strategies and configurations you can implement: • Optimize Deepgram STT for Turn Detection: `from livekit.plugins.deepgram import stt deepgram_stt = stt.STT( model="nova-3", interim_results=True, punctuate=True, filler_words=True, # Improves turn detection smart_format=True, endpointing_ms=25, # Silence duration to consider end of speech language="en-US", )`Enable filler words and punctuation to enhance turn detection. Adjust the
endpointing_ms
parameter to fine-tune the silence duration for detecting the end of speech. • Implement Advanced Turn Detection: Use a custom turn detector to monitor end-of-utterance (EOU) probabilities in real-time: `from livekit.plugins.turn_detector.multilingual import MultilingualModel class CustomTurnDetector: def __init__(self): self._model = MultilingualModel() async def predict_end_of_turn(self, chat_ctx): probability = await self._model.predict_end_of_turn(chat_ctx) print(f"EOU Probability: {probability * 100:.3f}%") return probability`This allows you to dynamically adjust turn-taking behavior based on confidence levels. • Configure Endpointing Delays: Adjust delays for high and low confidence turn detection: `session = AgentSession( turn_detection="stt", min_endpointing_delay=0.5, max_endpointing_delay=6.0 )`This ensures smoother transitions between turns. • Enhance TTS Naturalness: Use a TTS model that supports natural speech patterns: `from livekit.plugins.deepgram import tts deepgram_tts = tts.TTS( model="aura-asteria-en", )`If using a custom TTS, ensure it generates tokens incrementally to allow for pseudo-streaming and quicker responses. • Performance Optimization: The turn detection system achieves high accuracy (98.8% for completed turns and 87.5% for incomplete utterances). Monitor and tweak these configurations to suit your specific use case. Combining these strategies will help improve the naturalness and responsiveness of your calls. Let me know if you need further details or assistance with implementation.