dry-helmet-74553
07/28/2025, 4:25 PMAgentSession, but I still haven't found the secret sauce.
When I use the real-time model, it skips the greeting and jumps straight into the conversation. But when I switch to regular TTS/STT, the LLM response latency increases — I get replies with a 3–5 second delay.
Does anyone know how to fix or optimize these issues?tall-airline-16756
07/28/2025, 4:25 PMtall-airline-16756
07/28/2025, 4:26 PMAgentSession, here are some optimization strategies:
• 1. Reduce Latency:
◦ Ensure your agent is close (in terms of network latency) to your LLM, Speech-to-Text (STT), and Text-to-Speech (TTS) services.
◦ Choose STT, TTS, and LLM providers with low latency characteristics for your use case.
• *2. Enable Preemptive Generation:*`session = AgentSession(
preemptive_generation=True,
# ... STT, LLM, TTS, etc.
)`This allows the agent to speculatively begin LLM and TTS requests before detecting an end-of-turn.
◦ Configure your AgentSession with preemptive generation to reduce response latency:
• *3. Optimize Greeting Delivery:*`async def on_enter(self):
self.session.generate_reply(instructions="say hello to the user")`Alternatively, provide pre-synthesized audio for playback to skip the TTS step and reduce response time.
◦ Use session.say() to have the agent speak predefined messages:
• *4. Configure AgentSession Parameters:*`session = AgentSession(
min_interruption_duration=0.5, # Minimum speech length to register interruption
min_endpointing_delay=0.4, # Minimum wait after end-of-utterance signal
max_endpointing_delay=6.0, # Maximum wait before terminating turn
allow_interruptions=True, # Enable user interruptions
preemptive_generation=True, # Enable preemptive generation
# ... other components
)`
◦ Adjust key timing parameters for better real-time performance:
• *5. Monitor Performance Metrics:*`from livekit.agents.metrics import UsageCollector
usage_collector = UsageCollector()
session = AgentSession(
# ... other config
metrics_handler=usage_collector
)`
◦ Use the Agents metrics API to track and optimize latency metrics like Time To First Token (TTFT) and Time To First Byte (TTFB):
These steps should help reduce latency and ensure greetings are delivered reliably. Let me know if you need further clarification or assistance!
Sources: How can I reduce latency in voice agents using STT, TTS and LLM? | Voice Assistants | livekit/livekit_composite | DeepWiki | Agent speech and audio | LiveKit Docs | examples/voice_agents/preemptive_generation.py | Speech & Audio Processing | livekit/livekit_composite | DeepWiki