To decrease latency when using Gemini Beta TTS, consider the following strategies:
•
Network Proximity: Ensure your agent is close to the Text-to-Speech service in terms of network latency.
•
Metrics Monitoring: Use the Agents metrics API to monitor latency metrics like Time To First Byte (TTFB) and optimize based on these insights.
•
Configuration Options: Adjust parameters such as model selection (e.g., "gemini-2.5-flash-preview-tts"), voice selection, and instructions for style, tone, and pace. For VertexAI users, configure the location parameter (default is "us-central1").
•
Consider Alternatives: If latency is critical, you might explore using the Gemini Live API, which supports low-latency, real-time interactions. This API can also be configured with a different TTS provider for more control over speech output.
Note that Gemini Beta TTS is a non-streaming service, which inherently results in higher latency compared to streaming TTS services. If you need further assistance or specific guidance, let me know!
Sources:
AI Agents Framework | livekit/livekit_composite | DeepWiki |
examples/voice_agents/preemptive_generation.py |
Gemini Live API integration guide | LiveKit Docs |
Rime integration guide | LiveKit Docs |
How can I reduce latency in voice agents using STT, TTS and LLM?