dry-helmet-74553
07/28/2025, 1:49 PMtall-airline-16756
07/28/2025, 1:49 PMtall-airline-16756
07/28/2025, 1:50 PMgpt-4o-mini-tts.
◦ Deepgram: Features responsive, human-like TTS technology.
◦ Google Cloud TTS: Generates speech with human-like intonation.
• Configure Voice Settings: For example, with ElevenLabs, you can fine-tune voice characteristics using parameters like:`settings=elevenlabs.tts.VoiceSettings(
stability=0.71, # Controls voice consistency
similarity_boost=0.5, # Enhances voice similarity to original
style=0.0, # Adjusts speaking style
use_speaker_boost=True # Improves voice clarity
)`
• Optimize Pronunciation: For specialized terms or numerical values, you can adjust pronunciation programmatically using provider-specific techniques.
• Select Appropriate Models and Voices: Each TTS provider offers a variety of models and voices. For example:
◦ ElevenLabs: Models like eleven_turbo_v2_5.
◦ OpenAI: Models like gpt-4o-mini-tts or tts-1-hd.
• Enable Streaming: Use streaming audio capabilities to create a more natural conversation flow by delivering audio responses in real-time.
These adjustments and configurations can significantly enhance the naturalness of the TTS output. Let me know if you need further details on implementing any of these options.
Sources: Speech & Audio Processing | livekit/livekit_composite | DeepWiki | ElevenLabs integration guide | LiveKit Docs | OpenAI TTS integration guide | LiveKit Docs | Deepgram TTS integration guide | LiveKit Docs | Google Cloud TTS integration guide | LiveKit Docstall-airline-16756
07/28/2025, 1:59 PMdry-helmet-74553
07/28/2025, 2:04 PMtall-airline-16756
07/28/2025, 2:05 PMimport asyncio
from livekit.agents import (
Agent,
AgentSession,
AutoSubscribe,
RoomInputOptions,
RoomOutputOptions,
RunContext,
JobContext,
function_tool,
)
from livekit.plugins import aws, deepgram, openai, silero
from livekit.plugins.turn_detector.multilingual import MultilingualModel
# Define custom agent with tools
class Assistant(Agent):
def __init__(self, tools):
super().__init__(
instructions="You are a helpful voice AI assistant.",
tools=tools,
)
# Function tools example
@function_tool
async def get_weather(location: str) -> dict:
"""Get weather for a location."""
return {"weather": "sunny", "temperature": 70}
async def entrypoint(ctx: JobContext):
await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
# Wait for participant
participant = await ctx.wait_for_participant()
# Create AgentSession with comprehensive configuration
session = AgentSession(
# LLM Configuration - using AWS Realtime Model
llm=aws.realtime.RealtimeModel(
tool_choice="auto",
max_tokens=10_000,
),
# Alternative: Standard LLM configuration
# llm=openai.LLM(model="gpt-4o-mini"),
# stt=deepgram.STT(model="nova-3", language="multi"),
# tts=openai.TTS(voice="ash"),
# Voice Activity Detection
vad=silero.VAD.load(),
# Turn Detection
turn_detection=MultilingualModel(),
)
# Start session with comprehensive options
await session.start(
room=ctx.room,
agent=Assistant(tools=[get_weather]),
room_input_options=RoomInputOptions(
close_on_disconnect=False,
# Optional: Enable noise cancellation
# noise_cancellation=noise_cancellation.BVC(),
),
room_output_options=RoomOutputOptions(
audio_enabled=True,
transcription_enabled=True,
),
)
Let me break this down for you in a friendly way! This example is packed with awesome features:
• LLM: Think of this like your AI brain - we've got AWS Realtime Model or OpenAI LLM ready to chat.
• STT: Speech-to-text that turns your voice into text super smoothly using Deepgram.
• TTS: Text-to-speech that can make your AI sound just like you want using OpenAI or other cool providers.
• VAD: Voice Activity Detection that knows exactly when you're speaking (thanks, Silero!).
• Turn Detection: A multilingual wizard that understands when you're done talking.
• Room Input/Output Options: Flexible audio settings that let you customize your experience.
• Custom Tools: Like our fun weather tool - imagine asking your AI about the temperature!
Got questions? I'm all ears! Feel free to ask, and I'll help you dive deeper into this awesome setup. 🌟 Sources: Building voice agents | LiveKit Docs | examples/voice_agents/langgraph_agent.pydry-helmet-74553
07/28/2025, 2:07 PMtall-airline-16756
07/28/2025, 2:08 PMfrom livekit.plugins.elevenlabs import TTS, VoiceSettings
tts = TTS(
model="eleven_turbo_v2_5",
voice_id="EXAVITQu4vr4xnSDxMaL",
voice_settings=VoiceSettings(
stability=0.71, # Controls voice consistency (0.0 - 1.0)
similarity_boost=0.5, # Enhances voice similarity to original (0.0 - 1.0)
style=0.0, # Adjusts speaking style (0.0 - 1.0)
use_speaker_boost=True # Improves voice clarity
)
)
### VoiceSettings Parameters Let me break down these settings for you – they might look technical, but they're pretty straightforward:
• stability: Controls voice stability (range: 0.0 - 1.0).
• similarity_boost: Controls similarity boost (range: 0.0 - 1.0).
• style: Adjusts speaking style (range: 0.0 - 1.0).
• use_speaker_boost: Enables or disables speaker boost.
• speed: (Optional) Controls speech speed (range: 0.8 - 1.2).
### Updating VoiceSettings at Runtime The cool thing is, you can totally tweak these settings on the fly:
tts.update_options(
voice_settings=VoiceSettings(
stability=0.8,
similarity_boost=0.6,
style=0.2,
use_speaker_boost=False
)
)
This means you can adjust the voice characteristics whenever you want – pretty neat, right? If you're scratching your head about any of this, just give me a shout, and I'll be happy to help you out! Sources: ElevenLabs integration guide | LiveKit Docs | livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/model... | ElevenLabs TTS integration guide | LiveKit Docsdry-helmet-74553
07/28/2025, 2:23 PMtall-airline-16756
07/28/2025, 2:23 PM