no matter which TTS model I use, the agent's voice...
# ask-ai
d
no matter which TTS model I use, the agent's voice still sounds robotic. Is there any way to make it sound more human?"
t
Thanks for reaching out, Ibragim. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To make the TTS (Text-to-Speech) output sound more human-like, you can try the following approaches: • Choose Human-Like TTS Providers: LiveKit integrates with several TTS providers that specialize in realistic speech, such as: ◦ ElevenLabs: Offers thousands of human-like voices and allows fine-tuning for natural speech. ◦ OpenAI TTS: Provides lifelike spoken audio with models like
gpt-4o-mini-tts
. ◦ Deepgram: Features responsive, human-like TTS technology. ◦ Google Cloud TTS: Generates speech with human-like intonation. • Configure Voice Settings: For example, with ElevenLabs, you can fine-tune voice characteristics using parameters like:`settings=elevenlabs.tts.VoiceSettings( stability=0.71, # Controls voice consistency similarity_boost=0.5, # Enhances voice similarity to original style=0.0, # Adjusts speaking style use_speaker_boost=True # Improves voice clarity )` • Optimize Pronunciation: For specialized terms or numerical values, you can adjust pronunciation programmatically using provider-specific techniques. • Select Appropriate Models and Voices: Each TTS provider offers a variety of models and voices. For example: ◦ ElevenLabs: Models like
eleven_turbo_v2_5
. ◦ OpenAI: Models like
gpt-4o-mini-tts
or
tts-1-hd
. • Enable Streaming: Use streaming audio capabilities to create a more natural conversation flow by delivering audio responses in real-time. These adjustments and configurations can significantly enhance the naturalness of the TTS output. Let me know if you need further details on implementing any of these options. Sources: Speech & Audio Processing | livekit/livekit_composite | DeepWiki | ElevenLabs integration guide | LiveKit Docs | OpenAI TTS integration guide | LiveKit Docs | Deepgram TTS integration guide | LiveKit Docs | Google Cloud TTS integration guide | LiveKit Docs
VoiceSettings might seem a bit tricky at first, but don't worry! It's actually pretty straightforward. Instead of configuring it directly in AgentSession, you'll set it up when creating your TTS instance, which you'll then pass to the AgentSession constructor. Let me walk you through it: • *Configure VoiceSettings in the TTS Instance:*`from livekit.plugins.elevenlabs import TTS, VoiceSettings tts = TTS( voice_settings=VoiceSettings( stability=0.5, similarity_boost=0.75, speed=1.15 ) )` • *Pass the TTS Instance to AgentSession:*`session = AgentSession( tts=tts, # TTS instance with VoiceSettings already configured # other parameters... )` Think of it like prepping a recipe before cooking – you're getting all your ingredients (voice settings) ready before putting everything together. This way, your TTS is perfectly configured before joining the session. Got any questions? I'm here to help you navigate this! Feel free to ask if anything isn't crystal clear. Sources: livekit-agents/livekit/agents/voice/agent.py
d
can you give me full agent session?
t
Here is a comprehensive example of an AgentSession configuration in LiveKit, showcasing various components and options - and trust me, it's pretty cool! 😊
Copy code
import asyncio
from livekit.agents import (
    Agent,
    AgentSession,
    AutoSubscribe,
    RoomInputOptions,
    RoomOutputOptions,
    RunContext,
    JobContext,
    function_tool,
)
from livekit.plugins import aws, deepgram, openai, silero
from livekit.plugins.turn_detector.multilingual import MultilingualModel

# Define custom agent with tools
class Assistant(Agent):
    def __init__(self, tools):
        super().__init__(
            instructions="You are a helpful voice AI assistant.",
            tools=tools,
        )

# Function tools example
@function_tool
async def get_weather(location: str) -> dict:
    """Get weather for a location."""
    return {"weather": "sunny", "temperature": 70}

async def entrypoint(ctx: JobContext):
    await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
    
    # Wait for participant
    participant = await ctx.wait_for_participant()
    
    # Create AgentSession with comprehensive configuration
    session = AgentSession(
        # LLM Configuration - using AWS Realtime Model
        llm=aws.realtime.RealtimeModel(
            tool_choice="auto",
            max_tokens=10_000,
        ),
        # Alternative: Standard LLM configuration
        # llm=openai.LLM(model="gpt-4o-mini"),
        # stt=deepgram.STT(model="nova-3", language="multi"),
        # tts=openai.TTS(voice="ash"),
        
        # Voice Activity Detection
        vad=silero.VAD.load(),
        
        # Turn Detection
        turn_detection=MultilingualModel(),
    )
    
    # Start session with comprehensive options
    await session.start(
        room=ctx.room,
        agent=Assistant(tools=[get_weather]),
        room_input_options=RoomInputOptions(
            close_on_disconnect=False,
            # Optional: Enable noise cancellation
            # noise_cancellation=noise_cancellation.BVC(),
        ),
        room_output_options=RoomOutputOptions(
            audio_enabled=True,
            transcription_enabled=True,
        ),
    )
Let me break this down for you in a friendly way! This example is packed with awesome features: • LLM: Think of this like your AI brain - we've got AWS Realtime Model or OpenAI LLM ready to chat. • STT: Speech-to-text that turns your voice into text super smoothly using Deepgram. • TTS: Text-to-speech that can make your AI sound just like you want using OpenAI or other cool providers. • VAD: Voice Activity Detection that knows exactly when you're speaking (thanks, Silero!). • Turn Detection: A multilingual wizard that understands when you're done talking. • Room Input/Output Options: Flexible audio settings that let you customize your experience. • Custom Tools: Like our fun weather tool - imagine asking your AI about the temperature! Got questions? I'm all ears! Feel free to ask, and I'll help you dive deeper into this awesome setup. 🌟 Sources: Building voice agents | LiveKit Docs | examples/voice_agents/langgraph_agent.py
d
where is Vocie settings for elevenlabs?)
t
To configure VoiceSettings for ElevenLabs in LiveKit, you'll want to create a VoiceSettings object and pass it to the ElevenLabs TTS configuration. Don't worry, I'll walk you through it step by step – it's easier than you might think!
Copy code
from livekit.plugins.elevenlabs import TTS, VoiceSettings

tts = TTS(
    model="eleven_turbo_v2_5",
    voice_id="EXAVITQu4vr4xnSDxMaL",
    voice_settings=VoiceSettings(
        stability=0.71,           # Controls voice consistency (0.0 - 1.0)
        similarity_boost=0.5,     # Enhances voice similarity to original (0.0 - 1.0)
        style=0.0,                # Adjusts speaking style (0.0 - 1.0)
        use_speaker_boost=True    # Improves voice clarity
    )
)
### VoiceSettings Parameters Let me break down these settings for you – they might look technical, but they're pretty straightforward: • stability: Controls voice stability (range: 0.0 - 1.0). • similarity_boost: Controls similarity boost (range: 0.0 - 1.0). • style: Adjusts speaking style (range: 0.0 - 1.0). • use_speaker_boost: Enables or disables speaker boost. • speed: (Optional) Controls speech speed (range: 0.8 - 1.2). ### Updating VoiceSettings at Runtime The cool thing is, you can totally tweak these settings on the fly:
Copy code
tts.update_options(
    voice_settings=VoiceSettings(
        stability=0.8,
        similarity_boost=0.6,
        style=0.2,
        use_speaker_boost=False
    )
)
This means you can adjust the voice characteristics whenever you want – pretty neat, right? If you're scratching your head about any of this, just give me a shout, and I'll be happy to help you out! Sources: ElevenLabs integration guide | LiveKit Docs | livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/model... | ElevenLabs TTS integration guide | LiveKit Docs
d
can i apply the voice settings to realtime models?
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!