llm=openai.realtime.RealtimeModel( model: str ...
# ask-ai
b
llm=openai.realtime.RealtimeModel( model: str = "gpt-4o-realtime-preview", voice: str = "alloy", modalities: NotGivenOr[list[Literal['text', 'audio']]] = NOT_GIVEN, input_audio_transcription: NotGivenOr[InputAudioTranscription | None] = NOT_GIVEN, input_audio_noise_reduction: InputAudioNoiseReduction | None = None, turn_detection: NotGivenOr[TurnDetection | None] = NOT_GIVEN, temperature: NotGivenOr[float] = NOT_GIVEN, tool_choice: NotGivenOr[ToolChoice | None] = NOT_GIVEN, speed: NotGivenOr[float] = NOT_GIVEN, tracing: NotGivenOr[Tracing | None] = NOT_GIVEN, api_key: str | None = None, base_url: NotGivenOr[str] = NOT_GIVEN, http_session: ClientSession | None = None, max_session_duration: NotGivenOr[float | None] = NOT_GIVEN, conn_options: APIConnectOptions = DEFAULT_API_CONNECT_OPTIONS )
t
Thanks for reaching out, Reuban. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Could you clarify what assistance you need with the code snippet you provided? For example, are you encountering an error, looking for implementation guidance, or something else?
b
looking for implementation guidance,
t
Hey there! To get started with the
RealtimeModel
class, here’s a simple guide you can follow depending on what you need and which provider you're working with: Python Implementation for OpenAI
Copy code
from livekit.plugins.openai.realtime import RealtimeModel

# Here’s a basic setup
model = RealtimeModel(
    model="gpt-4o-realtime-preview",  # This is the default model
    voice="alloy",  # Default voice choice
    modalities=["text", "audio"],  # Supports both text and audio
    temperature=0.8,  # Controls creativity level
    input_audio_transcription=None,
    turn_detection=None,
    tool_choice=None,
    speed=None,
    api_key="your-api-key"  # Or you can set the OPENAI_API_KEY environment variable
)
JavaScript/TypeScript Implementation
Copy code
import { RealtimeModel } from '@livekit/agents-plugin-openai';

const model = new RealtimeModel({
    model: 'gpt-4o-realtime-preview-2024-10-01',
    voice: 'alloy',
    modalities: ['text', 'audio'],
    temperature: 0.8,
    inputAudioFormat: 'pcm16',
    outputAudioFormat: 'pcm16',
    inputAudioTranscription: { model: 'whisper-1' },
    turnDetection: { type: 'server_vad' },
    maxResponseOutputTokens: Infinity,
    apiKey: process.env.OPENAI_API_KEY
});
Quick rundown of key parameters • modalities: This tells the model what kinds of inputs and outputs it should handle, like "text", "audio", or both. • voice: Picks the voice style for audio replies (options like "alloy", "echo", or "shimmer" if you’re using OpenAI). • temperature: This sets how creative or focused the responses are. Higher means more creative, lower means more precise. • instructions: Use this to add custom instructions that act like a 'system prompt' for the model. If you want me to walk you through anything else or need examples for a different provider or parameter, just give me a shout! Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime... | plugins/openai/src/realtime/realtime_model.ts | Adjusting Model Parameters | LiveKit Docs | Working with the MultimodalAgent class | LiveKit Docs
b
preemptive_generation "async def entrypoint(ctx: agents.JobContext): session = AgentSession( stt=deepgram.STT( model="nova-3", language="multi", detect_language=True, interim_results=True, no_delay=True, sample_rate=16000 ), llm=openai.LLM( model="gpt-4o-mini", temperature=0.8 ), tts=cartesia.TTS( model="sonic-2", voice="f786b574-daa5-4673-aa0c-cbe3e8534c02", language="en", speed=1.0 ), vad=silero.VAD.load( min_speech_duration=0.05, min_silence_duration=0.3, activation_threshold=0.5, force_cpu=True ), turn_detection=MultilingualModel( min_endpointing_delay=300 ), preemptive_generation=True, allow_interruptions=True )"
t
Hey there! Could you share a bit more about what kind of help you’re looking for with the new code snippet? Are you after some tips on how to implement it, need a hand troubleshooting, or something else? I’m here to help!
b
I LIKE TO KNOW MORE ABOUT preemptive_generation=True, allow_interruptions=True
t
Hey there! Let me break down those two parameters you asked about in a simple way: preemptive_generation=True • This means the agent starts crafting a response as soon as it catches even part of what you're saying, without waiting for you to finish completely. • It helps make the conversation flow smoother by cutting down the wait time between your input and the agent’s reply. • If you happen to change your mind or correct yourself mid-sentence, the agent will stop the old response and start fresh based on what you just said. • This is super handy when the speech-to-text system picks up your words faster than the system can tell when you’ve stopped talking. allow_interruptions=True • This lets you jump in and interrupt the agent while it’s talking, making the chat feel more natural and flexible. • It's perfect for real-time chats where you might want to steer the conversation or fix something without waiting for the agent to finish. • If you set this to
False
, the agent will always finish its response before listening to anything new you say. Both of these settings are all about making your conversations with the agent quicker and more interactive. If you want me to go over anything else or give some examples, just shout! Sources: tests/test_agent_session.py | Turn detection and interruptions | LiveKit Docs | Agent speech and audio | LiveKit Docs | Agent speech and audio | LiveKit Docs
b
set a smoother interruption
t
If you're looking to make interruptions in LiveKit's
AgentSession
feel a bit smoother when
allow_interruptions=True
, here are some handy tweaks you can try out: Core Interruption Parameters • min_interruption_duration: This sets the minimum speech length that counts as an interruption. It’s
0.5
seconds by default. To avoid interruptions from quick sounds or pauses, bump it up to around
0.8
or
1.0
seconds. • min_interruption_words: If you’re using STT, this controls how many words need to be spoken before it’s considered an interruption. The default is
0
, but setting it to
2
or
3
helps make sure it’s real speech before interrupting. Turn Detection Parameters • min_endpointing_delay: This is how long the agent waits after thinking the user might be done speaking before it actually ends their turn. The default is
0.4
seconds, but increasing it to
0.6
or
0.8
seconds gives users a bit more breathing room to keep talking. • max_endpointing_delay: This sets the max wait time before the agent ends the turn. You can tweak this depending on how natural or quick you want the flow to be. Example Configuration
Copy code
from livekit.agents import AgentSession

session = AgentSession(
    # Enable interruptions but make them smoother
    allow_interruptions=True,
    
    # Require longer speech duration before interrupting
    min_interruption_duration=0.8,  # Increased from default 0.5
    
    # Require at least 2 words for interruption (if using STT)
    min_interruption_words=2,  # Increased from default 0
    
    # Give more time before considering turn complete
    min_endpointing_delay=0.6,  # Increased from default 0.4
    
    # Other parameters...
    stt=your_stt,
    llm=your_llm,
    tts=your_tts,
)
VAD Configuration for Smoother Detection • activationThreshold: Setting this higher (like
0.6
to
0.7
) makes voice activity detection more cautious. • minSpeechDuration: Increase this to require longer speech segments before detecting activity. • minSilenceDuration: Adjust this to fine-tune how sensitive it is to silence. The best approach is to start with these conservative settings, see how it feels with real users, and then tweak as needed to find the right balance between being responsive and avoiding accidental interruptions. If you want a hand with this or have any questions, just let me know—I’m here to help! Sources: livekit-agents/livekit/agents/voice/agent_activity.py | Configuring turn detection and interruptions | LiveKit Docs
b
turn_detection=MultilingualModel( min_endpointing_delay=300), is this correct
t
Hey there! Just a quick heads-up: using
turn_detection=MultilingualModel(min_endpointing_delay=300)
isn’t quite right. The
min_endpointing_delay
parameter shouldn’t be passed directly to
MultilingualModel
. The better way to do this is to pass
min_endpointing_delay
to
AgentSession
instead, not to
MultilingualModel
. Here’s how you can set it up:
Copy code
from livekit.plugins.turn_detector.multilingual import MultilingualModel

session = AgentSession(
    turn_detection=MultilingualModel(),
    min_endpointing_delay=0.3,  # 300ms as 0.3 seconds
    # ... other parameters
)
Just so you know,
MultilingualModel
doesn’t take any parameters when you create it. So if you want a 300ms delay, using
min_endpointing_delay=0.3
(in seconds) like above is the way to go. If you have any questions or want me to walk you through it, feel free to ask—I’m here to help! Sources: LiveKit turn detector plugin | LiveKit Docs | examples/voice_agents/mcp/mcp-agent.py
b
list me all the possible parameters for stt, llm, tts, vad, turn detection
t
Hey there! Here's a handy list of possible parameters you can use to configure STT, LLM, TTS, VAD, and turn detection in LiveKit's
AgentSession
. Hopefully, this makes things a bit clearer for you: STT (Speech-to-Text) Parameters • model: The name of the STT model you'd like to use (for example, "whisper-large-v3-turbo"). • language: The language for speech recognition (like "en" for English). • keyterms: A list of keywords to help improve recognition accuracy. • mip_opt_out: An option to opt out of certain processing features if you prefer. LLM (Large Language Model) Parameters • model: The LLM model you want to use (such as "gpt-4o-mini"). • temperature: Controls how random the generation is (default is 1.0). • max_tokens: The maximum number of tokens to generate. • parallel_tool_calls: Whether to allow parallel tool calls or not. TTS (Text-to-Speech) Parameters • model: The TTS model you want to use (like "sonic-2"). • voice: The voice ID for the speech synthesis. • speed: Control the speech rate (e.g., "fast"). VAD (Voice Activity Detection) Parameters • activationThreshold: How sensitive the system is to detecting voice activity (e.g., 0.6). • minSpeechDuration: The minimum length of speech to consider it active. • minSilenceDuration: The minimum silence length to consider it inactive. Turn Detection Parameters • allow_interruptions: Can the user interrupt the agent mid-sentence? (default: True) • min_interruption_duration: The shortest speech length to count as an interruption (default: 0.5 seconds). • min_interruption_words: Minimum words needed to consider it an interruption (default: 0). • min_endpointing_delay: How long to wait after a possible end-of-utterance before deciding the user's turn is done (default: 0.4 seconds). • max_endpointing_delay: The longest to wait before ending the turn (default: 6.0 seconds). Additional Parameters • preemptive_generation: Whether to start LLM and TTS requests early, before the turn ends (default: False). • discard_audio_if_uninterruptible: Whether to drop buffered audio if the agent is speaking and can't be interrupted (default: True). • user_away_timeout: Time after which the user is marked as "away" if both user and agent are silent (default: 15.0 seconds). Just a heads up: the exact parameters you can use for STT, LLM, TTS, and VAD might vary depending on the provider you pick (like OpenAI, Groq, Anthropic, Cartesia, Deepgram, and so on). If you want me to dive deeper into any specific parameter or provider, just let me know – I’m here to help! Sources: Voice AI quickstart | LiveKit Docs | livekit-agents/livekit/agents/voice/agent.py | Groq STT integration guide | LiveKit Docs | complex-agents/drive-thru/agent_config.py | Groq LLM integration guide | LiveKit Docs
b
write an example code for the above list of all parameters
t
Hey there! Here’s a friendly example code snippet that shows you how to set up all the parameters for
STT
,
LLM
,
TTS
,
VAD
, and turn detection in `AgentSession`:
Copy code
from livekit.agents import AgentSession
from livekit.plugins.stt import DeepgramSTT
from livekit.plugins.llm import OpenAILLM
from livekit.plugins.tts import CartesiaTTS
from livekit.plugins.vad import SileroVAD
from livekit.plugins.turn_detector.multilingual import MultilingualModel

# Configure STT
stt = DeepgramSTT(
    model="nova-3",
    language="en",
    detect_language=True,
    interim_results=True,
    no_delay=True,
    sample_rate=16000
)

# Configure LLM
llm = OpenAILLM(
    model="gpt-4o-mini",
    temperature=0.8,
    max_tokens=1000,
    parallel_tool_calls=True
)

# Configure TTS
tts = CartesiaTTS(
    model="sonic-2",
    voice="f786b574-daa5-4673-aa0c-cbe3e8534c02",
    language="en",
    speed=1.0
)

# Configure VAD
vad = SileroVAD(
    activation_threshold=0.5,
    min_speech_duration=0.05,
    min_silence_duration=0.3,
    force_cpu=True
)

# Configure AgentSession with all parameters
session = AgentSession(
    stt=stt,
    llm=llm,
    tts=tts,
    vad=vad,
    turn_detection=MultilingualModel(),
    allow_interruptions=True,
    discard_audio_if_uninterruptible=True,
    min_interruption_duration=0.8,
    min_interruption_words=2,
    min_endpointing_delay=0.6,
    max_endpointing_delay=6.0,
    preemptive_generation=True,
    user_away_timeout=15.0
)
So, this example walks you through setting up each piece with the right parameters and then bringing them all together in an
AgentSession
. If anything’s unclear or you want me to share more examples, just give me a shout—I’m here to help!