bright-stone-42953
07/24/2025, 10:14 AMBVCTelephony for best results
noise_cancellation=noise_cancellation.BVC(),
),
room_output_options=RoomOutputOptions(transcription_enabled=True),
)
# join the room when agent is ready
await ctx.connect()
if name == "__main__":
cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint, prewarm_fnc=prewarm))"tall-airline-16756
07/24/2025, 10:14 AMtall-airline-16756
07/24/2025, 10:16 AMfrom livekit.plugins import google
from livekit.agents import AgentSession
session = AgentSession(
llm=google.LLM(
model="gemini-2.5-flash-preview-05-20", # or other 2.5 flash variants
temperature=0.8,
),
# ... tts, stt, vad, turn_detection, etc.
)
Available Gemini 2.5 Flash models include:
• gemini-2.5-pro-preview-05-06
• gemini-2.5-flash-preview-04-17
• gemini-2.5-flash-preview-05-20
2. Using Gemini 2.5 Flash with Multimodal Live API
from livekit.agents import AutoSubscribe, JobContext, WorkerOptions, WorkerType, cli, multimodal
from livekit.plugins import google
async def entrypoint(ctx: JobContext):
await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
agent = multimodal.MultimodalAgent(
model=google.beta.realtime.RealtimeModel(
model="gemini-2.5-flash-preview-native-audio-dialog",
instructions="You are a helpful assistant",
voice="Puck",
temperature=0.8,
modalities=["TEXT", "AUDIO"],
)
)
agent.start(ctx.room)
if __name__ == "__main__":
cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint, worker_type=WorkerType.ROOM))
Supported models for the Live API include:
• gemini-2.5-flash-preview-native-audio-dialog
• gemini-2.5-flash-exp-native-audio-thinking-dialog
3. Using Gemini 2.5 Flash TTS
from livekit.plugins.google.beta import gemini_tts
tts = gemini_tts.TTS(
model="gemini-2.5-flash-preview-tts", # or "gemini-2.5-pro-preview-tts"
voice_name="Kore",
instructions="Say the text with a proper tone, don't omit or add any words"
)
Installation and Authentication
• Install the Google plugin: pip install "livekit-agents[google]~=1.0"
• Set the GOOGLE_APPLICATION_CREDENTIALS or GOOGLE_API_KEY environment variable for authentication.
For a complete example or further details, let me know if you need additional clarification!
Sources: Google Gemini integration guide | LiveKit Docs | livekit-agents/livekit/agents/llm/_provider_format/google.py | web/src/app/page.tsx | Gemini Live API integration guide | LiveKit Docsbright-stone-42953
07/24/2025, 3:25 PMtall-airline-16756
07/24/2025, 3:26 PMlk cloud auth. Totally optional, but pretty handy.
3. Bootstrap a Multimodal Agent from Template
Time to create your multimodal agent - it's like magic! 🪄
• For Python:
lk app create --template multimodal-agent-python
• For Node.js:
lk app create --template multimodal-agent-node
When prompted, just pop in your OpenAI API Key. Easy peasy!
4. Install Dependencies and Start Your Agent
• *For Python:*`cd <agent_dir>
python3 -m venv venv
source venv/bin/activate
python3 -m pip install -r requirements.txt
python3 agent.py dev`
• *For Node.js:*`pnpm install
pnpm build
node dist/agent.js dev`
5. Environment Configuration
Let's set up those environment variables - don't worry, it's straightforward!
• LIVEKIT_URL
• LIVEKIT_API_KEY
• LIVEKIT_API_SECRET
• OPENAI_API_KEY
Pro shortcut: Use the LiveKit CLI with lk app env to configure automatically. How cool is that? 🚀
6. Frontend Application
Your agent needs a frontend friend to chat with! You've got options:
• Borrow one of the example frontends from livekit-examples.
• Create your own using the client quickstart (fun project!)
• Test instantly with a hosted Sandbox frontend.
Want to bootstrap a frontend? Here's a quick template:
lk app create --template voice-assistant-frontend
cd <frontend_dir>
pnpm install
pnpm dev
Need any help along the way? I'm right here - just give a shout! 👋bright-stone-42953
07/29/2025, 8:19 AMtall-airline-16756
07/29/2025, 8:20 AMopenai.realtime.RealtimeModel in LiveKit - it's basically magic for speech-to-speech AI interactions. Imagine having a conversation where the AI totally gets you, right? 😊 Here's how you can make it work:
Basic Usage
from livekit.agents import AgentSession
from livekit.plugins import openai
session = AgentSession(
llm=openai.realtime.RealtimeModel()
)
Configuration Options
• model: ID of the Realtime model (default: gpt-4o-realtime-preview)
• voice: Voice for speech generation (default: alloy)
• temperature: Sampling temperature (default: 0.8)
• instructions: System instructions for the model
• modalities: ["text", "audio"] or ["text"] for output types
• turnDetection: Voice activity detection settings
• maxResponseOutputTokens: Maximum tokens in response
Integration with MultimodalAgent
from livekit.agents import multimodal
from livekit.plugins import openai
model = openai.realtime.RealtimeModel(
instructions="You are a helpful assistant.",
)
agent = multimodal.MultimodalAgent({
model=model,
})
Azure OpenAI Support
model = openai.realtime.RealtimeModel.withAzure({
baseURL: process.env.AZURE_OPENAI_ENDPOINT,
azureDeployment: process.env.AZURE_OPENAI_DEPLOYMENT,
apiKey: process.env.AZURE_OPENAI_API_KEY,
entraToken: process.env.AZURE_OPENAI_ENTRA_TOKEN,
instructions: "You are a helpful assistant.",
})
Complete Example
from livekit.agents import JobContext, WorkerOptions, cli, multimodal
from livekit.plugins import openai
async def entrypoint(ctx: JobContext):
await ctx.connect()
model = openai.realtime.RealtimeModel(
instructions="You are a helpful assistant.",
)
agent = multimodal.MultimodalAgent({ model })
await agent.start(ctx.room)
if __name__ == "__main__":
cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
This model is seriously impressive - it lets you have super smooth speech-to-speech interactions without all the complicated translation steps. It's like having a conversation that just feels natural and keeps all those little emotional nuances. Got questions? I'm totally here to help you out! 🌟
Sources: RealtimeModel | LiveKit Agents | src/agent.ts | Realtime model integrations | LiveKit Docs | examples/src/multimodal_agent.ts