chilly-ghost-74326
07/28/2025, 1:11 PMtall-airline-16756
07/28/2025, 1:11 PMtall-airline-16756
07/28/2025, 1:12 PMAgentSession or MultimodalAgent. Below are the steps for each approach:
1. Using AgentSession (Basic Approach):
from dotenv import load_dotenv
from livekit import agents
from livekit.agents.voice import AgentSession, Agent
from livekit.plugins import openai, silero
load_dotenv()
class Assistant(Agent):
def __init__(self):
super().__init__(instructions="You are a helpful voice AI assistant.")
async def entrypoint(ctx: agents.JobContext):
session = AgentSession(
llm=openai.realtime.RealtimeModel(),
vad=silero.VAD.load()
)
await session.start(
room=ctx.room,
agent=Assistant()
)
await session.generate_reply()
if __name__ == "__main__":
agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint))
2. Using MultimodalAgent (Advanced Approach):
from livekit import rtc
from livekit.agents import JobContext, AutoSubscribe, llm
from livekit.agents.multimodal import MultimodalAgent
from livekit.plugins import openai
async def entrypoint(ctx: JobContext):
await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
participant = await ctx.wait_for_participant()
model = openai.realtime.RealtimeModel(
instructions=(
"You are a voice assistant created by LiveKit. Your interface with users will be voice. "
"You should use short and concise responses, avoiding usage of unpronounceable punctuation."
),
modalities=["audio", "text"],
)
chat_ctx = llm.ChatContext()
chat_ctx.append(
text="Context about the user: you are talking to a software engineer who's building voice AI applications.",
role="assistant",
)
agent = MultimodalAgent(
model=model,
chat_ctx=chat_ctx,
)
agent.start(ctx.room, participant)
agent.generate_reply()
Key Components:
• RealtimeModel: Integrates with OpenAI's Realtime API.
• AgentSession: Manages the voice interaction session.
• MultimodalAgent: Enables complex multimodal interactions.
• VAD (Voice Activity Detection): Often uses Silero for detecting speech.
Choose the approach that best fits your requirements. The AgentSession is simpler, while the MultimodalAgent offers advanced multimodal capabilities.
Sources: realtime/openai-realtime-pitch-shift.py | agent.py | agent/main.py | Agent speech | LiveKit Docs | livekit-agents/README.mdchilly-ghost-74326
07/28/2025, 4:06 PMtall-airline-16756
07/28/2025, 4:06 PM