I want a voice agent in open ai real time. How can...
# ask-ai
c
I want a voice agent in open ai real time. How can I initialize livekit agent?
t
Thanks for reaching out, Rashmi. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To initialize a LiveKit voice agent for OpenAI real-time applications, you can use one of two main approaches:
AgentSession
or
MultimodalAgent
. Below are the steps for each approach: 1. Using AgentSession (Basic Approach):
Copy code
from dotenv import load_dotenv
from livekit import agents
from livekit.agents.voice import AgentSession, Agent
from livekit.plugins import openai, silero

load_dotenv()

class Assistant(Agent):
    def __init__(self):
        super().__init__(instructions="You are a helpful voice AI assistant.")

async def entrypoint(ctx: agents.JobContext):
    session = AgentSession(
        llm=openai.realtime.RealtimeModel(),
        vad=silero.VAD.load()
    )

    await session.start(
        room=ctx.room,
        agent=Assistant()
    )

    await session.generate_reply()

if __name__ == "__main__":
    agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint))
2. Using MultimodalAgent (Advanced Approach):
Copy code
from livekit import rtc
from livekit.agents import JobContext, AutoSubscribe, llm
from livekit.agents.multimodal import MultimodalAgent
from livekit.plugins import openai

async def entrypoint(ctx: JobContext):
    await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
    participant = await ctx.wait_for_participant()

    model = openai.realtime.RealtimeModel(
        instructions=(
            "You are a voice assistant created by LiveKit. Your interface with users will be voice. "
            "You should use short and concise responses, avoiding usage of unpronounceable punctuation."
        ),
        modalities=["audio", "text"],
    )

    chat_ctx = llm.ChatContext()
    chat_ctx.append(
        text="Context about the user: you are talking to a software engineer who's building voice AI applications.",
        role="assistant",
    )

    agent = MultimodalAgent(
        model=model,
        chat_ctx=chat_ctx,
    )
    agent.start(ctx.room, participant)
    agent.generate_reply()
Key Components: • RealtimeModel: Integrates with OpenAI's Realtime API. • AgentSession: Manages the voice interaction session. • MultimodalAgent: Enables complex multimodal interactions. • VAD (Voice Activity Detection): Often uses Silero for detecting speech. Choose the approach that best fits your requirements. The
AgentSession
is simpler, while the
MultimodalAgent
offers advanced multimodal capabilities. Sources: realtime/openai-realtime-pitch-shift.py | agent.py | agent/main.py | Agent speech | LiveKit Docs | livekit-agents/README.md
c
Hello
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!