Are there any working multi-avatar examples?
# ask-ai
l
Are there any working multi-avatar examples?
t
Thanks for reaching out, Derek. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I'm glad you asked! LiveKit offers some fantastic multi-avatar examples through its Agents Framework, working seamlessly with avatar providers like Anam, Beyond Presence (Bey), BitHuman, Hedra, Simli, and Tavus. These examples let AI voice agents showcase synchronized animated avatars in real time, which is pretty exciting. You can explore these examples in the LiveKit Agents GitHub repository, specifically in the
examples/avatar_agents/
directory. Each one comes with clear setup instructions and helpful code samples to guide you through the integration process. If you'd like to dive deeper, the Virtual Avatar Integrations Guide is a wonderful resource to check out as well. Sources: examples/avatar_agents/bey/README.md | Virtual avatar integrations | LiveKit Docs | avatars/hedra/education_avatar/education-frontend/components/livekit/media-...
l
Can I have multiple agent sessions with virtual avatars in a single room?
t
Absolutely, you can have multiple agent sessions with virtual avatars all within a single LiveKit room. Each agent session can have its own virtual avatar, acting as a separate participant (avatar worker) in the room. This way, multiple agents and avatars can interact together smoothly in the same space. To make this happen, you’d set up multiple
AgentSession
instances, each connected to a different
AvatarSession
. These avatars will then publish their own synchronized audio and video tracks independently to the room. This setup is especially handy for scenarios like warm transfers or when multiple agents need to collaborate. If you’d like, I’m here to help guide you through the setup process—just let me know! Sources: Workflows | LiveKit Docs | README.md | Anatomy of an Agent | LiveKit Docs | Anatomy of an Agent | LiveKit Docs | Virtual avatar integrations | LiveKit Docs
l
Can I load both sessions up with virtual avatars?
t
Absolutely! You can definitely load both agent sessions in a LiveKit room, each with their own unique virtual avatars. Each
AgentSession
can be connected to its own
AvatarSession
, which means every agent gets their own avatar worker that shares synchronized audio and video tracks in the room. This way, multiple agents with different avatars can comfortably coexist and interact within the same space. If you'd like, I’m here to help you set this up—just let me know! Sources: Virtual avatar integrations | LiveKit Docs | examples/avatar_agents/bey/README.md
l
Give me an example
t
I'd love to help you out! Could you please share a bit more about the kind of example you're looking for? For example: • A code snippet that shows how to configure multiple
AgentSession
and
AvatarSession
instances? • A friendly step-by-step guide to set up multiple avatars in a room? • Or maybe some details on how the frontend should handle multiple avatars? Just let me know what works best for you, and I'll make sure to provide the most helpful example!
l
Can only one agent session be active at the same time?
t
Absolutely! You can have multiple agent sessions running at the same time in LiveKit. Each session operates in its own separate process, and a single worker is capable of handling several sessions simultaneously. The total number of sessions you can run at once depends on your system or hosting plan limits, but these can be adjusted to better fit what you need. If you have any specific questions or would like a hand setting this up, please don’t hesitate to reach out—I’m here to help! Sources: README.md | Anatomy of an Agent | LiveKit Docs | Worker lifecycle | LiveKit Docs | Inside a session | LiveKit Docs | LiveKit Pricing
l
Can I generate reply for both agents at the same time?
t
Thanks so much for reaching out! I wanted to share that LiveKit doesn’t currently support generating replies from multiple agents at the exact same time, which means their speech won’t overlap or play concurrently. Instead, the design encourages agents to take turns, so only one agent speaks at a time, helping keep the conversation clear and smooth. If you’re interested in creating a custom setup where, for example, one agent offers background commentary while another leads the chat, just let me know! I’d be happy to walk you through some ideas and help you get that going. Sources: examples/voice_agents/silent_function_call.py | Dispatching agents | LiveKit Docs | Agent speech and audio | LiveKit Docs
l
Depending on my users input I want one agent or the other to reply how do I set this up.
t
To help you set up conditional replies based on user input, you can use the
session.generate_reply()
method along with some conditional logic. Here’s a friendly guide on how to do it: 1. Keep track of the user’s state using
self.session.userdata
with a custom data class that holds all the important info. 2. Use conditional logic inside methods like
on_enter()
or
@function_tool
functions to decide how the agent should respond, depending on the stored state or what the user says. 3. Call
generate_reply()
with parameters such as
user_input
or
instructions
to tailor the response dynamically. For instance, you might have one agent ask a question and then, based on the user’s answer, either continue with the same agent or switch to another one. If you'd like, I’d be happy to share a specific code example to make this clearer—just let me know! Sources: examples/src/basic_tool_call_agent.ts | examples/drive-thru/test_agent.py | Agent speech and audio | LiveKit Docs | Agent speech and audio | LiveKit Docs | Workflows | LiveKit Docs
l
Give me an example
t
Here’s a helpful example of setting up conditional replies where one agent or another responds based on what the user chooses:
Copy code
from dataclasses import dataclass
from livekit.agents import Agent, AgentSession

@dataclass
class UserData:
    preference: str = ""

class AgentA(Agent):
    async def on_enter(self) -> None:
        await self.session.say("You selected Agent A. How can I assist you?")

class AgentB(Agent):
    async def on_enter(self) -> None:
        await self.session.say("You selected Agent B. What do you need help with?")

class MainAgent(Agent):
    async def on_enter(self) -> None:
        await self.session.say("Do you prefer Agent A or Agent B?")

    @function_tool
    async def handle_preference(self, choice: str) -> Agent:
        userdata: UserData = self.session.userdata
        userdata.preference = choice
        if choice.lower() == "a":
            return AgentA()
        elif choice.lower() == "b":
            return AgentB()
        else:
            await self.session.say("Invalid choice. Please select A or B.")
            return self

# Initialize the session
session = AgentSession[UserData](userdata=UserData())
main_agent = MainAgent()
session.run(main_agent)
To break it down: • The
MainAgent
kindly asks the user to pick between Agent A or Agent B. • Depending on what the user chooses, the session smoothly switches to either
AgentA
or
AgentB
. • We also save the user’s preference in
userdata
, so it’s easy to refer back to later. If you’d like, I’m here to walk you through any part of this or provide more examples—just let me know!
l
Can the main agent just be a silent router?
t
Absolutely! A main agent can definitely act as a silent router in LiveKit by making routing decisions quietly, without generating any audible responses. Here’s how you can do it: 1. Use function tools that don’t return any value, so no verbal reply is triggered. 2. Or, explicitly cancel tool replies by using
ev.cancel_tool_reply()
within the
function_tools_executed
event handler. For instance, you might define a routing function like this:
Copy code
@function_tool()
async def route_to_support_agent(self):
    return llm.handoff({
        "agent": SupportAgent(),
        "returns": "Routing completed"
    })

@session.on("function_tools_executed")
def on_function_tools_executed(ev):
    if "route_to_support_agent" in [fnc.name for fnc in ev.function_calls]:
        ev.cancel_tool_reply()
This way, the routing happens smoothly and silently, with no audible response. If you have any questions or need a hand with this, please don’t hesitate to ask—I’m here to help! Sources: examples/voice_agents/README.md | examples/src/comprehensive_test.ts
l
Can the main agent route via regex based on the stt
t
Thanks for reaching out! Implementing regex-based routing with STT (speech-to-text) usually involves working with the transcribed text once the user has finished speaking. To make sure I can guide you in the best way possible, could you please share a bit more? • Are you aiming to route based on the final transcription after the user finishes, or would you like to work with real-time interim transcription? • Are you using the Python or Node.js SDK for this? • Is there a particular regex pattern or routing logic you'd like some help with? Just let me know, and I’ll be happy to provide an example that fits your needs perfectly! Sources: agents/src/voice/audio_recognition.ts | livekit-plugins/livekit-plugins-speechmatics/livekit/plugins/speechmatics/s... | Working with plugins | LiveKit Docs
l
I am using python I want to work with the transcribed text once the user has finished speaking
t
Here’s a friendly example to help you implement regex-based routing in Python, using the transcribed text after the user finishes speaking:
Copy code
import re
from livekit.agents import Agent, AgentSession

class MainAgent(Agent):
    async def on_user_turn_completed(self, chat_context, new_message):
        # Access the transcribed text
        transcript = new_message.content

        # Apply regex to determine routing
        if re.match(r"^route to support", transcript, re.IGNORECASE):
            await self.session.update_agent(SupportAgent())
        elif re.match(r"^route to sales", transcript, re.IGNORECASE):
            await self.session.update_agent(SalesAgent())
        else:
            await self.session.say("I didn’t understand your request. Could you clarify?")

class SupportAgent(Agent):
    async def on_enter(self):
        await self.session.say("You’ve been routed to Support. How can I assist you?")

class SalesAgent(Agent):
    async def on_enter(self):
        await self.session.say("You’ve been routed to Sales. What can I help you with?")

# Initialize the session
session = AgentSession()
main_agent = MainAgent()
session.run(main_agent)
Here’s a quick overview: • The
on_user_turn_completed
method handles the final transcription once the user finishes speaking. • Regex patterns help match specific phrases like "route to support" or "route to sales". • Depending on the match, the session smoothly transitions to the right agent, such as
SupportAgent
or
SalesAgent
. • If nothing matches, the main agent kindly asks the user to clarify their request. Please don’t hesitate to reach out if you’d like any more help or if there’s anything specific you want to customize further—I’m here to support you!
l
How does the virtual avatars tie into this? I want the support and sales agent to have their own unique TTS model
t
I'm happy to guide you through assigning unique TTS models and virtual avatars to your Support and Sales agents. You can easily configure each agent's
AgentSession
with the specific TTS and avatar settings that suit their roles. Let’s walk through the steps together: 1. Assigning Unique TTS Models For instance, you might choose OpenAI TTS for your Support team and Cartesia TTS for Sales. Here's a simple example:
Copy code
from livekit.plugins import openai, cartesia

# Support Agent TTS
support_tts = openai.TTS(
    model="gpt-4o-mini-tts",
    voice="ash",
    instructions="Speak in a friendly tone."
)

# Sales Agent TTS
sales_tts = cartesia.TTS(
    model="sonic-2",
    voice="f786b574-daa5-4673-aa0c-cbe3e8534c02"
)
2. Assigning Virtual Avatars Similarly, you can assign avatars like Tavus for Support and bitHuman for Sales. Here's how you might set that up:
Copy code
from livekit.plugins import tavus, bithuman

# Support Agent Avatar
support_avatar = tavus.AvatarSession(
    replica_id="r4c41453d2",
    persona_id="p2fbd605"
)

# Sales Agent Avatar
sales_avatar = bithuman.AvatarSession(
    avatar_id="your-avatar-id"
)
3. Integrating with Agent Sessions Next, you'll want to attach these TTS and avatar configurations to each agent’s session. Here’s an example:
Copy code
from livekit.agents import AgentSession

# Support Agent Session
support_session = AgentSession(tts=support_tts)
await support_avatar.start(support_session, room=ctx.room)

# Sales Agent Session
sales_session = AgentSession(tts=sales_tts)
await sales_avatar.start(sales_session, room=ctx.room)
4. Frontend Rendering The avatars will publish video tracks to the room, which you can render using LiveKit's client SDKs. If you’d like, I’m here to help you with the frontend setup or answer any questions you have along the way! Sources: Bringing AI avatars to voice agents | bitHuman virtual avatar integration guide | LiveKit Docs | LiveKit Agents integrations | LiveKit Docs | OpenAI TTS integration guide | LiveKit Docs | Groq TTS integration guide | LiveKit Docs
l
Give me the complete code example
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!
l
Can i switch between sessions if I have multiple sessions?