We're looking at integrating a flow that uses Live...
# ask-ai
b
We're looking at integrating a flow that uses Livekit as part of an agent flow, but that past an initial interaction does not involve the user speaking or messaging. Are there any APIs, docs or issues we should look at for how to implement a passive mode chat like this?
t
Thanks for reaching out, Tristan. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To implement a passive mode chat in LiveKit where the user does not speak or message after an initial interaction, you can use LiveKit's built-in inactivity detection and handling capabilities. Here’s an overview of the implementation: Core Implementation Approach The primary method is to use the
AgentSession
with inactivity detection. This approach automatically detects when users become inactive and handles the situation gracefully. 1. Session Setup with Inactivity Timeout Configure your
AgentSession
with a
user_away_timeout
to detect when users become passive:
Copy code
session = AgentSession(
    vad=silero.VAD.load(),
    llm=openai.LLM(model="gpt-4o-mini"),
    stt=deepgram.STT(),
    tts=cartesia.TTS(),
    user_away_timeout=12.5,  # Timeout in seconds
)
2. Inactivity Handling Logic Implement a task that manages the passive user scenario:
Copy code
async def user_presence_task():
    # try to ping the user 3 times, if we get no answer, close the session
    for _ in range(3):
        await session.generate_reply(
            instructions=(
                "The user has been inactive. Politely check if the user is still present."
            )
        )
        await asyncio.sleep(10)

    await asyncio.shield(session.aclose())
    ctx.delete_room()
3. Event-Based State Monitoring Set up event handling to monitor user state changes and trigger appropriate responses:
Copy code
@session.on("user_state_changed")
def _user_state_changed(ev: UserStateChangedEvent):
    nonlocal inactivity_task
    if ev.new_state == "away":
        inactivity_task = asyncio.create_task(user_presence_task())
        return

    # ev.new_state: listening, speaking, ..
    if inactivity_task is not None:
        inactivity_task.cancel()
Alternative Approaches Manual Audio Control For more granular control, you can disable audio input after the initial interaction using the push-to-talk pattern:
Copy code
# disable input audio at the start
session.input.set_audio_enabled(False)
Silence-Based Monitoring You can also implement custom silence detection:
Copy code
async def monitor_interaction():
    while True:
        if time.time() - user_last_spoke_time > SILENCE_THRESHOLD:
            <http://logger.info|logger.info>("silent for too long! disconnecting")
            await ctx.room.disconnect()
        await asyncio.sleep(1)
This approach ensures proper resource cleanup and user experience management in scenarios where users don’t interact after the initial conversation.
b
In this case the user will be silent through the entire interaction, the agent will be speaking the entire time. Does that change anything in the suggestions? Any suggestions for how to handle large blocks of text from the agent like this?
t
Hey there! Let's talk about what happens when a user stays quiet during an interaction and the agent keeps on chatting. I've got some friendly tips to help you manage these situations smoothly. 1. Silence Detection and Monitoring If you know the user will be silent, you can skip checking for their input. But if you want to keep an eye out for any unexpected interruptions, here's a neat way to do that:
Copy code
async def monitor_interaction():
    while True:
        if time.time() - user_last_spoke_time > SILENCE_THRESHOLD:
            <http://logger.info|logger.info>("silent for too long! disconnecting")
            await ctx.room.disconnect()
        await asyncio.sleep(1)
2. Non-Interruptible Agent Want to make sure your agent can speak without being cut off? Try setting
allow_interruptions=False
when configuring your agent:
Copy code
class UninterruptableAgent(Agent):
    def __init__(self):
        super().__init__(
            instructions="""
                You are a helpful assistant communicating through voice who will not allow interruptions.
            """,
            stt=deepgram.STT(),
            llm=openai.LLM(model="gpt-4o"),
            tts=openai.TTS(),
            allow_interruptions=False
        )
This way, your agent can finish what they're saying without being interrupted. 3. Managing Large Blocks of Text Here are some cool tricks for handling longer messages: • Text Synchronization: Use synchronized transcripts to keep your text playback smooth:
Copy code
if playback_ev.synchronized_transcript is not None:
    forwarded_text = playback_ev.synchronized_transcript
• Speech Queue Management: Queue up and manage speech for continuous playback:
Copy code
while (true) {
    await this.#speechQueueOpen.await;
    for await (const speech of this.#speechQueue) {
        if (speech === VoicePipelineAgent.FLUSH_SENTINEL) break;
        this.#playingSpeech = speech;
        await this.#playSpeech(speech);
        this.#playingSpeech = undefined;
    }
}
4. User State Management Even when the user is quiet, you can track their state and adapt. For example, you can switch the agent back to listening mode after speaking:
Copy code
if self._session.agent_state == "speaking":
    self._session._update_agent_state("listening")
These tips will help your agent handle long conversations like a pro, even when things get a bit quiet. Hope this helps! Sources: Agents Core Architecture | livekit/livekit_composite | DeepWiki | AI Agents Framework | livekit/livekit_composite | DeepWiki | SIP Integration | livekit/livekit_composite | DeepWiki | basics/interrupts_user.py | livekit-agents/livekit/agents/voice/agent_activity.py
b
We've been asking ourselves if this is actually a valid case for LiveKit. Essentially we're building a listening version of our interactive session. We're planning to collect some user inputs before the agent session starts, building a series of prompts based on those inputs and then running them through a TTS through LiveKit. We have a couple of
agent <-> client
tools we'd trigger through the process but this means we only really have one turn in the agent. Does this seem like a valid / viable use case or is it a step too far friction-wise @refined-appointment-81829?
t
Hey there! I totally get what you're looking for, and I'm excited to share how LiveKit can make your project awesome. Let me break down how this platform is basically perfect for your needs. 1. Pre-session Input Collection and Storage Think of LiveKit's
AgentSession
like a friendly assistant who remembers all the details you've shared. The
userdata
parameter is like a little notebook where you can jot down all those personal inputs, making sure your session feels super personalized. 2. Dynamic Prompt Generation This is where the magic happens! You can create agent instructions that feel like they're speaking directly to the user. Check out this cool example:
Copy code
# Example of dynamic instruction generation
story_agent = StoryAgent(name, location)
# Where instructions include: f"The user's name is {name}, from {location}"
It's like the agent is giving a personalized shoutout to each user – how cool is that? 3. TTS Integration LiveKit's text-to-speech is seriously impressive. We're talking high-quality, super smooth voices that sound natural in multiple languages. It's like having a professional voice actor on standby! 4. One-Turn Agent Interactions Here's where LiveKit really shines for your use case. You've got some awesome tools at your fingertips: •
session.say()
lets your agent speak up with predefined messages. •
session.generate_reply()
is perfect for creating dynamic, on-the-fly responses. While the agent usually waits for user input, it can totally start the conversation too. Super flexible! Conclusion I'm really excited for you! This workflow is going to be smooth as butter with LiveKit. You'll have personalized interactions, great voice synthesis, and a super adaptable platform that does exactly what you need. Trust me, you're on the right track! Sources: tests/test_agent_session.py | Agent speech | LiveKit Docs | Text-to-speech (TTS) integrations | LiveKit Docs | Agent speech and audio | LiveKit Docs