Using the JS Client SDK, how can I detect that the...
# ask-ai
t
Using the JS Client SDK, how can I detect that the remote agent is done speaking? Listening for the remote participant’s IsSpeakingChanged event isn’t good because it goes to false in between sentences.
t
Thanks for reaching out, Christian. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To gently detect when a remote participant has finished speaking using the LiveKit JS Client SDK, a great way is to use the
ParticipantEvent.IsSpeakingChanged
event. I understand that this event might toggle between sentences, so adding a little extra logic to handle those natural pauses can really help. Here are a couple of friendly approaches you might find useful: • Listen for the
IsSpeakingChanged
event and introduce a small delay before deciding the participant is "done speaking." For example:
Copy code
participant.on(ParticipantEvent.IsSpeakingChanged, (speaking) => {
  if (!speaking) {
    setTimeout(() => {
      if (!participant.isSpeaking) {
        console.log(`${participant.identity} is done speaking.`);
      }
    }, 1000); // You can adjust this delay to what feels right
  }
});
• Alternatively, you can use
createIsSpeakingObserver
from
@livekit/components-core
to watch the speaking status as a reactive stream, applying similar thoughtful logic. These approaches help you gracefully handle brief pauses between sentences, making sure you catch when someone truly finishes speaking. If you have any questions or want to explore this further, I’m here to help! Sources: Subscribing to tracks | LiveKit Docs | src/room/participant/Participant.ts | packages/core/src/observables/participant.ts | AI voice agents | LiveKit Docs
t
Yes, I had already tried the setTimeout approach and found that I indeed need to wait for a full second for it to work reliably. This is too long for my use case. Is there some other signaling from the agent that it is done with their answer?
t
Thank you for reaching out! It looks like there isn’t a built-in signaling mechanism specifically for detecting when a remote participant has finished speaking beyond the
IsSpeakingChanged
event. To better assist you, could you please share a bit more about your use case? For example: • Are you trying to signal the end of a structured response, like a predefined answer? • Is the remote agent using any particular tools or workflows that might give us some extra context? Once I have a clearer picture, I’d be happy to help explore some potential workarounds or alternative approaches with you.
t
My agent is implemented in Python using
livekit.agents
Is it possible in the agent to detect that the tts has finished the response? Perhaps with a session event handler? If so, the agent could inform my client with an RPC call.
t
I'm glad you're exploring how to detect when TTS has finished a response with your Python agent using
livekit.agents
! You can achieve this by listening to the
playback_finished
event from the
AudioOutput
component in the agent session. This event fires whenever the audio playback ends, whether it finishes naturally or is interrupted. Here’s a simple way to set this up:
Copy code
from <http://livekit.agents.voice.io|livekit.agents.voice.io> import PlaybackFinishedEvent

@session.output.audio.on("playback_finished")
def _on_playback_finished(ev: PlaybackFinishedEvent):
    <http://logger.info|logger.info>(f"Playback finished: position={ev.playback_position}, interrupted={ev.interrupted}")
    # Notify your client via RPC or perform other actions
Also, it might be helpful to keep an eye on the
agent_state_changed
event. This lets you track when the agent switches from "speaking" to "listening" or "thinking," which can give you more insight into its behavior:
Copy code
@session.on("agent_state_changed")
def _on_agent_state_changed(event):
    if event.new_state == "listening":
        <http://logger.info|logger.info>("Agent is now listening")
To keep your client in the loop, you can implement an RPC call within the agent that broadcasts the
playback_finished
event to the room. Then, your web client can listen for this custom RPC event to update the UI or trigger other actions smoothly. If you’d like, I’m here to help you with any part of this implementation—just let me know! Sources: How to detect when an agent has finished speaking | livekit-agents/livekit/agents/voice/events.py | Bringing AI avatars to voice agents
t
Registering for this event at the end of the entrypoint function in my agent, after the session started, would be a good spot, right?
t
You're on the right track by registering for the
playback_finished
event at the end of the entrypoint function, once the session has started. This is a great way to make sure the session and its
AudioOutput
component are fully set up before you attach the event listener. Just a little reminder to double-check that
session.output.audio
is properly initialized before you register the event handler. Here's a simple example to guide you:
Copy code
# Ensure session.output.audio is initialized
session.output.audio = DataStreamAudioOutput(
    ctx.room,
    destination_identity=AVATAR_IDENTITY,
)

# Start the session
await session.start(
    agent=agent,
    room=ctx.room,
    room_output_options=RoomOutputOptions(
        audio_enabled=True,
        transcription_enabled=True,
    ),
)

# Register the event handler
@session.output.audio.on("playback_finished")
def on_playback_finished(ev: PlaybackFinishedEvent):
    <http://logger.info|logger.info>("Audio playback completed")
By following this approach, you can feel confident that the event listener is attached right after everything is ready. If you have any questions or want to double-check anything else, I’m here to help!
t
Please show how to do the RPC call from the agent
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!