boundless-truck-87206
07/24/2025, 8:37 AMtall-airline-16756
07/24/2025, 8:37 AMtall-airline-16756
07/24/2025, 8:38 AMAgentSession with manual turn detection and disable audio input initially:
1. Push-to-Talk Implementation:`# When user starts speaking (button press)
@ctx.room.local_participant.register_rpc_method("start_turn")
async def start_turn(data: rtc.RpcInvocationData):
session.interrupt() # Stop any current agent speech
session.clear_user_turn() # Clear any previous input
session.input.set_audio_enabled(True) # Start listening
# When user finishes speaking (button release)
@ctx.room.local_participant.register_rpc_method("end_turn")
async def end_turn(data: rtc.RpcInvocationData):
session.input.set_audio_enabled(False) # Stop listening
session.commit_user_turn(
transcript_timeout=10.0, # Timeout for final transcript
)
# When user cancels their turn
@ctx.room.local_participant.register_rpc_method("cancel_turn")
async def cancel_turn(data: rtc.RpcInvocationData):
session.input.set_audio_enabled(False) # Stop listening
session.clear_user_turn() # Discard the input`
◦ Use RPC methods to control the audio input lifecycle:
1. Audio Processing as a Whole:
◦ The commit_user_turn() method ensures that the entire audio segment is processed by the STT after the button is released. It waits for the final transcript and processes the input as a complete segment.
1. Complete Example:`async def entrypoint(ctx: JobContext):
session = AgentSession(turn_detection="manual")
room_io = RoomIO(session, room=ctx.room)
await room_io.start()
agent = MyAgent() # Your agent implementation
await session.start(agent=agent)
session.input.set_audio_enabled(False) # Disable input audio at the start
@ctx.room.local_participant.register_rpc_method("start_turn")
async def start_turn(data: rtc.RpcInvocationData):
session.interrupt()
session.clear_user_turn()
room_io.set_participant(data.caller_identity) # Listen to the caller
session.input.set_audio_enabled(True)
@ctx.room.local_participant.register_rpc_method("end_turn")
async def end_turn(data: rtc.RpcInvocationData):
session.input.set_audio_enabled(False)
session.commit_user_turn(transcript_timeout=10.0)`
◦ Here’s a full implementation:
This setup ensures that the STT processes the complete audio segment captured between button press and release, providing accurate transcription of the entire user utterance before generating a response.
Sources: Turn detection and interruptions | LiveKit Docs | livekit-plugins/livekit-plugins-turn-detector/README.mdboundless-truck-87206
07/24/2025, 8:44 AMtall-airline-16756
07/24/2025, 8:47 AM