We want the agent to speak step by step. For each ...
# ask-ai
r
We want the agent to speak step by step. For each step it has to first do function call and then speak about that step. Now giving it once in a prompt won't be ideal since it will do all the function calls quickly and not wait for each step speech to finish before moving on to next step. So how to do this? We are using gpt-realtime to do this.
t
Thanks for reaching out, Utsav. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To help you make GPT-realtime perform step-by-step function calls with speech output, you can use the Multimodal Agent API within LiveKit. This is a wonderful way to enable real-time voice interactions where each step involves a function call followed by speech output. Here’s a friendly guide on how you can achieve this: • Use the MultimodalAgent Class: This combines OpenAI’s Realtime API with LiveKit, allowing both audio and text input/output. Audio is streamed via WebRTC to GPT-4o, which then processes the input and makes function calls. • Handle Function Calls in Realtime: Function calls are handled asynchronously, and their results are added back into the chat context. After each function runs, speech generation is triggered. • Generate Speech After Execution: The system creates a speech task based on the function result, streaming it back to you through LiveKit’s WebRTC infrastructure. • Ensure Real-Time Synchronization: LiveKit keeps everything synchronized with GPT-4o, even if interruptions happen, so your interactions stay smooth. You can get started by exploring the LiveKit Playground or following the LiveKit + OpenAI guide for detailed setup steps. If you’d like, I’m here to help you with any specific configurations or questions about the tools—just let me know! Sources: OpenAI and LiveKit partner to turn Advanced Voice into an API | agents/src/voice/generation.ts
r
how to step by step tool call speech then tool call after speech end
t
I'm glad you're exploring how to manage step-by-step tool calls with speech output! LiveKit actually has built-in features that handle this smoothly by default. It waits for the speech to finish before moving on to the next tool call, which helps keep the conversation feeling natural and easy. If you'd like more hands-on control, you can use the
RunContext.wait_for_playout()
method inside your tool function. This makes sure your tool logic pauses until the speech finishes playing. Here's a simple example to guide you:
Copy code
@function_tool
async def my_tool(self, run_ctx: RunContext) -> str:
    # Wait for the speech tied to this function call to finish
    await run_ctx.wait_for_playout()
    
    # Tool logic executes only after speech playout completes
    return "Tool completed after speech"
This approach helps avoid any circular waiting and keeps everything running smoothly. If you have any questions or want to dive deeper into the implementation, just let me know—I’m here to help! Sources: Bringing AI avatars to voice agents | Bringing AI avatars to voice agents | Bringing AI avatars to voice agents | examples/voice_agents/long_running_function.py | agents/src/voice/agent_activity.ts
r
Give me source or docs for wait_for_playout()
t
The
RunContext.wait_for_playout()
method is a helpful part of the LiveKit Agents framework that lets us patiently wait for the assistant's spoken response (just for the current function call step) to finish before moving forward. Here’s the source code implementation for you to take a look:
Copy code
async def wait_for_playout(self) -> None:
    """Waits for the speech playout corresponding to this function call step.

    Unlike `SpeechHandle.wait_for_playout`, which waits for the full
    assistant turn to complete (including all function tools),
    this method only waits for the assistant's spoken response prior running
    this tool to finish playing."""
    await self.speech_handle._wait_for_generation(step_idx=self._initial_step_idx)
For example, if you have a tool that ends a call, you could use it like this:
Copy code
@function_tool
async def end_call(self, ctx: RunContext):
    """Use this tool when the user has signaled they wish to end the current call."""
    await ctx.wait_for_playout()  # let the agent finish speaking
    await hangup_call()  # then end the call
This way, you make sure the agent completes its speech before any other actions take place. If you'd like, I'm here to help with more details or any other questions you might have! Sources: livekit-agents/livekit/agents/voice/agent_activity.py | Agent speech and audio | LiveKit Docs | Agents telephony integration | LiveKit Docs