jolly-autumn-53818
07/28/2025, 12:13 AMtall-airline-16756
07/28/2025, 12:13 AMtall-airline-16756
07/28/2025, 12:14 AMMultimodalAgent and PipelineVoiceAgent. These agents use OpenAI's multimodal model and real-time API to process user audio and generate audio responses while interacting with external functions or tools.
Here’s an overview of how it works:
• Function Calling Basics: Function calling (also known as "tool calling") allows AI models to interact with external functions to retrieve additional context or take real-world actions before generating a response.
• Implementation in Python: You can define functions using the @function_tool decorator within your agent class. For example:
from livekit.agents.llm import function_tool
class MyAgent(Agent):
def __init__(self):
super().__init__(instructions="Your agent instructions here")
@function_tool()
async def your_function_name(self):
"""Function description that the LLM will use."""
# Your function logic here
return "Function result"
• Key Features:
◦ Function calls are automatically detected by the LLM during conversation.
◦ Functions can be executed while maintaining the speech-to-speech flow.
◦ The agent can use function results to inform its audio responses.
◦ Both synchronous and asynchronous function execution is supported.
For JavaScript implementations, function call events can be handled and executed within the speech-to-speech context.
If you need further details or specific examples, let me know!
Sources: AI voice agents | LiveKit Docs | Function calling with Voice Agents | LiveKit Docs | tool_calling/call_function_tool.py | agents/src/pipeline/pipeline_agent.ts