How do you do function calling with speech to spee...
# ask-ai
j
How do you do function calling with speech to speech models
t
Thanks for reaching out, David. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Function calling with speech-to-speech models in LiveKit is supported through the
MultimodalAgent
and
PipelineVoiceAgent
. These agents use OpenAI's multimodal model and real-time API to process user audio and generate audio responses while interacting with external functions or tools. Here’s an overview of how it works: • Function Calling Basics: Function calling (also known as "tool calling") allows AI models to interact with external functions to retrieve additional context or take real-world actions before generating a response. • Implementation in Python: You can define functions using the
@function_tool
decorator within your agent class. For example:
Copy code
from livekit.agents.llm import function_tool

class MyAgent(Agent):
    def __init__(self):
        super().__init__(instructions="Your agent instructions here")

    @function_tool()
    async def your_function_name(self):
        """Function description that the LLM will use."""
        # Your function logic here
        return "Function result"
• Key Features: ◦ Function calls are automatically detected by the LLM during conversation. ◦ Functions can be executed while maintaining the speech-to-speech flow. ◦ The agent can use function results to inform its audio responses. ◦ Both synchronous and asynchronous function execution is supported. For JavaScript implementations, function call events can be handled and executed within the speech-to-speech context. If you need further details or specific examples, let me know! Sources: AI voice agents | LiveKit Docs | Function calling with Voice Agents | LiveKit Docs | tool_calling/call_function_tool.py | agents/src/pipeline/pipeline_agent.ts