how to manually invoke a tool call?
# ask-ai
w
how to manually invoke a tool call?
t
Thanks for reaching out, Test. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Thanks for reaching out! Could you please help me understand a bit better by clarifying which tool or feature you're referring to when you mention manually invoking a tool call? For instance, are you asking about a specific API, SDK, or another LiveKit functionality? I'm here to help you every step of the way!
w
i mean in workflow i have an agent and i want as soon as that agent enters a room a tool call should be invoked
t
Hey there! If you want to have a tool call triggered right when an agent enters a room, you can use the
onEnter()
lifecycle hook in your custom agent class. This handy hook fires as soon as the agent becomes active in the session, so you can run your tools right away. Direct Tool Invocation: • First, define your tool using
llm.tool()
for Node.js or
@function_tool
if you’re using Python. • Then, inside the
onEnter()
method, you can call the tool’s
execute()
method directly with the parameters it needs. Example (Node.js):
Copy code
class CustomAgent extends voice.Agent {
  async onEnter() {
    const result = await myTool.execute({ param: value }, { ctx: this.ctx });
    this.session.say(result);
  }
}
Alternative: LLM-Triggered Tool Call • Another way is to use
generateReply()
inside
onEnter()
, letting the LLM decide and call a tool based on your instructions. If you’d like, I’m here to help you walk through the implementation or answer any questions you have—just let me know! Sources: Authentication | LiveKit Docs | examples/src/multi_agent.ts | tool_calling/call_function_tool.py | Workflows | LiveKit Docs | examples/agent-dispatch/index.ts
w
but generateReply is saying a reply the a similar reply is said by agent is as well
t
When you use
generateReply
, the framework kindly takes care of cancelling any ongoing generation if a new
generateReply
call comes in. This way, it makes sure only the most recent request is processed, helping to avoid any duplicate or conflicting responses. That said, if you’re still seeing duplicate replies, it might be because multiple
generateReply
calls are happening at the same time in your agent logic. To help with this, you might want to: • Double-check that
generateReply
isn’t being called from several asynchronous flows simultaneously. • Use the returned
SpeechHandle
object to keep track of and manage the speech lifecycle, which helps keep everything in sync and prevents overlapping actions. If you continue to experience this issue, please don’t hesitate to reach out—I’m here to help you work through it! Sources: livekit-agents/livekit/agents/voice/agent_activity.py | README.md | Agent speech and audio | LiveKit Docs | Agent speech and audio | LiveKit Docs
w
i have called single generate reply still it says two times
t
Sometimes, duplicate responses from
generateReply
can happen because of overlapping triggers in the agent's workflow. No worries though—here are a few friendly steps to help you sort this out: • Disable Server-Side Turn Detection: If you have server-side turn detection turned on, it might be causing automatic responses. Try setting
turn_detection=None
in your RealtimeModel configuration to stop this from happening. • Avoid Manual Conflicts: Double-check that
generateReply
isn’t being called inside
onUserTurnCompleted
or similar places where it might already be triggered automatically. • Track Speech State: Use the returned
SpeechHandle
to keep an eye on and cancel any overlapping speech actions. • Log and Monitor: Turning on logging can really help you see when and why
generateReply
gets triggered more than once. If you’d like, I’m here to help you implement these changes or troubleshoot further—just let me know! Sources: tests/test_evals.py | agents/src/voice/agent_activity.ts | Voice activity detection (VAD) parameters | LiveKit Docs | Agent speech and audio | LiveKit Docs
w
class CustomAgent extends voice.Agent { async onEnter() { const result = await myTool.execute({ param: value }, { ctx: this.ctx }); this.session.say(result); } } in python
t
I'd love to help you implement your custom agent with
onEnter
and tool execution in Python! Here's a straightforward example to get you started: Example:
Copy code
from livekit.agents import Agent, function_tool
from livekit.agents.voice import AgentSession

class MyCustomAgent(Agent):
    async def on_enter(self):
        # This is called when the agent becomes active
        await self.session.say("Hello! I am now active. How can I assist you?")

    @function_tool
    async def my_tool(self, param: str) -> str:
        # Define your tool logic here
        return f"You passed: {param}"
Steps: • Override the
on_enter
method to define actions when the agent enters a room. • Use the
@function_tool
decorator to define tools that can be executed by the agent. • Call
self.session.say()
to send a message or response to the user. If you have any questions or need more guidance as you work through this, please don't hesitate to reach out. I'm here to help! Sources: main.py | tool_calling/call_function_tool.py | examples/src/gemini_realtime_agent.ts | Pipeline nodes and hooks | LiveKit Docs | examples/voice_agents/basic_agent.py
w
invoke tool call manually
t
To manually invoke a tool call in Python, you can simply call the tool's method and pass the parameters it needs. Here's a little example to help you out:
Copy code
class MyCustomAgent(Agent):
    async def on_enter(self):
        # Manually invoke the tool call
        result = await self.my_tool("example parameter")
        await self.session.say(result)

    @function_tool
    async def my_tool(self, param: str) -> str:
        return f"You passed: {param}"
In this example, the tool
my_tool
is called right inside the
on_enter
method, and its result is shared using
self.session.say()
. If you have any questions or want me to walk you through it more, just let me know—I'm here to help!
w
i want llm to execute it
t
If you'd like the LLM to execute the tool, you can use the
generate_reply
method and make sure the tool is registered with the agent. Here's a simple example to guide you:
Copy code
class MyCustomAgent(Agent):
    async def on_enter(self):
        # This triggers the LLM to decide and execute the tool
        await self.session.generate_reply(user_input="Execute the tool with parameter example")

    @function_tool
    async def my_tool(self, param: str) -> str:
        return f"You passed: {param}"
In this setup: • The
generate_reply
method gently prompts the LLM to decide the next step, which can include calling the tool. • The tool
my_tool
is registered with the agent using the
@function_tool
decorator, so the LLM knows it can use it. If you have any questions or need more help, please don’t hesitate to ask—I’m here to support you!
w
but generae reply will produce a speech as well
t
To use
generateReply
without producing speech output, you can simply disable audio output in the session configuration. Here’s a friendly guide to help you with this: 1. Disable Audio Output During Session Setup:
Copy code
# Python Example
session = AgentSession(
    # Other configurations (stt, llm, tts, vad, etc.)
    room_output_options=RoomOutputOptions(audio_enabled=False),
)
Copy code
// Node.js Example
const session = new voice.AgentSession({
    // Other configurations (vad, stt, tts, llm, etc.)
    outputOptions: {
        audioEnabled: false,
    },
});
2. Dynamically Disable Audio Output: If you want to toggle audio output during an active session, you can do it like this:
Copy code
# Python
session.output.set_audio_enabled(False)
Copy code
// Node.js
session.output.setAudioEnabled(false)
Once audio output is disabled, the agent won’t send audio tracks to the room, and instead, text responses will come through the
lk.transcription
text stream topic. If you have any questions or need a hand with this setup, please don’t hesitate to ask—I’m here to help! Sources: Agent speech and audio | LiveKit Docs | Agent speech and audio | LiveKit Docs | Docs::Agents | How to Set Up Speech-to-Text Input with Text-Only Output in LiveKit Agents | livekit-agents/livekit/agents/voice/agent_activity.py
w
i want speech in whole session just i want to call a specific tool manually via llm and also it should not generate any speech reply
t
Hi there! If you'd like to call a specific tool manually via the LLM without generating a speech reply, here’s a friendly guide to help you out: 1. Use a Tool Without a Return Value: You can define your tool so that it performs the action you need but doesn’t return a value. This way, tools without return values won’t trigger any automatic replies.
Copy code
@function_tool
async def my_tool(self, param: str) -> None:
    # Perform the action without returning a value
    print(f"Tool executed with param: {param}")
2. Cancel Tool Replies Programmatically: In case your tool does return a value but you’d prefer to suppress the reply, you can easily cancel it in the
function_tools_executed
event handler like this:
Copy code
def on_function_tools_executed(self, ev):
    ev.cancel_tool_reply()
This way, your tool runs smoothly without generating any speech response. If anything’s unclear or you’d like a hand implementing this, please don’t hesitate to ask—I’m here to help!
w
def on_function_tools_executed(self, ev): ev.cancel_tool_reply() but when i will use generate_reply that will generate the speech
t
I’m happy to help you with using
generate_reply
to invoke a tool via the LLM while keeping the speech output silent. Here are a couple of friendly approaches you can try: 1. Cancel Speech in the Tool Execution Context: You can use the
on_function_tools_executed
event handler to stop the speech reply that the tool might generate:
Copy code
def on_function_tools_executed(self, ev):
    ev.cancel_tool_reply()
2. Suppress Speech Dynamically: Another way is to temporarily turn off audio output just before calling `generate_reply`:
Copy code
# Python Example
self.session.output.set_audio_enabled(False)
await self.session.generate_reply(user_input="Execute the tool with parameter example")
self.session.output.set_audio_enabled(True)
This way, the tool runs through the LLM, but no speech will be produced during this particular interaction. If anything is unclear or you’d like more help, please don’t hesitate to reach out. I’m here for you!
w
but will it add the output of generate_reply to chat ctx?
t
Absolutely! When you use
generate_reply
, the output is added to the chat context. This way, the agent's response becomes part of the conversation history, helping to keep everything connected for future interactions. Feel free to reach out if you'd like me to explain anything further or if you need any more help—I’m here for you! Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime... | Agent speech and audio | LiveKit Docs | Agent speech and audio | LiveKit Docs
w
self.session.output.set_audio_enabled(False) similar like this i want it should not be added to chat ctx for that particular case
t
I understand you're looking for a way to prevent the output of
generate_reply
from being added to the chat context in certain cases. Right now, the function is designed to append its output to the chat context to keep our conversations flowing smoothly. That said, a helpful workaround is to run the tool directly without using
generate_reply
if you'd prefer not to add the output to the chat context. Another option is to adjust the chat context after the reply is generated by programmatically removing or modifying the entry. Please feel free to reach out if you'd like some guidance on how to try these approaches—I’m here to help!
w
That said, a helpful workaround is to run the tool directly without using
generate_reply
how to do this in python and with worflow agents
t
I'd be happy to help you run a tool directly in Python with workflow agents! You can do this by defining your tool with the
@function_tool
decorator and then calling it programmatically within your agent logic. Here's a simple example to guide you:
Copy code
from livekit.agents import Agent, function_tool

class MyWorkflowAgent(Agent):
    @function_tool
    async def my_tool(self, param: str) -> str:
        # Define the tool logic here
        return f"Tool executed with param: {param}"

    async def on_enter(self):
        # Directly invoke the tool without using generate_reply
        result = await self.my_tool("example parameter")
        print(result)  # Optionally log or process the result
Here’s how you can approach it: • Start by defining the tool using
@function_tool
inside your agent class. • Then, call the tool directly with
await self.my_tool(parameters)
in your workflow logic, like within
on_enter
. • Finally, you can handle the result however you need, without involving the LLM or adding it to the chat context. If you have any questions or want me to walk you through any part of this, just let me know—I’m here to help! Sources: Tool definition and use | LiveKit Docs | Tool definition and use | LiveKit Docs | basics/function_calling.py | livekit-agents/livekit/agents/voice/agent_activity.py
w
how to stop generate_reply to add the message in chat ctx
t
I understand you're looking for a way to prevent
generate_reply
from adding its output to the chat context. While there isn't a direct method for this right now, I’m happy to share some helpful workarounds you can try: 1. Use
session.say()
with `add_to_chat_ctx=False`:
This lets the agent send a message without adding it to the chat context, which can be quite handy:
Copy code
await session.say("Hello", add_to_chat_ctx=False)
2. Cancel Replies in `on_user_turn_completed`: You can override the
on_user_turn_completed
method and raise
StopResponse()
to stop the agent from generating a reply. Here’s how it looks:
Copy code
async def on_user_turn_completed(self, turn_ctx, new_message):
    raise StopResponse()
3. Use Silent Function Tools: Another approach is to define tools that don’t return values or to cancel replies programmatically. For example:
Copy code
@function_tool()
async def silent_tool(self):
    # Perform logic without returning a value
    pass

@session.on("function_tools_executed")
def on_tool_executed(ev):
    ev.cancel_tool_reply()
If you’d like, I’m here to help you implement any of these methods or answer any other questions you might have. Just let me know! Sources: Agent speech and audio | LiveKit Docs | Pipeline nodes and hooks | LiveKit Docs | livekit-plugins/livekit-plugins-aws/livekit/plugins/aws/experimental/realti...
w
i want to use generate_reply to invoke a tool call via llm but do not want the generated message from generate reply to be added to chat ctx
t
Currently,
generate_reply
automatically adds its generated message to the chat context, and there isn’t a direct way to turn off this behavior. But don’t worry! You can still achieve what you want with a simple workaround: 1. Use
generate_reply
to Invoke the Tool:
Just let the LLM invoke the tool as it normally does with
generate_reply
. 2. Remove the Generated Message from the Chat Context: Once the tool is invoked, you can programmatically remove the generated message from the chat context like this:
Copy code
async def on_user_turn_completed(self, turn_ctx, new_message):
    # Call generate_reply
    await self.session.generate_reply(user_input="Invoke the tool")

    # Remove the last assistant message from the chat context
    if turn_ctx.chat_ctx.messages[-1]["role"] == "assistant":
        turn_ctx.chat_ctx.messages.pop()
This way, the tool gets invoked through the LLM, but the generated message won’t stay in the chat context. If anything’s unclear or you’d like me to walk you through it step-by-step, just let me know—I’m here to help!