Is it possible to use prerecorded messages with a ...
# ask-ai
i
Is it possible to use prerecorded messages with a multimodal agent to cut costs for common phrases on api usage?
t
Thanks for reaching out, Matthew. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Absolutely, you can use prerecorded messages with a multimodal agent in LiveKit to help reduce API usage and save on costs, especially for those common phrases. It’s a smart way to play audio files for frequently used phrases instead of generating each response through TTS APIs. For instance, you might use the
play_wav_once("let_me_check_that.wav", ctx.room)
function to play a quick audio clip while the system is processing data. Plus, the
session.say()
method supports playing pre-synthesized audio, which skips the TTS step and helps keep costs down. This method works really well for keeping interactions natural and responsive, all while optimizing your API usage. If you’d like, I’m here to help with more details or guide you through the implementation! Sources: Agent speech and audio | LiveKit Docs | sip-refer/agent.py | examples/voice_agents/fast-preresponse.py | Working with the MultimodalAgent class | LiveKit Docs | Agent speech and audio | LiveKit Docs
i
My agent calls knowledge banks based on query string data, each knowledge bank would need its own definition of predefined audio clips to play. How can this be implemented without the realtime api agent speaking over or getting interrupted by the prerecorded parts?
t
I'm glad you're looking to implement predefined audio clips smoothly alongside your real-time API agent's speech. Here are some friendly strategies to help you avoid interruptions or overlap: 1. Disable Interruptions: Try using the
allow_interruptions=False
parameter in
session.say()
. This way, your prerecorded audio clips can play fully without being cut off by user input or other system events. 2. Discard Incoming Audio: Enabling the
discard_audio_if_uninterruptible
option can be really helpful. It prevents user audio from being processed during those important non-interruptible playback moments, so you won’t get any unintended triggers. 3. Play Pre-Synthesized Audio: Consider using an
AsyncIterable[rtc.AudioFrame]
to feed prerecorded audio directly into
session.say()
. This bypasses text-to-speech synthesis, making playback smoother and more efficient. 4. Wait for Playout Completion: Using
SpeechHandle.wait_for_playout()
ensures that any next steps—like playing another clip or responding to the user—only happen after the current audio finishes. It’s a great way to keep things in order. 5. Manage Speech Priorities: Assign priority levels to your speech tasks. This helps control the playback sequence and makes sure important messages aren’t delayed. These approaches should help your prerecorded clips and real-time speech work together seamlessly. If you'd like, I’m here to provide more details or even some code examples tailored to your setup. Just let me know! Sources: agents/src/voice/speech_handle.ts | Agent speech and audio | LiveKit Docs | livekit-agents/livekit/agents/voice/agent_activity.py | Bringing AI avatars to voice agents