does the TTS pipeline node get called before or af...
# ask-ai
c
does the TTS pipeline node get called before or after the utterance is sent to the user?
t
Thanks for reaching out, Faraz. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Thanks for reaching out! Let me walk you through how the TTS pipeline node works. It’s called before the utterance reaches you, synthesizing speech from the text that the LLM node generates. This audio is then sent to you as part of the agent’s reply. Here’s the sequence in the STT-LLM-TTS pipeline to give you a clearer picture: 1.
stt_node()
transcribes your audio into text. 2.
llm_node()
takes that text and crafts a response. 3.
tts_node()
turns the LLM’s text response into audio. 4. Finally, the synthesized audio is sent your way. Also, for realtime models, there’s a handy
realtime_audio_output_node
that lets us tweak the audio output before you receive it. This just goes to show that the audio processing happens before it gets delivered to you. Sources: Pipeline nodes and hooks | LiveKit Docs | Agents 0.x migration guide | LiveKit Docs
c
how does the tts node work by default? how does it synthesize my text to audio?
t
I'm happy to walk you through how the TTS node in LiveKit works! It takes text segments and turns them into speech, using the Text-To-Speech (TTS) capability provided by the agent by default. If the TTS implementation doesn’t support streaming right away, no worries! The node smartly breaks the text into smaller chunks using a sentence tokenizer, so it can synthesize audio bit by bit. This approach helps keep things responsive and efficient, especially when you’re working with real-time applications. The TTS node handles an asynchronous stream of text segments and produces
rtc.AudioFrame
objects, which are the pieces of synthesized audio. You also have the flexibility to customize how it works—whether that’s changing how text is chunked, plugging in your own TTS engine, tweaking pronunciation, or adding other special audio processing. For instance, when working with providers like ElevenLabs or OpenAI, the TTS node connects to their APIs to convert text to speech. With ElevenLabs, it streams text over a WebSocket connection, processes it as it comes in, and turns the audio data into
AudioFrame
objects in real time. Sources: Docs::Agents | Pipeline nodes and hooks | LiveKit Docs | Pipeline nodes and hooks | LiveKit Docs | livekit-plugins/livekit-plugins-lmnt/livekit/plugins/lmnt/tts.py | plugins/elevenlabs/src/tts.ts