crooked-tiger-31253
09/20/2025, 6:49 PMtall-airline-16756
09/20/2025, 6:49 PMtall-airline-16756
09/20/2025, 6:49 PMstt_node() transcribes your audio into text.
2. llm_node() takes that text and crafts a response.
3. tts_node() turns the LLM’s text response into audio.
4. Finally, the synthesized audio is sent your way.
Also, for realtime models, there’s a handy realtime_audio_output_node that lets us tweak the audio output before you receive it. This just goes to show that the audio processing happens before it gets delivered to you.
Sources: Pipeline nodes and hooks | LiveKit Docs | Agents 0.x migration guide | LiveKit Docscrooked-tiger-31253
09/20/2025, 6:50 PMtall-airline-16756
09/20/2025, 6:51 PMrtc.AudioFrame objects, which are the pieces of synthesized audio. You also have the flexibility to customize how it works—whether that’s changing how text is chunked, plugging in your own TTS engine, tweaking pronunciation, or adding other special audio processing.
For instance, when working with providers like ElevenLabs or OpenAI, the TTS node connects to their APIs to convert text to speech. With ElevenLabs, it streams text over a WebSocket connection, processes it as it comes in, and turns the audio data into AudioFrame objects in real time.
Sources: Docs::Agents | Pipeline nodes and hooks | LiveKit Docs | Pipeline nodes and hooks | LiveKit Docs | livekit-plugins/livekit-plugins-lmnt/livekit/plugins/lmnt/tts.py | plugins/elevenlabs/src/tts.ts