brash-barista-66564
07/24/2025, 9:40 PMtall-airline-16756
07/24/2025, 9:40 PMtall-airline-16756
07/24/2025, 9:41 PM# Create audio source with TTS sample rate and channels
source = rtc.AudioSource(tts_11labs.sample_rate, tts_11labs.num_channels)
track = rtc.LocalAudioTrack.create_audio_track("agent-mic", source)
options = rtc.TrackPublishOptions(source=rtc.TrackSource.SOURCE_MICROPHONE)
# Connect and publish
await ctx.connect(auto_subscribe=AutoSubscribe.SUBSCRIBE_NONE)
publication = await ctx.room.local_participant.publish_track(track, options)
await publication.wait_for_subscription()
2. Connection Sequence
The connection sequence is critical. Ensure the following steps are followed:
1. Connect to the room first.
2. Publish the audio track.
3. Wait for subscription before sending audio.
4. Start capturing TTS frames to the audio source.
3. Audio Frame Capture
Even if TTS generates audio chunks correctly, you need to capture them to the audio source. Use the following example:
async def _playout_task(
tts_forwarder: transcription.TTSSegmentsForwarder,
playout_q: asyncio.Queue,
audio_source: rtc.AudioSource,
) -> None:
"""Playout audio frames from the queue to the audio source"""
tts_forwarder.segment_playout_started()
while True:
frame = await playout_q.get()
if frame is None:
break
await audio_source.capture_frame(frame)
4. TTS Audio Frame Processing
Ensure that the TTS system generates SynthesizedAudio objects containing rtc.AudioFrame data. The AudioEmitter class should process TTS output and create properly formatted audio frames with the correct sample rate, channels, and format.
5. Debugging Steps
• Verify that the audio track is published and subscribed to.
• Confirm the rtc.AudioSource is created with matching TTS sample rate and channels.
• Ensure source.capture_frame(frame) is being called for each TTS audio frame.
• Check the connection timing (connect → publish → wait for subscription → send audio).
• Verify that the TTS audio format matches what the audio source expects.
By addressing these areas, you should be able to resolve the issue and ensure the agent's audio is heard by the user.
Sources: livekit-plugins/livekit-plugins-deepgram/livekit/plugins/deepgram/tts.pybrash-barista-66564
07/24/2025, 9:43 PMtall-airline-16756
07/24/2025, 9:44 PMsession.start() with room_output_options that has audio_enabled=True, the session does all the heavy lifting for you. It automatically creates a RoomIO object that handles audio track publishing. Easy peasy!
2. Does AgentSession.start() handle the audio track publishing automatically when audio_enabled=True?
Absolutely! When you set audio_enabled=True in the room_output_options, AgentSession.start() takes care of audio track publishing automatically. The RoomIO utility class acts like a friendly bridge between your agent session and the LiveKit room, making sure all room participants can tune into the audio tracks.
3. What's the correct connection sequence when using AgentSession?
Here's a simple step-by-step guide:
1. Get your agent started and set the state to "initializing."
2. If a room is provided and you don't have a custom RoomIO, create a RoomIO object.
3. Start the RoomIO task to manage those tracks.
4. Connect to the job context, which will automatically connect to the room when RoomIO is in play.
Here's an example to help you out:
async def entrypoint(ctx: JobContext):
session = AgentSession(
stt=deepgram.STT(),
llm=openai.LLM(),
tts=cartesia.TTS(),
)
await session.start(
agent=MyAgent(),
room=ctx.room, # Room connection handled automatically
room_output_options=RoomOutputOptions(audio_enabled=True)
)
4. Could the issue be related to my custom TTS implementation not properly integrating with the agent session's audio pipeline?
You might be onto something! If you're using a custom TTS implementation, here are some tips to make sure everything runs smoothly:
• Set up your custom audio output before starting the session by configuring session.output.audio.
• Disable RoomIO audio by setting audio_enabled=False in room_output_options to avoid any potential conflicts.
• If you disable RoomIO audio, you'll need to handle track publishing manually.
Here's an example to help you out:
# Custom audio output configured
session.output.audio = DataStreamAudioOutput(ctx.room, destination_identity=AVATAR_IDENTITY)
await session.start(
agent=agent,
room=ctx.room,
room_output_options=RoomOutputOptions(
audio_enabled=False, # Disabled because custom output is used
transcription_enabled=True,
),
)
The key is to make sure only one audio output mechanism is active to prevent any conflicts between the automatic RoomIO publishing and your custom TTS setup. Hope this helps!
Sources: livekit-agents/livekit/agents/voice/agent_activity.py | Building voice agents | LiveKit Docsbrash-barista-66564
07/24/2025, 9:46 PMtall-airline-16756
07/24/2025, 9:48 PMstart_segment() before pushing your audio
• Make sure to wrap up each segment with end_segment()
• Give each segment its own unique ID - think of it like giving each audio clip a name tag
• Create a fresh stream for each segment (pro tip: reusing streams is now old school)
3. Could the "Previous segment still active" warnings indicate why audio isn't being heard?
You bet! These warnings are like little red flags saying your segments might be stepping on each other's toes. If one segment isn't properly closed before the next one starts, your audio might get a bit shy and not play. The fix? Just make sure you're closing each segment completely before starting the next.
4. Do I need to implement any specific synchronization in my custom TTS when used with AgentSession?
Yep, synchronization is your new best friend! Here's how to keep things in sync:
• Get cozy with the TextAudioSynchronizer class to keep your text and audio playing nicely together
• Use markAudioSegmentEnd() and markTextSegmentEnd() to give clear signals about your segment boundaries
• Handle those segment transitions like a pro to avoid any awkward audio conflicts
Follow these tips, and you'll have your custom TTS working with AgentSession smoother than butter. You've got this!
Sources: livekit-agents/livekit/agents/voice/agent_activity.py | agents/src/tts/tts.ts