I'm having an issue where my LiveKit agent can rec...
# ask-ai
b
I'm having an issue where my LiveKit agent can receive user audio and generate TTS responses, but the user cannot hear the agent's audio. Looking at the logs: 1. The agent successfully starts with audio_enabled=True in RoomOutputOptions 2. The agent receives user speech ("Hello", "Hello, hello") 3. The TTS generates audio chunks and processes them 4. The agent appears to be creating multiple segments (I see "Previous segment still active, ending it first" warnings) However, the user reports not hearing any audio from the agent. The agent is using a custom TTS implementation that creates one segment per synthesis request. My current configuration: await session.start( agent=briseis_agent, room=ctx.room, room_input_options=RoomInputOptions( noise_cancellation=noise_cancellation.BVC(), ), room_output_options=RoomOutputOptions( audio_enabled=True, transcription_enabled=True ), ) # Then connecting with: await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY) What could be preventing the agent's audio from being heard by the user? Is there something wrong with my connection sequence or audio publishing setup?
t
Thanks for reaching out, Andrew. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To resolve the issue where the LiveKit agent's TTS audio is not heard by the user, despite generating audio chunks and segments, please review the following key areas: 1. Audio Publishing Setup Ensure that the audio track is being published correctly. Here is an example of the correct sequence:
Copy code
# Create audio source with TTS sample rate and channels
source = rtc.AudioSource(tts_11labs.sample_rate, tts_11labs.num_channels)
track = rtc.LocalAudioTrack.create_audio_track("agent-mic", source)
options = rtc.TrackPublishOptions(source=rtc.TrackSource.SOURCE_MICROPHONE)

# Connect and publish
await ctx.connect(auto_subscribe=AutoSubscribe.SUBSCRIBE_NONE)
publication = await ctx.room.local_participant.publish_track(track, options)
await publication.wait_for_subscription()
2. Connection Sequence The connection sequence is critical. Ensure the following steps are followed: 1. Connect to the room first. 2. Publish the audio track. 3. Wait for subscription before sending audio. 4. Start capturing TTS frames to the audio source. 3. Audio Frame Capture Even if TTS generates audio chunks correctly, you need to capture them to the audio source. Use the following example:
Copy code
async def _playout_task(
    tts_forwarder: transcription.TTSSegmentsForwarder,
    playout_q: asyncio.Queue,
    audio_source: rtc.AudioSource,
) -> None:
    """Playout audio frames from the queue to the audio source"""
    tts_forwarder.segment_playout_started()
    while True:
        frame = await playout_q.get()
        if frame is None:
            break

        await audio_source.capture_frame(frame)
4. TTS Audio Frame Processing Ensure that the TTS system generates
SynthesizedAudio
objects containing
rtc.AudioFrame
data. The
AudioEmitter
class should process TTS output and create properly formatted audio frames with the correct sample rate, channels, and format. 5. Debugging Steps • Verify that the audio track is published and subscribed to. • Confirm the
rtc.AudioSource
is created with matching TTS sample rate and channels. • Ensure
source.capture_frame(frame)
is being called for each TTS audio frame. • Check the connection timing (connect → publish → wait for subscription → send audio). • Verify that the TTS audio format matches what the audio source expects. By addressing these areas, you should be able to resolve the issue and ensure the agent's audio is heard by the user. Sources: livekit-plugins/livekit-plugins-deepgram/livekit/plugins/deepgram/tts.py
b
I'm using the LiveKit Python agents framework with AgentSession. The agent receives audio but users can't hear the agent's responses. Looking at your previous response about creating an AudioSource and LocalAudioTrack, I'm confused because I'm using the high-level AgentSession API, not the low-level track publishing. My code: session = AgentSession( llm=openai.LLM(model="gpt-4o"), stt=openai.STT(model="gpt-4o-transcribe", language="en"), tts=tts_instance, # Custom TTS turn_detection=EnglishModel(), vad=ctx.proc.userdata["vad"], ) await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY) await session.start( agent=briseis_agent, room=ctx.room, room_input_options=RoomInputOptions(noise_cancellation=noise_cancellation.BVC()), room_output_options=RoomOutputOptions( audio_enabled=True, transcription_enabled=True ), ) 1. When using AgentSession, do I still need to manually create and publish audio tracks? 2. Does AgentSession.start() handle the audio track publishing automatically when audio_enabled=True? 3. What's the correct connection sequence when using AgentSession? Should I connect before or after calling session.start()? 4. Could the issue be related to my custom TTS implementation not properly integrating with the agent session's audio pipeline?
t
Hey there! I totally get your questions about using the LiveKit Python AgentSession API. Let me break it down for you in a friendly way: 1. Do I still need to manually create and publish audio tracks when using AgentSession? Great news! No, you don't have to worry about manually creating and publishing audio tracks. When you call
session.start()
with
room_output_options
that has
audio_enabled=True
, the session does all the heavy lifting for you. It automatically creates a
RoomIO
object that handles audio track publishing. Easy peasy! 2. Does
AgentSession.start()
handle the audio track publishing automatically when
audio_enabled=True
?
Absolutely! When you set
audio_enabled=True
in the
room_output_options
,
AgentSession.start()
takes care of audio track publishing automatically. The
RoomIO
utility class acts like a friendly bridge between your agent session and the LiveKit room, making sure all room participants can tune into the audio tracks. 3. What's the correct connection sequence when using AgentSession? Here's a simple step-by-step guide: 1. Get your agent started and set the state to "initializing." 2. If a room is provided and you don't have a custom
RoomIO
, create a
RoomIO
object. 3. Start the
RoomIO
task to manage those tracks. 4. Connect to the job context, which will automatically connect to the room when
RoomIO
is in play. Here's an example to help you out:
Copy code
async def entrypoint(ctx: JobContext):
    session = AgentSession(
        stt=deepgram.STT(),
        llm=openai.LLM(),
        tts=cartesia.TTS(),
    )

    await session.start(
        agent=MyAgent(),
        room=ctx.room,  # Room connection handled automatically
        room_output_options=RoomOutputOptions(audio_enabled=True)
    )
4. Could the issue be related to my custom TTS implementation not properly integrating with the agent session's audio pipeline? You might be onto something! If you're using a custom TTS implementation, here are some tips to make sure everything runs smoothly: • Set up your custom audio output before starting the session by configuring
session.output.audio
. • Disable
RoomIO
audio by setting
audio_enabled=False
in
room_output_options
to avoid any potential conflicts. • If you disable
RoomIO
audio, you'll need to handle track publishing manually. Here's an example to help you out:
Copy code
# Custom audio output configured
session.output.audio = DataStreamAudioOutput(ctx.room, destination_identity=AVATAR_IDENTITY)

await session.start(
    agent=agent,
    room=ctx.room,
    room_output_options=RoomOutputOptions(
        audio_enabled=False,  # Disabled because custom output is used
        transcription_enabled=True,
    ),
)
The key is to make sure only one audio output mechanism is active to prevent any conflicts between the automatic
RoomIO
publishing and your custom TTS setup. Hope this helps! Sources: livekit-agents/livekit/agents/voice/agent_activity.py | Building voice agents | LiveKit Docs
b
My agent is still not producing audible output despite using AgentSession with audio_enabled=True. I notice in the logs that my custom TTS is generating audio chunks, but I'm getting warnings like "Previous segment still active, ending it first". My custom TTS extends tts.TTS and implements a streaming interface that: 1. Creates ONE segment per synthesis request using output_emitter.start_segment() 2. Pushes audio chunks with output_emitter.push() 3. Ends the segment with output_emitter.end_segment() However, I'm seeing multiple synthesis requests happening simultaneously, causing segment conflicts. Questions: 1. Should the AgentSession automatically queue TTS synthesis requests to prevent overlapping segments? 2. Is there a specific way I should implement segment handling in my custom TTS to work properly with AgentSession? 3. Could the "Previous segment still active" warnings indicate why audio isn't being heard? 4. Do I need to implement any specific synchronization in my custom TTS when used with AgentSession? Here's a snippet of the warnings from logs: 2025-07-24 183924,838 - WARNING custom-baseten-tts - Previous segment still active, ending it first 2025-07-24 183924,839 - WARNING custom-baseten-tts - Previous segment still active, ending it first
t
Hey there! Let me help you navigate the world of TTS synchronization with LiveKit AgentSession. I totally understand how tricky these technical details can be, so I'll break it down in a friendly way. 1. Should the AgentSession automatically queue TTS synthesis requests to prevent overlapping segments? Not quite! The AgentSession won't automatically line up your TTS requests for you. You'll need to take the wheel and manage those segments yourself. Just a heads up - if you try to start a new segment before the previous one is done, LiveKit will give you a friendly error nudge. 2. Is there a specific way I should implement segment handling in my custom TTS to work properly with AgentSession? Absolutely! Here are some pro tips to keep things smooth: • Always kick things off with
start_segment()
before pushing your audio • Make sure to wrap up each segment with
end_segment()
• Give each segment its own unique ID - think of it like giving each audio clip a name tag • Create a fresh stream for each segment (pro tip: reusing streams is now old school) 3. Could the "Previous segment still active" warnings indicate why audio isn't being heard? You bet! These warnings are like little red flags saying your segments might be stepping on each other's toes. If one segment isn't properly closed before the next one starts, your audio might get a bit shy and not play. The fix? Just make sure you're closing each segment completely before starting the next. 4. Do I need to implement any specific synchronization in my custom TTS when used with AgentSession? Yep, synchronization is your new best friend! Here's how to keep things in sync: • Get cozy with the
TextAudioSynchronizer
class to keep your text and audio playing nicely together • Use
markAudioSegmentEnd()
and
markTextSegmentEnd()
to give clear signals about your segment boundaries • Handle those segment transitions like a pro to avoid any awkward audio conflicts Follow these tips, and you'll have your custom TTS working with AgentSession smoother than butter. You've got this! Sources: livekit-agents/livekit/agents/voice/agent_activity.py | agents/src/tts/tts.ts