This message was deleted.
# ask-ai
s
This message was deleted.
t
Thanks for reaching out, Andrew. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Here are the answers to your questions about using the LiveKit Python framework with `AgentSession`: 1. Do you need to manually create and publish audio tracks? No, you do not need to manually create and publish audio tracks when using
AgentSession
. If
audio_enabled=True
is set in
RoomOutputOptions
, the session automatically handles audio track creation and publishing. This includes creating a
LocalAudioTrack
and publishing it as a microphone source. 2. Does
AgentSession.start()
handle audio track publishing automatically?
Yes,
AgentSession.start()
automatically handles audio track publishing when
audio_enabled=True
. It sets up the audio output pipeline and publishes the track without requiring additional manual steps. 3. What is the correct connection sequence? The correct sequence is: a. Create the
AgentSession
instance with the required components (e.g., TTS, STT, VAD). b. Connect the job context using
await ctx.connect()
if applicable. c. Call
session.start()
with the agent and room. This method handles all internal setup, including connecting to the room and initializing audio/video tasks. 1. Could the issue be related to your custom TTS implementation? Yes, a custom TTS implementation could cause issues if it does not properly yield
rtc.AudioFrame
objects or handle streaming correctly. Ensure that your TTS implementation: If you suspect the custom TTS is causing the issue, verify its output and test it independently to ensure compatibility. ◦ Yields properly formatted
rtc.AudioFrame
objects. ◦ Handles streaming efficiently to avoid delays or interruptions. ◦ Includes robust error handling and timeout management. ◦ Integrates seamlessly with the
AgentSession
audio pipeline. Let me know if you need further clarification or assistance! Sources: livekit-agents/livekit/agents/voice/room_io/_output.py | evaluating-agents/agent_evals.py | livekit-agents/livekit/agents/worker.py | agent.py | Pipeline nodes and hooks | LiveKit Docs
b
My agent is still not producing audible output despite using AgentSession with audio_enabled=True. I notice in the logs that my custom TTS is generating audio chunks, but I'm getting warnings like "Previous segment still active, ending it first". My custom TTS extends tts.TTS and implements a streaming interface that: 1. Creates ONE segment per synthesis request using output_emitter.start_segment() 2. Pushes audio chunks with output_emitter.push() 3. Ends the segment with output_emitter.end_segment() However, I'm seeing multiple synthesis requests happening simultaneously, causing segment conflicts. Questions: 1. Should the AgentSession automatically queue TTS synthesis requests to prevent overlapping segments? 2. Is there a specific way I should implement segment handling in my custom TTS to work properly with AgentSession? 3. Could the "Previous segment still active" warnings indicate why audio isn't being heard? 4. Do I need to implement any specific synchronization in my custom TTS when used with AgentSession? Here's a snippet of the warnings from logs: 2025-07-24 183924,838 - WARNING custom-baseten-tts - Previous segment still active, ending it first 2025-07-24 183924,839 - WARNING custom-baseten-tts - Previous segment still active, ending it first
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!