brash-barista-66564
07/24/2025, 12:01 PMtall-airline-16756
07/24/2025, 12:01 PMtall-airline-16756
07/24/2025, 12:03 PMoutput_emitter with streaming enabled:
• Start a segment before pushing audio data:
• Push audio data:
• End the segment when complete:
2. Difference between ChunkedStream and SynthesizeStream:
• ChunkedStream: Processes a complete text input in one operation. Use this when you have the full text upfront and latency is not critical.
• SynthesizeStream: Allows real-time streaming synthesis with incremental text input. Use this for low-latency applications or when text is generated dynamically.
3. Methods to call on output_emitter in the _run() method of SynthesizeStream:
• Initialize the emitter:
• Start a segment:
• Push audio data:
• End the segment:
Error Explanation:
• The error RuntimeError: start_segment() must be called before pushing audio data occurs because start_segment() was not called before push().
• The error TypeError: start_segment() missing 1 required keyword-only argument: 'segment_id' indicates that the segment_id argument is mandatory. Ensure you pass it as a keyword argument, e.g., start_segment(segment_id="your_segment_id").
Let me know if you need further clarification or assistance!
Sources: livekit-plugins/livekit-plugins-playai/livekit/plugins/playai/tts.py | livekit-plugins/livekit-plugins-google/livekit/plugins/google/tts.py | plugins/openai/src/tts.ts | livekit-plugins/livekit-plugins-sarvam/livekit/plugins/sarvam/tts.pybrash-barista-66564
07/24/2025, 12:06 PMtall-airline-16756
07/24/2025, 12:11 PMpush(chunk)!
2. Initialize Parameters:
• For output_emitter.initialize(), mime_type="audio/pcm" is totally the way to go. While "audio/raw" works too, "audio/pcm" is the standard that everyone uses.
3. Error Handling in Streaming:
• If something goes wrong mid-stream, you'll want to use the _emit_error method to signal a TTSError event. It'll give all the juicy details like error type, when it happened, and if it's something that can be fixed.
• For errors that can be recovered, LiveKit's got your back with automatic retries. For the trickier ones, you can set up custom error handling, like playing a backup audio message.
• If you're using a FallbackAdapter, it'll automatically switch to a backup TTS provider if the main one stumbles, as long as no audio has been pushed for the current segment.
4. Buffering and Chunk Size:
• There's no hard and fast rule for chunk size with LiveKit TTS, but a good thumb rule is to keep chunks under 15kB and aim for 50-250ms of audio for smooth playback.
• Buffering smaller chunks isn't a must, but it might help keep your audio output nice and consistent.
5. Text Buffering Logic:
• For most apps, it's best to buffer text until you hit a punctuation mark or finish a sentence. This helps keep the speech sounding natural and avoids choppy audio.
• Sending text right away can speed things up, but it might make the speech sound a bit robotic. Only go this route if your app is super time-sensitive.
6. The _FlushSentinel Pattern:
• The _FlushSentinel is like a signal flag for segment boundaries in streaming TTS. It kicks in when you call the flush() method.
• It makes sure all the text in a segment is processed before moving to the next one. It's crucial for managing resources and keeping everything in sync.
• Different TTS systems handle it by flushing or resetting things like tokenizer streams when the sentinel shows up.
Need more info or just want to chat about it? I'm here for you! Feel free to ask anything.
Sources: livekit-agents/livekit/agents/tts/tts.py | livekit-plugins/livekit-plugins-inworld/livekit/plugins/inworld/tts.py | livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/gemini_t... | Events and error handling | LiveKit Docs | Resolving "InvalidState - failed to capture frame" errors in AudioSourcebrash-barista-66564
07/24/2025, 3:57 PMtall-airline-16756
07/24/2025, 3:58 PM"track_published" event.
Here's a friendly walkthrough of how to implement it:
@room.on("track_published")
def on_track_published(
publication: rtc.RemoteTrackPublication,
participant: rtc.RemoteParticipant
):
# Check if it's an audio track from a microphone
if (publication.kind == rtc.TrackKind.KIND_AUDIO and
publication.source == rtc.TrackSource.SOURCE_MICROPHONE):
<http://logger.info|logger.info>(
"Microphone track published: %s from participant %s (%s)",
publication.sid,
participant.sid,
participant.identity,
)
# Handle the microphone track publication here
Let me break down the key points for you:
• The "track_published" event is triggered when a remote participant publishes a track.
• To make sure it's a microphone audio track, you'll want to check:
◦ publication.kind == rtc.TrackKind.KIND_AUDIO
◦ publication.source == rtc.TrackSource.SOURCE_MICROPHONE
Pro tip: If you need to dive deeper and process the audio data, keep an eye out for the "track_subscribed" event. That's when you successfully subscribe to the track. Need any more help? I'm here for you!
Sources: livekit-agents/livekit/agents/voice/room_io/_input.py | livekit-rtc/livekit/rtc/room.pybrash-barista-66564
07/24/2025, 7:20 PMtall-airline-16756
07/24/2025, 7:27 PMbrash-barista-66564
07/24/2025, 7:50 PM