numerous-whale-53652
07/23/2025, 8:14 PMtall-airline-16756
07/23/2025, 8:14 PMtall-airline-16756
07/23/2025, 8:17 PMfrom livekit.agents import tts
import asyncio
from dataclasses import dataclass
@dataclass
class CustomTTSOptions:
sample_rate: int = 24000
voice: str = "default"
api_key: str = None
class CustomTTS(tts.TTS):
def __init__(self, **options):
self._opts = CustomTTSOptions(**options)
super().__init__(
capabilities=tts.TTSCapabilities(streaming=True),
sample_rate=self._opts.sample_rate,
num_channels=1,
)
if not self._opts.api_key:
raise ValueError("API key is required")
def synthesize(self, text: str) -> tts.ChunkedStream:
return CustomChunkedStream(tts=self, input_text=text, opts=self._opts)
def stream(self) -> tts.SynthesizeStream:
return CustomSynthesizeStream(tts=self, opts=self._opts)
JavaScript/TypeScript Implementation:
import { tts, AudioByteStream } from '@livekit/agents';
import { randomUUID } from 'node:crypto';
export class CustomTTS extends tts.TTS {
constructor(opts) {
super(opts.sampleRate || 24000, 1, { streaming: true });
this.opts = opts;
}
synthesize(text) {
return new CustomChunkedStream(this, text, this.opts);
}
stream() {
return new CustomSynthesizeStream(this, this.opts);
}
}
Key Components:
• Implement the synthesize() method to handle text-to-audio conversion.
• Use AudioByteStream for audio data conversion.
• Ensure proper error handling for API failures and invalid configurations.
Usage in Applications:
from livekit.agents import AgentSession
session = AgentSession(
tts=CustomTTS(
api_key="your-api-key",
voice="your-voice-id",
sample_rate=24000
)
)
For more details, you can refer to the contribution guidelines for Python and Node.js repositories:
• Python Contribution Guidelines
• Node.js Contribution Guidelines
This guide provides the structure and implementation details needed to create a custom TTS plugin for LiveKit, including error handling, streaming support, and integration patterns.
Sources: Text-to-speech (TTS) integrations | LiveKit Docs | Working with plugins | LiveKit Docs | TTS | LiveKit Agentsnumerous-whale-53652
07/28/2025, 10:24 AMtall-airline-16756
07/28/2025, 10:25 AMoutput_emitter needs to be set up just right before you start pushing data.
• Network/API failures: External TTS services can sometimes be a bit temperamental.
Here's a Friendly Implementation Guide:
Let's make sure your custom TTS plugin is doing things correctly in the _run method:
# 1. Initialize the output emitter
output_emitter.initialize(
request_id=utils.shortuuid(),
sample_rate=self._tts.sample_rate,
num_channels=NUM_CHANNELS,
mime_type=format,
)
# 2. Push actual audio data
async for data, _ in resp.content.iter_chunks():
output_emitter.push(data) # This is your key moment!
Let's Troubleshoot Together:
1. Verify API response: Double-check that your TTS service is actually sending back audio data.
2. Check data format: Make sure the audio data looks exactly like you expect before pushing.
3. Add logging: Keep track of your audio data's size and presence before calling output_emitter.push().
4. Test with known working data: Try pushing some sample audio to pinpoint where things might be going wrong.
I'm confident that by following these steps, we'll get to the bottom of this "no audio frames were pushed" mystery. You've got this, and I'm here to help!
Sources: livekit-agents/livekit/agents/voice/agent.pynumerous-whale-53652
07/28/2025, 10:28 AMtall-airline-16756
07/28/2025, 10:30 AM# Check text before synthesis
text = text.strip()
if not text:
# Handle empty text case
return
# Proceed with synthesis
tts_stream = tts.synthesize(text=text)
1. TTS Service Configuration Issues: Sometimes it's just about getting your settings right. Let's test things out with a simple "Hello world":
# Example for different TTS providers
from livekit.plugins import openai
# Ensure API keys are set
tts = openai.TTS(
api_key="your-api-key",
voice="alloy"
)
# Test with simple text
try:
stream = tts.synthesize("Hello world")
async for audio in stream:
print(f"Generated audio: {audio.frame.duration}s")
except APIError as e:
print(f"TTS Error: {e}")
1. Use Retry Logic and Fallback: No worries if one service doesn't work – we've got backup plans!
from livekit.agents.tts import FallbackAdapter
# Use fallback adapter with multiple TTS providers
primary_tts = openai.TTS()
fallback_tts = elevenlabs.TTS()
fallback_adapter = FallbackAdapter([primary_tts, fallback_tts])
try:
stream = fallback_adapter.synthesize(text)
async for audio in stream:
# Process audio
pass
except APIError as e:
print(f"All TTS providers failed: {e}")
1. Check TTS Provider Limits and Quotas: Sometimes services get a bit overwhelmed – keep an eye on those limits!
2. Validate Audio Output: Let's make sure we're actually getting some sound:
async def validate_tts_output(tts, text):
try:
stream = tts.synthesize(text)
frames = []
async for audio in stream:
if audio.frame.duration > 0:
frames.append(audio.frame)
else:
print(f"Warning: Empty frame for text: {text}")
if not frames:
raise APIError(f"No valid audio frames generated for: {text}")
return frames
except APIError as e:
print(f"TTS validation failed: {e}")
raise
1. Debug with Logging: When in doubt, turn on those logs and see what's happening behind the scenes:
import logging
logging.basicConfig(level=logging.DEBUG)
# This will show detailed TTS processing information
stream = tts.synthesize(text)
I know tech troubles can be super annoying, but we'll get this sorted out together! If you're still stuck after trying these steps, don't hesitate to reach out. We'll make sure your text-to-speech is up and running in no time.
Sources: livekit-plugins/livekit-plugins-lmnt/livekit/plugins/lmnt/tts.py