how to use my own tts or how to make tts plugin
# ask-ai
n
how to use my own tts or how to make tts plugin
t
Thanks for reaching out, Mohamed. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To use your own Text-to-Speech (TTS) or create a TTS plugin in LiveKit, you can leverage its extensible plugin architecture. Below is a detailed guide: Using Custom TTS: • LiveKit supports custom TTS implementations through its plugin framework. You can integrate with various providers or load models for local inference. • By adopting the standard TTS interfaces, you can simplify switching between different providers in your agent code. Creating a TTS Plugin: To create a custom TTS plugin, you need to inherit from the base TTS class and implement the required methods. Below are examples for Python and JavaScript/TypeScript: Python Implementation:
Copy code
from livekit.agents import tts
import asyncio
from dataclasses import dataclass

@dataclass
class CustomTTSOptions:
    sample_rate: int = 24000
    voice: str = "default"
    api_key: str = None

class CustomTTS(tts.TTS):
    def __init__(self, **options):
        self._opts = CustomTTSOptions(**options)
        super().__init__(
            capabilities=tts.TTSCapabilities(streaming=True),
            sample_rate=self._opts.sample_rate,
            num_channels=1,
        )
        if not self._opts.api_key:
            raise ValueError("API key is required")

    def synthesize(self, text: str) -> tts.ChunkedStream:
        return CustomChunkedStream(tts=self, input_text=text, opts=self._opts)

    def stream(self) -> tts.SynthesizeStream:
        return CustomSynthesizeStream(tts=self, opts=self._opts)
JavaScript/TypeScript Implementation:
Copy code
import { tts, AudioByteStream } from '@livekit/agents';
import { randomUUID } from 'node:crypto';

export class CustomTTS extends tts.TTS {
  constructor(opts) {
    super(opts.sampleRate || 24000, 1, { streaming: true });
    this.opts = opts;
  }

  synthesize(text) {
    return new CustomChunkedStream(this, text, this.opts);
  }

  stream() {
    return new CustomSynthesizeStream(this, this.opts);
  }
}
Key Components: • Implement the
synthesize()
method to handle text-to-audio conversion. • Use
AudioByteStream
for audio data conversion. • Ensure proper error handling for API failures and invalid configurations. Usage in Applications:
Copy code
from livekit.agents import AgentSession

session = AgentSession(
    tts=CustomTTS(
        api_key="your-api-key",
        voice="your-voice-id",
        sample_rate=24000
    )
)
For more details, you can refer to the contribution guidelines for Python and Node.js repositories: • Python Contribution Guidelines • Node.js Contribution Guidelines This guide provides the structure and implementation details needed to create a custom TTS plugin for LiveKit, including error handling, streaming support, and integration patterns. Sources: Text-to-speech (TTS) integrations | LiveKit Docs | Working with plugins | LiveKit Docs | TTS | LiveKit Agents
n
But LiveKit still thinks "no audio frames were pushed
t
Hey there! Dealing with the "no audio frames were pushed" error can be frustrating, but don't worry—I'm here to help you troubleshoot this step by step. Think of this like a friendly chat about getting your audio working smoothly. What's Typically Going Wrong: • Missing audio data push: Sometimes your custom TTS plugin just forgets to send over the actual audio content. It's like preparing a meal but forgetting to serve it! • Empty audio response: Your TTS service might be coming back empty-handed, like a pizza delivery with no pizza. • Incorrect initialization: The
output_emitter
needs to be set up just right before you start pushing data. • Network/API failures: External TTS services can sometimes be a bit temperamental. Here's a Friendly Implementation Guide: Let's make sure your custom TTS plugin is doing things correctly in the
_run
method:
Copy code
# 1. Initialize the output emitter
output_emitter.initialize(
    request_id=utils.shortuuid(),
    sample_rate=self._tts.sample_rate,
    num_channels=NUM_CHANNELS,
    mime_type=format,
)

# 2. Push actual audio data
async for data, _ in resp.content.iter_chunks():
    output_emitter.push(data)  # This is your key moment!
Let's Troubleshoot Together: 1. Verify API response: Double-check that your TTS service is actually sending back audio data. 2. Check data format: Make sure the audio data looks exactly like you expect before pushing. 3. Add logging: Keep track of your audio data's size and presence before calling
output_emitter.push()
. 4. Test with known working data: Try pushing some sample audio to pinpoint where things might be going wrong. I'm confident that by following these steps, we'll get to the bottom of this "no audio frames were pushed" mystery. You've got this, and I'm here to help! Sources: livekit-agents/livekit/agents/voice/agent.py
n
2025-07-28 132815,590 - DEBUG riva-tts-plugin - Ending Riva TTS streaming synthesis {"room": "xirk-anz5"} 2025-07-28 132815,590 - ERROR livekit.agents - Error in _tts_inference_task {"room": "xirk-anz5"} Traceback (most recent call last): File "C:\Users\ASUS\Desktop\agents\examples\voice_agents\venv\Lib\site-packages\livekit\agents\tts\tts.py", line 517, in anext val = await self._event_aiter.__anext__() ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ StopAsyncIteration During handling of the above exception, another exception occurred: Traceback (most recent call last): File "C:\Users\ASUS\Desktop\agents\examples\voice_agents\venv\Lib\site-packages\livekit\agents\utils\log.py", line 16, in async_fn_logs return await fn(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\ASUS\Desktop\agents\examples\voice_agents\venv\Lib\site-packages\opentelemetry\util\_decorator.py", line 71, in async_wrapper return await func(*args, **kwargs) # type: ignore ^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\ASUS\Desktop\agents\examples\voice_agents\venv\Lib\site-packages\livekit\agents\voice\generation.py", line 212, in _tts_inference_task async for audio_frame in tts_node: File "C:\Users\ASUS\Desktop\agents\examples\voice_agents\venv\Lib\site-packages\livekit\agents\voice\agent.py", line 408, in tts_node async for ev in stream: File "C:\Users\ASUS\Desktop\agents\examples\voice_agents\venv\Lib\site-packages\livekit\agents\tts\tts.py", line 520, in anext raise exc # noqa: B904 ^^^^^^^^^ File "C:\Users\ASUS\Desktop\agents\examples\voice_agents\venv\Lib\site-packages\opentelemetry\util\_decorator.py", line 71, in async_wrapper return await func(*args, **kwargs) # type: ignore ^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\ASUS\Desktop\agents\examples\voice_agents\venv\Lib\site-packages\livekit\agents\tts\tts.py", line 356, in _main_task raise APIError(f"no audio frames were pushed for text: {self._pushed_text}") livekit.agents._exceptions.APIError: no audio frames were pushed for text: Great! What's on your mind? (body=None, retryable=True)
t
Hey there! Dealing with the "no audio frames were pushed for text" error can be frustrating, but don't worry – I'm here to help you troubleshoot this step by step. Think of this like a friendly chat about getting your text-to-speech working smoothly. Common Causes and Solutions: 1. Empty or Invalid Text Input: First things first, let's make sure you're sending some actual text to convert. Here's a quick way to check:
Copy code
# Check text before synthesis
text = text.strip()
if not text:
    # Handle empty text case
    return

# Proceed with synthesis
tts_stream = tts.synthesize(text=text)
1. TTS Service Configuration Issues: Sometimes it's just about getting your settings right. Let's test things out with a simple "Hello world":
Copy code
# Example for different TTS providers
from livekit.plugins import openai

# Ensure API keys are set
tts = openai.TTS(
    api_key="your-api-key",
    voice="alloy"
)

# Test with simple text
try:
    stream = tts.synthesize("Hello world")
    async for audio in stream:
        print(f"Generated audio: {audio.frame.duration}s")
except APIError as e:
    print(f"TTS Error: {e}")
1. Use Retry Logic and Fallback: No worries if one service doesn't work – we've got backup plans!
Copy code
from livekit.agents.tts import FallbackAdapter

# Use fallback adapter with multiple TTS providers
primary_tts = openai.TTS()
fallback_tts = elevenlabs.TTS()

fallback_adapter = FallbackAdapter([primary_tts, fallback_tts])

try:
    stream = fallback_adapter.synthesize(text)
    async for audio in stream:
        # Process audio
        pass
except APIError as e:
    print(f"All TTS providers failed: {e}")
1. Check TTS Provider Limits and Quotas: Sometimes services get a bit overwhelmed – keep an eye on those limits! 2. Validate Audio Output: Let's make sure we're actually getting some sound:
Copy code
async def validate_tts_output(tts, text):
    try:
        stream = tts.synthesize(text)
        frames = []
        
        async for audio in stream:
            if audio.frame.duration > 0:
                frames.append(audio.frame)
            else:
                print(f"Warning: Empty frame for text: {text}")
        
        if not frames:
            raise APIError(f"No valid audio frames generated for: {text}")
            
        return frames
    except APIError as e:
        print(f"TTS validation failed: {e}")
        raise
1. Debug with Logging: When in doubt, turn on those logs and see what's happening behind the scenes:
Copy code
import logging
logging.basicConfig(level=logging.DEBUG)

# This will show detailed TTS processing information
stream = tts.synthesize(text)
I know tech troubles can be super annoying, but we'll get this sorted out together! If you're still stuck after trying these steps, don't hesitate to reach out. We'll make sure your text-to-speech is up and running in no time. Sources: livekit-plugins/livekit-plugins-lmnt/livekit/plugins/lmnt/tts.py