future-shampoo-29056
09/23/2025, 7:41 AMtall-airline-16756
09/23/2025, 7:41 AMtall-airline-16756
09/23/2025, 7:42 AMstream() method in LiveKit’s TTS integrations, which lets you synthesize audio in real-time, bit by bit. Here’s a simple overview to guide you through the process:
• Initialize TTS and Audio Source: Start by creating a TTS instance along with an AudioSource that matches the sample rate and channels.
• Publish Audio Track: Then, publish the audio track to the room so everyone can hear the speech as it’s generated.
• Stream TTS: Use the stream() method to handle the text incrementally, sending audio frames to the room as soon as they’re ready.
• Push Text Chunks: Feed chunks of text into the TTS stream—like from a language model—and publish each audio frame in real-time.
This method really helps cut down on delays since you don’t have to wait for the entire TTS response before playing it back. If you’d like, I’m here to help with detailed examples or any questions you have along the way!
Sources: Text-to-speech (TTS) integrations | LiveKit Docs | agents/src/tts/tts.ts | livekit-plugins/livekit-plugins-spitch/livekit/plugins/spitch/tts.pyfuture-shampoo-29056
09/23/2025, 7:44 AMtall-airline-16756
09/23/2025, 7:44 AMtts.TTS class provided by LiveKit. Here’s a friendly, high-level overview of how you might approach this:
1. Define the Custom TTS Class: Start by inheriting from tts.TTS and override the stream() method so it returns a custom SynthesizeStream instance tailored for your needs.
2. Implement the Streaming Logic: Inside your SynthesizeStream subclass, you’ll want to connect to the external TTS API using WebSocket or SSE to handle real-time audio streaming smoothly.
3. Emit Audio Frames: As the API sends audio data, process it incrementally and emit these as AudioFrame objects to the LiveKit room, ensuring a seamless streaming experience.
Here’s a simple example to illustrate the structure:
class CustomTTS(tts.TTS):
def __init__(self, api_key, sample_rate=24000):
super().__init__(sample_rate, 1, streaming=True)
self.api_key = api_key
def stream(self):
return CustomSynthesizeStream(self)
class CustomSynthesizeStream(tts.SynthesizeStream):
def __init__(self, tts):
super().__init__(tts)
async def run(self):
# Connect to API and handle streaming
pass
This method gives you the flexibility to integrate with any API that supports streaming TTS responses. If you’d like, I’m here to help with a more detailed implementation or any specific questions you might have—just let me know!
Sources: plugins/openai/src/tts.tsfuture-shampoo-29056
09/23/2025, 7:51 AMtall-airline-16756
09/23/2025, 7:52 AMfuture-shampoo-29056
09/23/2025, 7:52 AMtall-airline-16756
09/23/2025, 7:53 AMfuture-shampoo-29056
09/25/2025, 9:21 AMtall-airline-16756
09/25/2025, 9:22 AMpcm_s16le) or compressed formats such as OGG or MP3, depending on the provider. LiveKit helps by automatically resampling audio to fit WebRTC standards (like 24kHz).
• Simulated Streaming: If a provider doesn’t support streaming directly, LiveKit uses a StreamAdapter to break text into chunks and process them one after another, creating a streaming-like experience.
• Real-Time Metrics: Streaming setups often track metrics such as Time to First Byte (TTFB) and synthesis duration to keep an eye on performance.
If you’re building your own streaming TTS, just make sure your API can handle incremental text input and audio output. Feel free to reach out if you’d like specific examples or need any more help—I’m here to support you!
Sources: livekit-plugins/livekit-plugins-inworld/livekit/plugins/inworld/tts.py | agents/src/tts/stream_adapter.tsfuture-shampoo-29056
09/25/2025, 9:23 AMtall-airline-16756
09/25/2025, 9:23 AMstreaming=True, there’s a helpful way to simulate streaming using LiveKit’s StreamAdapter. This tool breaks your text into smaller chunks, like sentences, and processes them step-by-step, giving you that nice streaming feel.
Here’s a simple way to set it up:
• Wrap OpenAI TTS: Use StreamAdapter to enable incremental audio synthesis.
• Push Text: Send your text chunks to the adapter, which handles them one at a time.
• Consume Audio: Receive audio frames as they’re generated, almost like real-time streaming.
Here’s an example to guide you:
from livekit.agents.tts import StreamAdapter
from livekit.agents.tokenize import SentenceTokenizer
streaming_tts = StreamAdapter(tts=openai_tts, sentence_tokenizer=SentenceTokenizer())
stream = streaming_tts.stream()
stream.push_text("Hello world.")
stream.flush()
This method helps reduce the delay you might notice and works smoothly with real-time applications. If you have any questions or want me to walk you through it, just let me know—I’m here to help!
Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/stt.py | agents/src/tts/tts.ts | Tool definition and use | LiveKit Docsfuture-shampoo-29056
09/25/2025, 9:30 AMtall-airline-16756
09/25/2025, 9:31 AMcapabilities by setting streaming=True. This will let your plugin know streaming is supported:
super().__init__(
capabilities=tts.TTSCapabilities(streaming=True),
sample_rate=sample_rate,
num_channels=1,
)
• Implement the stream() Method: Add a stream() method that returns a SynthesizeStream instance. This is key for streaming functionality:
def stream(self, *, conn_options=DEFAULT_API_CONNECT_OPTIONS):
return KittenTTSSynthesizeStream(tts=self, conn_options=conn_options)
• Create a Streaming Class: You’ll want to implement a KittenTTSSynthesizeStream class inheriting from tts.SynthesizeStream. This class will manage real-time text-to-speech synthesis by:
◦ Initializing the AudioEmitter to handle streaming audio.
◦ Sending text chunks to your TTS model step-by-step.
◦ Emitting audio frames as soon as they’re generated, so the streaming feels smooth.
Here’s a helpful example of the _run() method inside your streaming class:
async def _run(self, output_emitter):
output_emitter.initialize(
request_id=str(id(self)),
sample_rate=self._tts.sample_rate,
num_channels=1,
mime_type="audio/pcm",
stream=True,
)
for text_chunk in self.input_text.split(" "):
audio = self._tts.model.generate(text_chunk, voice=self._tts.voice)
audio_bytes = (audio * 32767).astype(np.int16).tobytes()
output_emitter.push(audio_bytes)
output_emitter.end_input()
These updates will enable your TTS plugin to support real-time streaming beautifully. If you have any questions or want me to walk you through any part, just let me know—I’m here to help!
Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/tts.pyfuture-shampoo-29056
09/25/2025, 9:34 AMtall-airline-16756
09/25/2025, 9:34 AMsynthesize() method, which processes the entire text input and generates audio frames all at once. Here’s a friendly breakdown of what happens:
• Non-Streaming Behavior: The TTS engine creates the full audio output for your text, and the audio frames come back as an asynchronous stream.
• Simulated Streaming: If you’d like a streaming-like experience, you can use a StreamAdapter to split your text into smaller chunks (like sentences) and process them step-by-step, which mimics real-time streaming nicely.
• Aligned Transcripts: By enabling use_tts_aligned_transcript=True, you can get sentence-level timestamps that help sync the transcript perfectly with the audio playback.
• Text Preprocessing: You can also use the before_tts_cb callback to clean up or adjust your text before synthesis, making sure the spoken output sounds just right.
This way, even if the TTS isn’t streaming by default, it can still fit smoothly into your real-time applications. If you want, I’m here to help you set these up or answer any questions you have!
Sources: agents/src/tts/tts.ts | examples/voice_agents/timed_agent_transcript.py | chain-of-thought-tts/README.mdoc | Pipeline nodes and hooks | LiveKit Docsfuture-shampoo-29056
09/25/2025, 9:41 AMtall-airline-16756
09/25/2025, 9:42 AMttft (Time to First Token): This measures the time to generate the very first token.
◦ `duration`: This is the total time taken to generate the entire response.
• TTS (Text-to-Speech):
◦ ttfb (Time to First Byte): This indicates the time to generate the first byte of audio.
◦ `duration`: This covers the total time for the TTS request.
◦ `audio_duration`: This reflects the duration of the generated audio output.
To get these metrics, just subscribe to the metrics_collected event in your agent code. From there, you can log them or even forward them to external systems for a deeper analysis. If you'd like, I’m here to help you implement this step-by-step—just let me know!
Sources: Logs, metrics, and telemetry | LiveKit Docs | metrics/metrics_stt.py | agents/src/metrics/base.ts | livekit-agents/livekit/agents/voice/agent.py | Capturing metrics | LiveKit Docsfuture-shampoo-29056
09/25/2025, 9:47 AMtall-airline-16756
09/25/2025, 9:47 AMmetrics_collected event in LiveKit! You can easily register an event listener on your agent session or voice pipeline agent. Here’s a simple way to do it:
• For AgentSession:
from livekit.agents import metrics, MetricsCollectedEvent
@session.on("metrics_collected")
def on_metrics_collected(ev: MetricsCollectedEvent):
metrics.log_metrics(ev.metrics)
• For VoicePipelineAgent:
from livekit.agents import metrics
@agent.on("metrics_collected")
def _on_metrics_collected(mtrcs: metrics.AgentMetrics):
metrics.log_metrics(mtrcs)
Additionally, if you'd like to gather metrics throughout a session and get a summary at the end, the UsageCollector class is a great tool for that:
usage_collector = metrics.UsageCollector()
@agent.on("metrics_collected")
def _on_metrics_collected(mtrcs: metrics.AgentMetrics):
usage_collector.collect(mtrcs)
summary = usage_collector.get_summary()
<http://logger.info|logger.info>(f"Usage: {summary}")
If you have any questions or need a hand with the implementation, please don’t hesitate to ask. I’m here to help!
Sources: agent_extensions/agent_extensions/handlers/wake_word_handler.py | agent-examples/agent_examples/pipeline_agent.py | Logs, metrics, and telemetry | LiveKit Docs | Capturing metrics | LiveKit Docs | agent/main.pyfuture-shampoo-29056
09/25/2025, 9:49 AMfrom livekit.agents import metrics, MetricsCollectedEvent
@session.on("metrics_collected")
def on_metrics_collected(ev: MetricsCollectedEvent):
metrics.log_metrics(ev.metrics)
any agent session will return the logs?tall-airline-16756
09/25/2025, 9:50 AM@session.on("metrics_collected"), the event listener will log metrics for any active agent session tied to that specific session object. Here’s what that means:
• If you’re working with multiple agent sessions, you’ll want to attach the listener to each session individually to capture their metrics.
• The listener only picks up metrics for the session it’s registered with, so it’s session-specific.
If you’re looking to monitor metrics across several sessions, you have a couple of great options:
• Register the listener separately for each session.
• Or, set up a centralized logging or aggregation system to gather metrics from all your sessions in one place.
Feel free to reach out if you'd like a hand setting this up for multiple sessions—I’m here to help!future-shampoo-29056
09/25/2025, 9:51 AMtall-airline-16756
09/25/2025, 9:52 AMmetrics_collected event listener to work, as long as your implementation fits within LiveKit’s framework and emits metrics. Here’s a quick checklist to guide you:
• Emit Metrics: Make sure your custom LLM or TTS emits metrics like ttft, duration, or ttfb. If you’re inheriting from LiveKit’s base classes like tts.TTS or llm.LLM, this usually happens automatically.
• Register the Listener: Don’t forget to attach the metrics_collected listener to your agent session or component. For example:
@session.on("metrics_collected")
def on_metrics_collected(ev):
metrics.log_metrics(ev.metrics)
• Custom Backend: If your LLM or TTS uses a non-standard API, just ensure it integrates smoothly with LiveKit’s metrics system by overriding methods like synthesize() or stream() to capture and emit those important metrics.
Please feel free to reach out if you’d like any help adapting your custom implementation to emit metrics—I’m here to support you!
Sources: Logs, metrics, and telemetry | LiveKit Docs | livekit-agents/livekit/agents/voice/agent_activity.py | agent/main.py | metrics/metrics_vad.py | chain-of-thought-tts/agent.pyfuture-shampoo-29056
09/25/2025, 10:03 AMtall-airline-16756
09/25/2025, 10:03 AMchat() or synthesize() to calculate and emit important metrics such as ttft (Time to First Token) or duration. For example:
from livekit.agents import metrics
start_time = time.monotonic()
# Perform the LLM or TTS operation
end_time = time.monotonic()
# Emit metrics
self.emit_metrics(metrics.LLMMetrics(ttft=end_time - start_time, ...))
• Emit Metrics in Streams: If your implementation uses streams, just make sure to emit metrics at the end of the stream or when key events happen.
• Register the Listener: After you emit metrics, the metrics_collected event listener will automatically pick them up for you.
If you’d like, I’m more than happy to help you step-by-step with adding metrics to your custom implementation—just let me know!future-shampoo-29056
09/25/2025, 10:37 AMtall-airline-16756
09/25/2025, 10:38 AM_session_connected_at attribute, which marks the moment your streaming session begins. However, this detail isn’t available through a public API just yet.
If you’re looking to keep an eye on the session time, a great approach is to add some custom logic in your application. You can record the start time right when the session kicks off and then calculate how much time has passed whenever you need it. Here’s a simple example to guide you:
import time
# Record the start time
session_start_time = time.time()
# Later, calculate the elapsed time
elapsed_time = time.time() - session_start_time
print(f"Session duration: {elapsed_time} seconds")
If you’d like, I’m here to help you implement this or answer any other questions you might have. Just let me know!
Sources: livekit-plugins/livekit-plugins-sarvam/livekit/plugins/sarvam/stt.py | Google Cloud STT integration guide | LiveKit Docsmost-vr-18651
10/27/2025, 2:27 PMtall-airline-16756
10/27/2025, 2:29 PMpip install livekit-plugins-aws
◦ Make sure to set your AWS credentials in your environment variables or pass them directly when you initialize the TTS instance.
◦ Here's how you can initialize the TTS instance:
• *Option 2: Implement a Custom TTS Class*`from livekit.agents import tts
import aiohttp
class CustomAWSTTS(tts.TTS):
def __init__(self, api_endpoint, api_key, voice_id="default", sample_rate=24000):
super().__init__(capabilities=tts.TTSCapabilities(streaming=True), sample_rate=sample_rate, num_channels=1)
self._api_endpoint = api_endpoint
self._api_key = api_key
self._voice_id = voice_id
def synthesize(self, text, conn_options):
return CustomChunkedStream(tts=self, text=text, conn_options=conn_options)
class CustomChunkedStream(tts.ChunkedStream):
async def _run(self, output_emitter):
async with aiohttp.ClientSession() as session:
async with session.post(self._tts._api_endpoint, json={"text": self._input_text, "voice_id": self._tts._voice_id}) as response:
if response.status == 200:
async for chunk in response.content.iter_chunked(8192):
output_emitter.push(chunk)`
◦ If you have your own TTS model deployed on AWS (like on EC2 or Lambda), you can create a custom TTS class like this:
• *Option 3: Integrate with AgentSession*`from livekit.agents.voice import AgentSession
session = AgentSession(
stt=google.STT(),
llm=openai.LLM(),
tts=custom_tts, # Your custom TTS
vad=silero.VAD.load()
)`
◦ After setting up your custom TTS, you can integrate it into an AgentSession like this:
If you run into any bumps or want me to walk you through any part of this, please don’t hesitate to reach out. I’m here to support you every step of the way!
Sources: Amazon Polly TTS plugin guide | LiveKit Docs | examples/voice_agents/error_callback.py | Text-to-speech (TTS) models | LiveKit Docs | Pipeline nodes and hooks | LiveKit Docs