is it possible to stream the response from tts to ...
# ask-ai
f
is it possible to stream the response from tts to the room the make the agent respond faster
t
Thanks for reaching out, Huy. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Absolutely, you can stream TTS (Text-to-Speech) responses directly to a LiveKit room to help agents respond more quickly. It’s a great way to make interactions feel more natural and immediate! This works by using the
stream()
method in LiveKit’s TTS integrations, which lets you synthesize audio in real-time, bit by bit. Here’s a simple overview to guide you through the process: • Initialize TTS and Audio Source: Start by creating a TTS instance along with an
AudioSource
that matches the sample rate and channels. • Publish Audio Track: Then, publish the audio track to the room so everyone can hear the speech as it’s generated. • Stream TTS: Use the
stream()
method to handle the text incrementally, sending audio frames to the room as soon as they’re ready. • Push Text Chunks: Feed chunks of text into the TTS stream—like from a language model—and publish each audio frame in real-time. This method really helps cut down on delays since you don’t have to wait for the entire TTS response before playing it back. If you’d like, I’m here to help with detailed examples or any questions you have along the way! Sources: Text-to-speech (TTS) integrations | LiveKit Docs | agents/src/tts/tts.ts | livekit-plugins/livekit-plugins-spitch/livekit/plugins/spitch/tts.py
f
can you write a custom tts class that inherit the base one but for an api calling tts that return as streaming
t
To help you create a custom TTS class that supports streaming responses from an API, you can extend the base
tts.TTS
class provided by LiveKit. Here’s a friendly, high-level overview of how you might approach this: 1. Define the Custom TTS Class: Start by inheriting from
tts.TTS
and override the
stream()
method so it returns a custom
SynthesizeStream
instance tailored for your needs. 2. Implement the Streaming Logic: Inside your
SynthesizeStream
subclass, you’ll want to connect to the external TTS API using WebSocket or SSE to handle real-time audio streaming smoothly. 3. Emit Audio Frames: As the API sends audio data, process it incrementally and emit these as
AudioFrame
objects to the LiveKit room, ensuring a seamless streaming experience. Here’s a simple example to illustrate the structure:
Copy code
class CustomTTS(tts.TTS):
    def __init__(self, api_key, sample_rate=24000):
        super().__init__(sample_rate, 1, streaming=True)
        self.api_key = api_key

    def stream(self):
        return CustomSynthesizeStream(self)

class CustomSynthesizeStream(tts.SynthesizeStream):
    def __init__(self, tts):
        super().__init__(tts)

    async def run(self):
        # Connect to API and handle streaming
        pass
This method gives you the flexibility to integrate with any API that supports streaming TTS responses. If you’d like, I’m here to help with a more detailed implementation or any specific questions you might have—just let me know! Sources: plugins/openai/src/tts.ts
f
if im trying to make the flow faster :stt,llm,tts i already did streaming what else can i do (for example pre warm something)
t
I'd love to help you optimize your STT, LLM, and TTS flow for better and faster performance! To get a clearer picture, could you please share a bit more about your current setup? For instance: • Are you working with any specific APIs or frameworks for STT, LLM, and TTS? • What kind of latency are you noticing right now? • Are there particular parts of the flow—like STT processing, LLM response generation, or TTS synthesis—that seem slower than you'd expect? Once I have these details, I’ll be able to suggest some targeted improvements, like pre-warming techniques or other optimizations, to help speed things up for you.
f
i create my own api for all 3 then write the custom stt,lmm,tts for each of them
t
To help you get the most out of your custom STT, LLM, and TTS APIs and make them run faster, here are some friendly tips you might find useful: • Reduce Latency: ◦ Try deploying your services close to each other—that way, you can cut down on network delays. ◦ Consider using providers known for low-latency inference, like Groq for LLMs or Google Cloud for STT/TTS, to speed things up. ◦ Processing audio in small chunks (around 20ms) can be a great way to balance speed and accuracy. • Increase Throughput: ◦ Enabling parallel STT and TTS operations can really help when you’re handling multiple streams or quick interactions. ◦ Streaming LLM tokens incrementally and feeding them into the TTS pipeline as they come can save you from waiting for the full response. • Pre-Warming Techniques: ◦ Keeping your STT, LLM, and TTS models loaded and ready means you won’t have to wait for initialization each time. ◦ Caching commonly used phrases or responses can also speed things along nicely. • Monitor Performance: ◦ Keeping an eye on metrics like Time To First Token (TTFT), Time To First Byte (TTFB), and overall latency can help you spot where things might slow down. I hope these ideas help you achieve faster and smoother performance! If you’d like, I’m here to guide you through any of these steps in more detail—just let me know! Sources: Live Conversations with AI using ChatGPT and WebRTC | Speech & Audio Processing | livekit/livekit_composite | DeepWiki | README.md | An open source stack for real-time multimodal AI
f
is there any format when it come to streaming tts
t
When it comes to streaming TTS, the format and implementation can vary depending on the provider and the specific needs of your application. Let’s walk through some key points together: • Streaming Protocol: Most TTS providers use WebSocket for real-time streaming. For example, Cartesia, Neuphonic, and Resemble AI utilize WebSocket connections to send text bit by bit and receive audio frames live. • Audio Format: The audio output is usually in raw PCM format (like
pcm_s16le
) or compressed formats such as OGG or MP3, depending on the provider. LiveKit helps by automatically resampling audio to fit WebRTC standards (like 24kHz). • Simulated Streaming: If a provider doesn’t support streaming directly, LiveKit uses a
StreamAdapter
to break text into chunks and process them one after another, creating a streaming-like experience. • Real-Time Metrics: Streaming setups often track metrics such as Time to First Byte (TTFB) and synthesis duration to keep an eye on performance. If you’re building your own streaming TTS, just make sure your API can handle incremental text input and audio output. Feel free to reach out if you’d like specific examples or need any more help—I’m here to support you! Sources: livekit-plugins/livekit-plugins-inworld/livekit/plugins/inworld/tts.py | agents/src/tts/stream_adapter.ts
f
how about openai with streaming =True
t
I understand you’re looking to create a streaming experience with OpenAI TTS. While it doesn’t natively support
streaming=True
, there’s a helpful way to simulate streaming using LiveKit’s
StreamAdapter
. This tool breaks your text into smaller chunks, like sentences, and processes them step-by-step, giving you that nice streaming feel. Here’s a simple way to set it up: • Wrap OpenAI TTS: Use
StreamAdapter
to enable incremental audio synthesis. • Push Text: Send your text chunks to the adapter, which handles them one at a time. • Consume Audio: Receive audio frames as they’re generated, almost like real-time streaming. Here’s an example to guide you:
Copy code
from livekit.agents.tts import StreamAdapter
from livekit.agents.tokenize import SentenceTokenizer

streaming_tts = StreamAdapter(tts=openai_tts, sentence_tokenizer=SentenceTokenizer())
stream = streaming_tts.stream()
stream.push_text("Hello world.")
stream.flush()
This method helps reduce the delay you might notice and works smoothly with real-time applications. If you have any questions or want me to walk you through it, just let me know—I’m here to help! Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/stt.py | agents/src/tts/tts.ts | Tool definition and use | LiveKit Docs
f
what do i need to change if i want tts to be streaming? from kittentts import KittenTTS import numpy as np from livekit.agents import tts from livekit.agents.types import DEFAULT_API_CONNECT_OPTIONS class KittenTTSPlugin(tts.TTS): def __init__(self, model_name="KittenML/kitten-tts-nano-0.1", voice="expr-voice-2-f", sample_rate=24000): # Initialize parent with required parameters (matching OpenAI TTS pattern) super().__init__( capabilities=tts.TTSCapabilities(streaming=False), sample_rate=sample_rate, num_channels=1, # Mono audio ) self.model = KittenTTS(model_name) self.voice = voice # Verify voice availability if voice not in self.model.available_voices: raise ValueError(f"Voice {voice} not available. Choose from {self.model.available_voices}") def synthesize(self, text: str, *, conn_options=DEFAULT_API_CONNECT_OPTIONS) -> tts.ChunkedStream: """ Synthesize text to speech using KittenTTS. Args: text: The text to synthesize conn_options: Connection options (ignored for KittenTTS) Returns: ChunkedStream: Stream of synthesized audio chunks """ return KittenTTSChunkedStream(tts=self, input_text=text, conn_options=conn_options) @property def name(self): return f"KittenTTS-{self.voice}" class KittenTTSChunkedStream(tts.ChunkedStream): def __init__(self, *, tts: KittenTTSPlugin, input_text: str, conn_options): super().__init__(tts=tts, input_text=input_text, conn_options=conn_options) self._tts: KittenTTSPlugin = tts self._audio_generated = False async def _run(self, output_emitter: tts.AudioEmitter): """ Generate audio using KittenTTS and emit it through the output emitter. This is the core method that must be implemented. """ if self._audio_generated: return try: # Add padding to avoid cut-offs padded_text = self.input_text.strip() + " ... " # Generate audio waveform (numpy float32 array) audio = self._tts.model.generate(padded_text, voice=self._tts.voice) # Convert float32 to int16 (LiveKit requirement) # Scale from [-1, 1] to [-32768, 32767] audio_int16 = (audio * 32767).astype(np.int16) # Convert to bytes audio_bytes = audio_int16.tobytes() # Initialize the output emitter with PCM format output_emitter.initialize( request_id=str(id(self)), sample_rate=self._tts.sample_rate, num_channels=1, # Mono mime_type="audio/pcm", # Use PCM format for raw audio ) # Push the audio data through the emitter output_emitter.push(audio_bytes) # Flush the emitter to indicate completion output_emitter.flush() self._audio_generated = True except Exception as e: # Handle any errors during synthesis raise RuntimeError(f"KittenTTS synthesis failed: {str(e)}") finally: # Ensure the emitter is properly closed try: output_emitter.end_input() except: pass
t
To help you modify your KittenTTS plugin to support streaming, here are some friendly steps you can follow: • Enable Streaming Capability: In the constructor, update the
capabilities
by setting
streaming=True
. This will let your plugin know streaming is supported:
Copy code
super().__init__(
    capabilities=tts.TTSCapabilities(streaming=True),
    sample_rate=sample_rate,
    num_channels=1,
)
• Implement the
stream()
Method:
Add a
stream()
method that returns a
SynthesizeStream
instance. This is key for streaming functionality:
Copy code
def stream(self, *, conn_options=DEFAULT_API_CONNECT_OPTIONS):
    return KittenTTSSynthesizeStream(tts=self, conn_options=conn_options)
• Create a Streaming Class: You’ll want to implement a
KittenTTSSynthesizeStream
class inheriting from
tts.SynthesizeStream
. This class will manage real-time text-to-speech synthesis by: ◦ Initializing the
AudioEmitter
to handle streaming audio. ◦ Sending text chunks to your TTS model step-by-step. ◦ Emitting audio frames as soon as they’re generated, so the streaming feels smooth. Here’s a helpful example of the
_run()
method inside your streaming class:
Copy code
async def _run(self, output_emitter):
    output_emitter.initialize(
        request_id=str(id(self)),
        sample_rate=self._tts.sample_rate,
        num_channels=1,
        mime_type="audio/pcm",
        stream=True,
    )

    for text_chunk in self.input_text.split(" "):
        audio = self._tts.model.generate(text_chunk, voice=self._tts.voice)
        audio_bytes = (audio * 32767).astype(np.int16).tobytes()
        output_emitter.push(audio_bytes)

    output_emitter.end_input()
These updates will enable your TTS plugin to support real-time streaming beautifully. If you have any questions or want me to walk you through any part, just let me know—I’m here to help! Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/tts.py
f
what if the output of my tts is not streaming what would happens then
t
I understand that your TTS output might not be streaming, and that's completely okay! LiveKit takes care of this by using the
synthesize()
method, which processes the entire text input and generates audio frames all at once. Here’s a friendly breakdown of what happens: • Non-Streaming Behavior: The TTS engine creates the full audio output for your text, and the audio frames come back as an asynchronous stream. • Simulated Streaming: If you’d like a streaming-like experience, you can use a
StreamAdapter
to split your text into smaller chunks (like sentences) and process them step-by-step, which mimics real-time streaming nicely. • Aligned Transcripts: By enabling
use_tts_aligned_transcript=True
, you can get sentence-level timestamps that help sync the transcript perfectly with the audio playback. • Text Preprocessing: You can also use the
before_tts_cb
callback to clean up or adjust your text before synthesis, making sure the spoken output sounds just right. This way, even if the TTS isn’t streaming by default, it can still fit smoothly into your real-time applications. If you want, I’m here to help you set these up or answer any questions you have! Sources: agents/src/tts/tts.ts | examples/voice_agents/timed_agent_transcript.py | chain-of-thought-tts/README.mdoc | Pipeline nodes and hooks | LiveKit Docs
f
how about how to check time used for each process (stt , llm. tts)
t
I'm glad you're interested in measuring the time used for each process (STT, LLM, TTS) in LiveKit! You can easily do this using the built-in metrics system. Here’s a friendly guide to help you track timing for each stage: • STT (Speech-to-Text): ◦ `duration`: This tells you the time taken to create the transcript (for non-streaming STT). ◦ `audio_duration`: This shows the duration of the audio input that was processed. • LLM (Large Language Model): ◦
ttft
(Time to First Token): This measures the time to generate the very first token. ◦ `duration`: This is the total time taken to generate the entire response. • TTS (Text-to-Speech): ◦
ttfb
(Time to First Byte): This indicates the time to generate the first byte of audio. ◦ `duration`: This covers the total time for the TTS request. ◦ `audio_duration`: This reflects the duration of the generated audio output. To get these metrics, just subscribe to the
metrics_collected
event in your agent code. From there, you can log them or even forward them to external systems for a deeper analysis. If you'd like, I’m here to help you implement this step-by-step—just let me know! Sources: Logs, metrics, and telemetry | LiveKit Docs | metrics/metrics_stt.py | agents/src/metrics/base.ts | livekit-agents/livekit/agents/voice/agent.py | Capturing metrics | LiveKit Docs
f
how to subsrcibe to the metrics coleected event
t
I'm happy to guide you through subscribing to the
metrics_collected
event in LiveKit! You can easily register an event listener on your agent session or voice pipeline agent. Here’s a simple way to do it: • For AgentSession:
Copy code
from livekit.agents import metrics, MetricsCollectedEvent

@session.on("metrics_collected")
def on_metrics_collected(ev: MetricsCollectedEvent):
    metrics.log_metrics(ev.metrics)
• For VoicePipelineAgent:
Copy code
from livekit.agents import metrics

@agent.on("metrics_collected")
def _on_metrics_collected(mtrcs: metrics.AgentMetrics):
    metrics.log_metrics(mtrcs)
Additionally, if you'd like to gather metrics throughout a session and get a summary at the end, the
UsageCollector
class is a great tool for that:
Copy code
usage_collector = metrics.UsageCollector()

@agent.on("metrics_collected")
def _on_metrics_collected(mtrcs: metrics.AgentMetrics):
    usage_collector.collect(mtrcs)

summary = usage_collector.get_summary()
<http://logger.info|logger.info>(f"Usage: {summary}")
If you have any questions or need a hand with the implementation, please don’t hesitate to ask. I’m here to help! Sources: agent_extensions/agent_extensions/handlers/wake_word_handler.py | agent-examples/agent_examples/pipeline_agent.py | Logs, metrics, and telemetry | LiveKit Docs | Capturing metrics | LiveKit Docs | agent/main.py
f
so when i do
Copy code
from livekit.agents import metrics, MetricsCollectedEvent

@session.on("metrics_collected")
def on_metrics_collected(ev: MetricsCollectedEvent):
    metrics.log_metrics(ev.metrics)
any agent session will return the logs?
t
Absolutely! When you use
@session.on("metrics_collected")
, the event listener will log metrics for any active agent session tied to that specific
session
object. Here’s what that means: • If you’re working with multiple agent sessions, you’ll want to attach the listener to each session individually to capture their metrics. • The listener only picks up metrics for the session it’s registered with, so it’s session-specific. If you’re looking to monitor metrics across several sessions, you have a couple of great options: • Register the listener separately for each session. • Or, set up a centralized logging or aggregation system to gather metrics from all your sessions in one place. Feel free to reach out if you'd like a hand setting this up for multiple sessions—I’m here to help!
f
if i write a custom llm or custom tts do i need to add antything in order for the event listener to listen
t
I’m glad you’re working on a custom LLM or TTS! You actually don’t need to add anything extra for the
metrics_collected
event listener to work, as long as your implementation fits within LiveKit’s framework and emits metrics. Here’s a quick checklist to guide you: • Emit Metrics: Make sure your custom LLM or TTS emits metrics like
ttft
,
duration
, or
ttfb
. If you’re inheriting from LiveKit’s base classes like
tts.TTS
or
llm.LLM
, this usually happens automatically. • Register the Listener: Don’t forget to attach the
metrics_collected
listener to your agent session or component. For example:
Copy code
@session.on("metrics_collected")
def on_metrics_collected(ev):
    metrics.log_metrics(ev.metrics)
• Custom Backend: If your LLM or TTS uses a non-standard API, just ensure it integrates smoothly with LiveKit’s metrics system by overriding methods like
synthesize()
or
stream()
to capture and emit those important metrics. Please feel free to reach out if you’d like any help adapting your custom implementation to emit metrics—I’m here to support you! Sources: Logs, metrics, and telemetry | LiveKit Docs | livekit-agents/livekit/agents/voice/agent_activity.py | agent/main.py | metrics/metrics_vad.py | chain-of-thought-tts/agent.py
f
it doesnt have metrics from future import annotations import aiohttp import asyncio import json import uuid from dataclasses import dataclass from typing import Any from livekit.agents import ( APIConnectionError, APIStatusError, APITimeoutError, llm, ) from livekit.agents.types import DEFAULT_API_CONNECT_OPTIONS, APIConnectOptions, NOT_GIVEN, NotGivenOr from livekit.agents.llm.chat_context import ChatContext, ChatRole from livekit.agents.llm.tool_context import FunctionTool, RawFunctionTool, ToolChoice @dataclass class _LLMOptions: api_base_url: str api_key: str model: str temperature: float max_tokens: int class CustomLLM(llm.LLM): def __init__( self, *, _api_base_url_: str = "https://api.openai.com/v1", _api_key_: str | None = None, _model_: str = "gpt-4o-mini", _temperature_: float = 0.7, _max_tokens_: int = 1000, ) -> None: """ Create a new instance of Custom LLM that makes HTTP calls to GPT-4o mini. Args: api_base_url: Base URL of the OpenAI API api_key: OpenAI API key model: Model name to use temperature: Sampling temperature max_tokens: Maximum tokens in response """ super().__init__() self._opts = _LLMOptions( _api_base_url_=_api_base_url_.rstrip('/'), _api_key_=_api_key_ or "", model=model, temperature=temperature, _max_tokens_=_max_tokens_, ) # Create HTTP client with proper timeout settings self._client_session: aiohttp.ClientSession | None = None async def _get_client_session(self) -> aiohttp.ClientSession: """Get or create HTTP client session""" if self._client_session is None or self._client_session.closed: timeout = aiohttp.ClientTimeout( total=60.0, connect=15.0, _sock_read_=30.0, _sock_connect_=5.0 ) self._client_session = aiohttp.ClientSession( timeout=timeout, headers={ "Authorization": f"Bearer {self._opts.api_key}", "Content-Type": "application/json", }, connector=aiohttp.TCPConnector( limit=50, _limit_per_host_=50, _keepalive_timeout_=120, _enable_cleanup_closed_=True ) ) return self._client_session @property def model(self) -> str: return self._opts.model def chat( self, *, _chat_ctx_: ChatContext, _tools_: list[FunctionTool | RawFunctionTool] | None = None, _conn_options_: APIConnectOptions = DEFAULT_API_CONNECT_OPTIONS, _parallel_tool_calls_: NotGivenOr[bool] = NOT_GIVEN, _tool_choice_: NotGivenOr[ToolChoice] = NOT_GIVEN, _extra_kwargs_: NotGivenOr[dict[str, Any]] = NOT_GIVEN, ) -> llm.LLMStream: """ Create a chat completion stream. """ return CustomLLMStream( llm=self, _chat_ctx_=_chat_ctx_, tools=tools or [], _conn_options_=_conn_options_, _parallel_tool_calls_=_parallel_tool_calls_, _tool_choice_=_tool_choice_, _extra_kwargs_=_extra_kwargs_, ) async def aclose(self) -> None: """Close the LLM instance and cleanup resources""" if self._client_session and not self._client_session.closed: await self._client_session.close() self._client_session = None class CustomLLMStream(llm.LLMStream): def __init__( self, _llm_: CustomLLM, *, _chat_ctx_: ChatContext, _tools_: list[FunctionTool | RawFunctionTool], _conn_options_: APIConnectOptions, _parallel_tool_calls_: NotGivenOr[bool] = NOT_GIVEN, _tool_choice_: NotGivenOr[ToolChoice] = NOT_GIVEN, _extra_kwargs_: NotGivenOr[dict[str, Any]] = NOT_GIVEN, ) -> None: super().__init__(llm=llm, _chat_ctx_=_chat_ctx_, tools=tools, _conn_options_=_conn_options_) self._llm: CustomLLM = llm self._parallel_tool_calls = _parallel_tool_calls_ self._tool_choice = _tool_choice_ self._extra_kwargs = _extra_kwargs_ def _convert_chat_context_to_messages(self, _chat_ctx_: ChatContext) -> list[dict[str, Any]]: """Convert ChatContext to OpenAI messages format""" _try_: # Use the official method like OpenAI LLM does messages, _ = _chat_ctx_.to_provider_format(format="openai") return messages except Exception: # Fallback: Ensure we always have at least a system message return [ { "role": "system", "content": "You are a helpful voice AI assistant that answers questions and helps with tasks in japanese." } ] async def _run(self) -> None: """ Make HTTP request to GPT-4o mini API (non-streaming). """ # Convert chat context to OpenAI format messages = self._convert_chat_context_to_messages(self._chat_ctx) # Prepare the API request payload (no streaming) payload = { "model": self._llm._opts.model, "messages": messages, "temperature": self._llm._opts.temperature, "max_tokens": self._llm._opts.max_tokens, "stream": False, # Disable streaming for simplicity } # Add tools if provided if self._tools: # Convert tools to OpenAI format (simplified) tools_payload = [] for tool in self._tools: if isinstance(tool, FunctionTool): tools_payload.append({ "type": "function", "function": { "name": tool.name, "description": tool.description or "", "parameters": tool.parameters or {}, } }) if tools_payload: payload["tools"] = tools_payload _try_: # Get HTTP client session session = await self._llm._get_client_session() # Make the non-streaming API request async with session.post( f"{self._llm._opts.api_base_url}/chat/completions", json=payload, timeout=aiohttp.ClientTimeout( total=60, connect=self._conn_options.timeout ) ) as response: if response.status != 200: error_text = await response.text() raise APIStatusError( f"Custom LLM API error: {error_text}", _status_code_=response.status, _request_id_=str(response.headers.get('request-id', '')), body=error_text ) # Get the complete response data = await response.json() request_id = str(uuid.uuid4()) # Extract the response content if 'choices' in data and len(data['choices']) > 0: choice = data['choices'][0] message = choice.get('message', {}) content = message.get('content', '') # Send the complete response as a single chunk chunk_delta = llm.ChoiceDelta( role="assistant", content=content, _tool_calls_=[] ) chunk = llm.ChatChunk( id=request_id, delta=chunk_delta, usage=None # Usage info if needed ) # Send the chunk self._event_ch.send_nowait(chunk) except (aiohttp.ClientError, asyncio.TimeoutError) as e: raise APITimeoutError() from None except APIStatusError: raise # Re-raise our custom status errors except Exception as e: raise APIConnectionError() from e
t
I understand that if your custom LLM or TTS implementation doesn’t currently emit metrics, you’ll need to add support for metrics collection yourself. Don’t worry, I’m here to guide you through it! Here’s a simple way to integrate metrics into your custom LLM: • Emit Metrics: You can override methods like
chat()
or
synthesize()
to calculate and emit important metrics such as
ttft
(Time to First Token) or
duration
. For example:
Copy code
from livekit.agents import metrics

start_time = time.monotonic()
# Perform the LLM or TTS operation
end_time = time.monotonic()

# Emit metrics
self.emit_metrics(metrics.LLMMetrics(ttft=end_time - start_time, ...))
• Emit Metrics in Streams: If your implementation uses streams, just make sure to emit metrics at the end of the stream or when key events happen. • Register the Listener: After you emit metrics, the
metrics_collected
event listener will automatically pick them up for you. If you’d like, I’m more than happy to help you step-by-step with adding metrics to your custom implementation—just let me know!
f
if i use google plugin for stt how can i get the time session = AgentSession( stt = google.STT( model="telephony", languages="ja-JP", _spoken_punctuation_=False, ),
t
I wanted to share a little insight about the Google STT plugin in LiveKit. It keeps track of the session connection time internally using the
_session_connected_at
attribute, which marks the moment your streaming session begins. However, this detail isn’t available through a public API just yet. If you’re looking to keep an eye on the session time, a great approach is to add some custom logic in your application. You can record the start time right when the session kicks off and then calculate how much time has passed whenever you need it. Here’s a simple example to guide you:
Copy code
import time

# Record the start time
session_start_time = time.time()

# Later, calculate the elapsed time
elapsed_time = time.time() - session_start_time
print(f"Session duration: {elapsed_time} seconds")
If you’d like, I’m here to help you implement this or answer any other questions you might have. Just let me know! Sources: livekit-plugins/livekit-plugins-sarvam/livekit/plugins/sarvam/stt.py | Google Cloud STT integration guide | LiveKit Docs
m
@future-shampoo-29056 were you able to implement the tts ? i want tp implement my own custom tts deployed on aws. How do i do that ?
t
I'm here to help you implement your own custom TTS deployed on AWS in LiveKit! You can follow these steps, and if you have any questions along the way, just let me know: • *Option 1: Use Amazon Polly TTS Plugin*`from livekit.plugins.aws import TTS tts = TTS( voice="Ruth", language="en-US", speech_engine="generative", sample_rate=16000, region="us-east-1" )` ◦ First, install the plugin by running:
pip install livekit-plugins-aws
◦ Make sure to set your AWS credentials in your environment variables or pass them directly when you initialize the TTS instance. ◦ Here's how you can initialize the TTS instance: • *Option 2: Implement a Custom TTS Class*`from livekit.agents import tts import aiohttp class CustomAWSTTS(tts.TTS): def __init__(self, api_endpoint, api_key, voice_id="default", sample_rate=24000): super().__init__(capabilities=tts.TTSCapabilities(streaming=True), sample_rate=sample_rate, num_channels=1) self._api_endpoint = api_endpoint self._api_key = api_key self._voice_id = voice_id def synthesize(self, text, conn_options): return CustomChunkedStream(tts=self, text=text, conn_options=conn_options) class CustomChunkedStream(tts.ChunkedStream): async def _run(self, output_emitter): async with aiohttp.ClientSession() as session: async with session.post(self._tts._api_endpoint, json={"text": self._input_text, "voice_id": self._tts._voice_id}) as response: if response.status == 200: async for chunk in response.content.iter_chunked(8192): output_emitter.push(chunk)` ◦ If you have your own TTS model deployed on AWS (like on EC2 or Lambda), you can create a custom TTS class like this: • *Option 3: Integrate with AgentSession*`from livekit.agents.voice import AgentSession session = AgentSession( stt=google.STT(), llm=openai.LLM(), tts=custom_tts, # Your custom TTS vad=silero.VAD.load() )` ◦ After setting up your custom TTS, you can integrate it into an
AgentSession
like this: If you run into any bumps or want me to walk you through any part of this, please don’t hesitate to reach out. I’m here to support you every step of the way! Sources: Amazon Polly TTS plugin guide | LiveKit Docs | examples/voice_agents/error_callback.py | Text-to-speech (TTS) models | LiveKit Docs | Pipeline nodes and hooks | LiveKit Docs