how do i modify the POST TTS audio frames in my vo...
# ask-ai
c
how do i modify the POST TTS audio frames in my voice agent pipeline?
t
Thanks for reaching out, Faraz. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I'm glad you're exploring how to modify POST TTS audio frames in your voice agent pipeline! To achieve this, you’ll want to override the
tts_node
method in your custom
Agent
class. This approach lets you apply your own audio processing to the synthesized audio frames right after TTS synthesis but before they’re outputted. Here’s a friendly example in Python that speeds up the output audio by tweaking the audio frames:
Copy code
from livekit import rtc
from livekit.agents import ModelSettings, Agent
import numpy as np
import librosa
from typing import AsyncIterable

class MyAgent(Agent):
    async def tts_node(self, text: AsyncIterable[str], model_settings: ModelSettings):
        # Get the default TTS audio stream
        audio_stream = Agent.default.tts_node(self, text, model_settings)
        
        # Apply custom post-processing
        return self._process_audio_stream(audio_stream)

    async def _process_audio_stream(self, audio: AsyncIterable[rtc.AudioFrame]):
        stream = None
        async for frame in audio:
            if stream is None:
                stream = utils.audio.AudioByteStream(
                    sample_rate=frame.sample_rate,
                    num_channels=frame.num_channels,
                    samples_per_channel=frame.sample_rate // 10,  # 100ms
                )
            # Process frame (e.g., speed up)
            for f in stream.push(frame.data):
                yield self._process_audio(f)

        if stream is not None:
            for f in stream.flush():
                yield self._process_audio(f)

    def _process_audio(self, frame: rtc.AudioFrame) -> rtc.AudioFrame:
        # Time-stretch audio without changing pitch
        audio_data = np.frombuffer(frame.data, dtype=np.int16)
        stretched = librosa.effects.time_stretch(
            audio_data.astype(np.float32) / np.iinfo(np.int16).max,
            rate=1.2  # Speed up by 20%
        )
        stretched = (stretched * np.iinfo(np.int16).max).astype(np.int16)
        return rtc.AudioFrame(
            data=stretched.tobytes(),
            sample_rate=frame.sample_rate,
            num_channels=frame.num_channels,
            samples_per_channel=stretched.shape[-1],
        )
If you’re working with Node.js, you can similarly modify the
ttsNode
method to process audio frames. Please don’t hesitate to reach out if you’d like me to provide further assistance or examples in another language—I’m here to help! Sources: Agent speech and audio | LiveKit Docs | Agents v0.x migration guide - Node.js | LiveKit Docs | Agents 0.x migration guide | LiveKit Docs | Docs::Agents | livekit-agents/livekit/agents/voice/agent.py