how to only change pronunciation when sending chun...
# ask-ai
b
how to only change pronunciation when sending chunks to tts in tts node ? I dont want the transcript to be changed, only the chunk being sent to tts will replaced with my custom pronunciation dictionary ?
t
Thanks for reaching out, Mustafa. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I'm happy to help you with changing pronunciation when sending chunks to the TTS node without altering the transcript! You can achieve this by preprocessing the text using either the before_tts_cb callback (in
VoicePipelineAgent
) or by overriding the
tts_node
method in a custom agent. These approaches let you fine-tune pronunciation while keeping the original transcript intact. Method 1: Using
before_tts_cb
This callback is a great way to modify text just before it goes to the TTS engine. Here's a simple example in Python:
Copy code
from livekit.agents import tokenize
from livekit.agents.pipeline import VoicePipelineAgent

def replace_words(assistant: VoicePipelineAgent, text: str | AsyncIterable[str]):
    return tokenize.utils.replace_words(
        text=text, replacements={"livekit": r"<<l|aɪ|v|k|ɪ|t|>>"}
    )

agent = VoicePipelineAgent(
    before_tts_cb=replace_words,
)
Method 2: Overriding
tts_node
If you want more control, you can override the
tts_node
method to process text chunks before they are sent to TTS. Here's how you might do that in Python:
Copy code
class MyAgent(Agent):
    async def tts_node(self, text: AsyncIterable[str], model_settings: ModelSettings) -> AsyncIterable[AudioFrame]:
        pronunciations = {"LiveKit": "Live Kit"}

        async def adjust_pronunciation(input_text: AsyncIterable[str]) -> AsyncIterable[str]:
            async for chunk in input_text:
                modified_chunk = chunk
                for term, pronunciation in pronunciations.items():
                    modified_chunk = re.sub(rf'\b{term}\b', pronunciation, modified_chunk, flags=re.IGNORECASE)
                yield modified_chunk

        async for frame in Agent.default.tts_node(self, adjust_pronunciation(text), model_settings):
            yield frame
These techniques ensure your original transcript stays just as you want it, while giving you precise control over how things sound when spoken. If you have any questions or want to explore this further, please don’t hesitate to ask—I’m here to help! Sources: VoicePipelineAgent | LiveKit Docs | VoicePipelineAgent | LiveKit Docs | Agent speech and audio | LiveKit Docs | Agent speech and audio | LiveKit Docs | Docs::Agents