boundless-painting-47811
09/23/2025, 12:31 PMtall-airline-16756
09/23/2025, 12:31 PMtall-airline-16756
09/23/2025, 12:31 PMVoicePipelineAgent) or by overriding the tts_node method in a custom agent. These approaches let you fine-tune pronunciation while keeping the original transcript intact.
Method 1: Using before_tts_cb
This callback is a great way to modify text just before it goes to the TTS engine. Here's a simple example in Python:
from livekit.agents import tokenize
from livekit.agents.pipeline import VoicePipelineAgent
def replace_words(assistant: VoicePipelineAgent, text: str | AsyncIterable[str]):
return tokenize.utils.replace_words(
text=text, replacements={"livekit": r"<<l|aɪ|v|k|ɪ|t|>>"}
)
agent = VoicePipelineAgent(
before_tts_cb=replace_words,
)
Method 2: Overriding tts_node
If you want more control, you can override the tts_node method to process text chunks before they are sent to TTS. Here's how you might do that in Python:
class MyAgent(Agent):
async def tts_node(self, text: AsyncIterable[str], model_settings: ModelSettings) -> AsyncIterable[AudioFrame]:
pronunciations = {"LiveKit": "Live Kit"}
async def adjust_pronunciation(input_text: AsyncIterable[str]) -> AsyncIterable[str]:
async for chunk in input_text:
modified_chunk = chunk
for term, pronunciation in pronunciations.items():
modified_chunk = re.sub(rf'\b{term}\b', pronunciation, modified_chunk, flags=re.IGNORECASE)
yield modified_chunk
async for frame in Agent.default.tts_node(self, adjust_pronunciation(text), model_settings):
yield frame
These techniques ensure your original transcript stays just as you want it, while giving you precise control over how things sound when spoken. If you have any questions or want to explore this further, please don’t hesitate to ask—I’m here to help!
Sources: VoicePipelineAgent | LiveKit Docs | VoicePipelineAgent | LiveKit Docs | Agent speech and audio | LiveKit Docs | Agent speech and audio | LiveKit Docs | Docs::Agents