How to improve TTS naturalness by controlling text...
# ask-ai
j
How to improve TTS naturalness by controlling text chunking in AgentSession? We're experiencing stilted/unnatural TTS speech with our voice agent using AgentSession. The issue appears to be that text is being sent to TTS in very small chunks, causing unnatural pauses and intonation. Current setup: session = AgentSession( llm=anthropic.LLM(model="claude-sonnet-4-20250514"), tts=deepgram.TTS(model="aura-2-cora-en"), # ... other settings ) What we're seeing in logs: Text is chunked unnaturally: ' course! I hear' ' you. I' ' without adding those' ' pauses an' '\n\nWhat' This creates: 1. Unnatural pauses between chunks 2. Stilted speech patterns 3. Deepgram STT interpreting fragments as "DOT" Questions: 1. Does AgentSession support sentence_tokenizer parameter like VoicePipelineAgent does? 2. If not, what's the recommended way to buffer text into complete sentences before sending to TTS? 3. Is there a way to control the LLM streaming chunk size to get more natural sentence boundaries? 4. Should we migrate to VoicePipelineAgent for better sentence handling, or is there a way to achieve this with AgentSession? We need the text to be buffered into complete sentences before TTS synthesis to create natural-sounding speech. Using livekit-agents 1.2.2
t
Thanks for reaching out, Albert. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To improve TTS naturalness and control text chunking in AgentSession, here are the key points and recommendations: 1. Sentence Tokenization in AgentSession: AgentSession does not directly support a
sentence_tokenizer
parameter like the deprecated VoicePipelineAgent. However, you can achieve similar functionality by customizing the
tts_node
to buffer text into complete sentences before sending it to TTS. 2. Customizing Text Chunking: You can override the
tts_node
in your AgentSession to implement custom text buffering and sentence chunking. Below is an example:
Copy code
from livekit.agents.voice import Agent
from livekit.agents import tokenize

class SentenceBufferedAgent(Agent):
    async def tts_node(self, text: AsyncIterable[str], model_settings: ModelSettings):
        async def process_with_sentence_buffering():
            buffer = ""
            sentence_tokenizer = tokenize.blingfire.SentenceTokenizer(
                min_sentence_len=15,
                retain_format=False
            )

            async for chunk in text:
                buffer += chunk

                # Tokenize accumulated text
                sentences = sentence_tokenizer.tokenize(buffer)

                # Yield complete sentences, keep incomplete text in buffer
                if len(sentences) > 1:
                    for sentence in sentences[:-1]:
                        yield sentence + " "
                    buffer = sentences[-1]  # Keep last incomplete sentence

            # Yield any remaining text
            if buffer.strip():
                yield buffer

        return Agent.default.tts_node(self, process_with_sentence_buffering(), model_settings)
This approach ensures that text is buffered into complete sentences before being sent to TTS, reducing unnatural pauses and improving speech naturalness. 3. Available Sentence Tokenizers: You can use various sentence tokenizers, such as: • Blingfire Tokenizer: Provides options for minimum sentence length, stream context length, and format retention. • NLTK Tokenizer: Uses the punkt tokenizer with configurable parameters. 4. Migration to VoicePipelineAgent: Since VoicePipelineAgent is deprecated, it is recommended to adapt your implementation to the new AgentSession architecture using the customization options mentioned above. By implementing these changes, you can achieve natural-sounding TTS output with improved sentence handling. Sources: Agents v0.x migration guide | LiveKit Docs | AgentTranscriptionOptions | LiveKit Agents | livekit-agents/livekit/agents/tokenize/blingfire.py | Voice pipeline nodes | LiveKit Docs | pipeline-tts/short_replies_only.py