I'm using a LiveKit voice agent with Deepgra...
# ask-ai
j
I'm using a LiveKit voice agent with Deepgram TTS. My LLM (Claude) sometimes returns stage directions in its responses like *thoughtful pause*, sighs, or (thinking). These are being read aloud by Deepgram TTS, which sounds unnatural. What's the best practice for filtering these out before sending to TTS? Should I: 1. Implement filtering in the Python agent before passing to Deepgram TTS? 2. Use a LiveKit plugin or middleware feature? 3. Configure Deepgram TTS to ignore certain patterns? My current setup: - Python LiveKit agent using livekit-agents SDK - Deepgram TTS with model aura-2 - Claude LLM for responses Code snippet of my current TTS setup: session = AgentSession( llm=anthropi c.LLM(model="cla ude-sonnet-4-202 50514"), tts=deepgram .TTS(model="aura -2"), # ... other config )
t
Thanks for reaching out, Albert. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To filter out stage directions or non-verbal cues like thoughtful pause or sighs from LLM responses before sending them to Deepgram TTS, here are the best practices for your LiveKit Python agent setup: 1. Use the
before_tts_cb
Callback:
This is the simplest and most effective method. You can define a callback function to process and clean the text before it is passed to the TTS engine. For example:
Copy code
async def _before_tts_cb(agent: VoicePipelineAgent, text: str | AsyncIterable[str]):
    if isinstance(text, str):
        # Handle non-streaming text
        result = text.replace("*thoughtful pause*", "")
        return result
    else:
        # Handle streaming text
        async def process_stream():
            async for chunk in text:
                processed = chunk.replace("*thoughtful pause*", "")
                yield processed
        return process_stream()

agent = VoicePipelineAgent(
    vad=ctx.proc.userdata["vad"],
    stt=openai.STT(),
    llm=openai.LLM.with_groq(model="deepseek-r1-distill-llama-70b"),
    tts=deepgram.TTS(model="aura-2"),
    before_tts_cb=_before_tts_cb,
    chat_ctx=initial_ctx
)
1. Override the
llm_node
Method:
For more control, you can create a custom agent class and filter the content as it streams from the LLM. This method is useful for real-time filtering of streaming text.
Copy code
class ChainOfThoughtAgent(Agent):
    async def llm_node(self, chat_ctx, tools, model_settings=None):
        async def process_stream():
            async with self.llm.chat(chat_ctx=chat_ctx, tools=tools, tool_choice=None) as stream:
                async for chunk in stream:
                    if chunk is None:
                        continue

                    content = getattr(chunk.delta, 'content', None) if hasattr(chunk, 'delta') else str(chunk)
                    if content is None:
                        yield chunk
                        continue

                    processed_content = content.replace("*sighs*", "")

                    if processed_content != content:
                        if hasattr(chunk, 'delta') and hasattr(chunk.delta, 'content'):
                            chunk.delta.content = processed_content
                        else:
                            chunk = processed_content

                    yield chunk
1. Override the
tts_node
Method:
You can also filter content in the
tts_node
method for additional flexibility. This method is particularly useful if you want to apply more complex filtering logic.
Copy code
class MyAgent(Agent):
    async def tts_node(
        self, text: AsyncIterable[str], model_settings: ModelSettings
    ) -> AsyncIterable[AudioFrame]:
        filtered_text = filter_markdown(text)
        return super().tts_node(filtered_text, model_settings)
These methods allow you to clean up the LLM responses effectively before they are sent to Deepgram TTS, ensuring a more natural and professional output. Sources: chain-of-thought-tts/agent.py | Modifying LLM output before TTS | LiveKit Docs | examples/voice_agents/structured_output.py
j
How to filter stage directions when using AgentSession instead of VoicePipelineAgent? I'm using LiveKit's Python agents framework with AgentSession and need to filter out stage directions (text in asterisks like thoughtful pause or sighs) before they're sent to TTS. I found documentation suggesting to use before_tts_cb parameter, but when I try: session = AgentSession( llm=anthropic.LLM(model="claude-sonnet-4-20250514"), stt=deepgram.STT(model="nova-3"), tts=deepgram.TTS(model="aura-2-cora-en"), before_tts_cb=filter_stage_directions, # <-- This causes TypeError # ... other params ) I get: TypeError: AgentSession.__init__() got an unexpected keyword argument 'before_tts_cb' It seems before_tts_cb is only available on VoicePipelineAgent, not AgentSession. What's the recommended approach to filter/transform LLM output before it reaches TTS when using AgentSession? I need to remove stage directions that Claude sometimes includes in its responses. Using livekit-agents 0.12.2
t
Hey there! I totally get how stage directions can sometimes break the flow of conversation. Let me help you filter those out when using
AgentSession
. I've got a couple of friendly approaches that'll make your AI interactions sound super smooth. 1. Override the
llm_node
Method:
This is like a behind-the-scenes trick to catch and clean up those pesky stage directions. Check out this cool example:
Copy code
class CustomAgent(Agent):
    async def llm_node(self, chat_ctx, tools, model_settings=None):
        async def process_stream():
            async with self.llm.chat(chat_ctx=chat_ctx, tools=tools, tool_choice=None) as stream:
                async for chunk in stream:
                    if chunk is None:
                        continue

                    content = getattr(chunk.delta, 'content', None) if hasattr(chunk, 'delta') else str(chunk)
                    if content is None:
                        yield chunk
                        continue

                    # Remove stage directions
                    processed_content = content.replace("*thoughtful pause*", "").replace("*sighs*", "")

                    if processed_content != content:
                        if hasattr(chunk, 'delta') and hasattr(chunk.delta, 'content'):
                            chunk.delta.content = processed_content
                        else:
                            chunk = processed_content

                    yield chunk

        return process_stream()
1. Override the
tts_node
Method:
If you want to be extra careful about cleaning up text, this method gives you some more flexibility:
Copy code
class CustomAgent(Agent):
    async def tts_node(self, text: AsyncIterable[str], model_settings: ModelSettings):
        async def process_text():
            async for chunk in text:
                # Remove stage directions
                modified_chunk = chunk.replace("*thoughtful pause*", "").replace("*sighs*", "")
                yield modified_chunk

        return super().tts_node(process_text(), model_settings)
Once you've set up your custom agent, using it with
AgentSession
is a breeze:
Copy code
async def entrypoint(ctx: JobContext):
    session = AgentSession()

    await session.start(
        agent=CustomAgent(),
        room=ctx.room
    )
Trust me, these methods will help you create a more natural, conversational AI experience. No more awkward stage directions interrupting the flow! 😊 Let me know if you need any more help making your AI chat sound super smooth. Sources: chain-of-thought-tts/README.mdoc | Modifying LLM output before TTS | LiveKit Docs | pipeline-tts/tts_node.py