jolly-eve-3916
07/30/2025, 3:17 AMtall-airline-16756
07/30/2025, 3:17 AMtall-airline-16756
07/30/2025, 3:18 AMbefore_tts_cb Callback: This is the simplest and most effective method. You can define a callback function to process and clean the text before it is passed to the TTS engine. For example:
async def _before_tts_cb(agent: VoicePipelineAgent, text: str | AsyncIterable[str]):
if isinstance(text, str):
# Handle non-streaming text
result = text.replace("*thoughtful pause*", "")
return result
else:
# Handle streaming text
async def process_stream():
async for chunk in text:
processed = chunk.replace("*thoughtful pause*", "")
yield processed
return process_stream()
agent = VoicePipelineAgent(
vad=ctx.proc.userdata["vad"],
stt=openai.STT(),
llm=openai.LLM.with_groq(model="deepseek-r1-distill-llama-70b"),
tts=deepgram.TTS(model="aura-2"),
before_tts_cb=_before_tts_cb,
chat_ctx=initial_ctx
)
1. Override the llm_node Method: For more control, you can create a custom agent class and filter the content as it streams from the LLM. This method is useful for real-time filtering of streaming text.
class ChainOfThoughtAgent(Agent):
async def llm_node(self, chat_ctx, tools, model_settings=None):
async def process_stream():
async with self.llm.chat(chat_ctx=chat_ctx, tools=tools, tool_choice=None) as stream:
async for chunk in stream:
if chunk is None:
continue
content = getattr(chunk.delta, 'content', None) if hasattr(chunk, 'delta') else str(chunk)
if content is None:
yield chunk
continue
processed_content = content.replace("*sighs*", "")
if processed_content != content:
if hasattr(chunk, 'delta') and hasattr(chunk.delta, 'content'):
chunk.delta.content = processed_content
else:
chunk = processed_content
yield chunk
1. Override the tts_node Method: You can also filter content in the tts_node method for additional flexibility. This method is particularly useful if you want to apply more complex filtering logic.
class MyAgent(Agent):
async def tts_node(
self, text: AsyncIterable[str], model_settings: ModelSettings
) -> AsyncIterable[AudioFrame]:
filtered_text = filter_markdown(text)
return super().tts_node(filtered_text, model_settings)
These methods allow you to clean up the LLM responses effectively before they are sent to Deepgram TTS, ensuring a more natural and professional output.
Sources: chain-of-thought-tts/agent.py | Modifying LLM output before TTS | LiveKit Docs | examples/voice_agents/structured_output.pyjolly-eve-3916
07/30/2025, 3:40 AMtall-airline-16756
07/30/2025, 3:41 AMAgentSession. I've got a couple of friendly approaches that'll make your AI interactions sound super smooth.
1. Override the llm_node Method: This is like a behind-the-scenes trick to catch and clean up those pesky stage directions. Check out this cool example:
class CustomAgent(Agent):
async def llm_node(self, chat_ctx, tools, model_settings=None):
async def process_stream():
async with self.llm.chat(chat_ctx=chat_ctx, tools=tools, tool_choice=None) as stream:
async for chunk in stream:
if chunk is None:
continue
content = getattr(chunk.delta, 'content', None) if hasattr(chunk, 'delta') else str(chunk)
if content is None:
yield chunk
continue
# Remove stage directions
processed_content = content.replace("*thoughtful pause*", "").replace("*sighs*", "")
if processed_content != content:
if hasattr(chunk, 'delta') and hasattr(chunk.delta, 'content'):
chunk.delta.content = processed_content
else:
chunk = processed_content
yield chunk
return process_stream()
1. Override the tts_node Method: If you want to be extra careful about cleaning up text, this method gives you some more flexibility:
class CustomAgent(Agent):
async def tts_node(self, text: AsyncIterable[str], model_settings: ModelSettings):
async def process_text():
async for chunk in text:
# Remove stage directions
modified_chunk = chunk.replace("*thoughtful pause*", "").replace("*sighs*", "")
yield modified_chunk
return super().tts_node(process_text(), model_settings)
Once you've set up your custom agent, using it with AgentSession is a breeze:
async def entrypoint(ctx: JobContext):
session = AgentSession()
await session.start(
agent=CustomAgent(),
room=ctx.room
)
Trust me, these methods will help you create a more natural, conversational AI experience. No more awkward stage directions interrupting the flow! 😊 Let me know if you need any more help making your AI chat sound super smooth.
Sources: chain-of-thought-tts/README.mdoc | Modifying LLM output before TTS | LiveKit Docs | pipeline-tts/tts_node.py