steep-balloon-41261
07/23/2025, 11:47 AMtall-airline-16756
07/23/2025, 11:47 AMtall-airline-16756
07/23/2025, 11:49 AMuse_tts_aligned_transcript=True in your AgentSession configuration:
session = AgentSession(
# ... stt, llm, tts, vad, etc...
use_tts_aligned_transcript=True,
)
This feature provides word-level synchronization with Cartesia and ElevenLabs TTS providers, and sentence-level alignment with others. The system dynamically adjusts timing based on the actual audio duration and text length, ensuring accurate transcript pacing even when the TTS output is longer than the original text.
For advanced or manual control, you can use TTSSegmentsForwarder to push text and audio segments explicitly, allowing you to fine-tune synchronization as needed.
These approaches will help keep your transcript display aligned with the agent's spoken output, even when processing extended speakable math expressions.
Sources: Text and transcriptions | LiveKit Docs | livekit-agents/livekit/agents/tts/tts.pyfast-computer-50491
07/23/2025, 11:52 AMtall-airline-16756
07/23/2025, 11:55 AMAgentSession just like this:
session = AgentSession(
stt=your_stt,
llm=your_llm,
tts=cartesia.TTS(), # or elevenlabs.TTS()
vad=your_vad,
use_tts_aligned_transcript=True,
)
1. Now, we'll override the transcription_node to keep your original LaTeX formula looking sharp.
2. Then, we'll tweak the tts_node to convert LaTeX into a more speakable version for your agent.
Here's the magic: your transcript will show the exact formula, while your agent speaks a more natural-sounding explanation. And the best part? Everything stays perfectly synchronized!
Check out this simplified example that shows you exactly how to do it:
class LaTeXAgent(Agent):
async def transcription_node(self, text, model_settings):
async for chunk in text:
yield chunk # Keep original LaTeX for transcript
async def tts_node(self, text, model_settings):
async def process_text():
async for chunk in text:
yield convert_latex_to_speech(chunk) # Expand LaTeX for TTS
async for frame in Agent.default.tts_node(self, process_text(), model_settings):
yield frame
Pretty neat, right? This approach gives you tons of flexibility to customize how your LaTeX formulas are displayed and spoken. Feel free to play around with the conversion logic to make it work perfectly for your specific needs!
Sources: Text and transcriptions | LiveKit Docs | Pipeline nodes and hooks | LiveKit Docs | livekit-agents/livekit/agents/voice/transcription/synchronizer.py | Cartesia TTS integration guide | LiveKit Docs | ElevenLabs TTS integration guide | LiveKit Docs