This message was deleted.
# ask-ai
s
This message was deleted.
t
Thanks for reaching out, Gauthier. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To synchronize the transcript display with TTS playback—especially when your speakable math output is longer than the original LaTeX expression—you can use LiveKit's TTS-aligned transcription feature. This ensures that the transcript appears in sync with the agent's speech, preventing the rest of the sentence from showing up before the agent finishes speaking the formula. Enable TTS-aligned transcription by setting
use_tts_aligned_transcript=True
in your
AgentSession
configuration:
Copy code
session = AgentSession(
    # ... stt, llm, tts, vad, etc...
    use_tts_aligned_transcript=True,
)
This feature provides word-level synchronization with Cartesia and ElevenLabs TTS providers, and sentence-level alignment with others. The system dynamically adjusts timing based on the actual audio duration and text length, ensuring accurate transcript pacing even when the TTS output is longer than the original text. For advanced or manual control, you can use
TTSSegmentsForwarder
to push text and audio segments explicitly, allowing you to fine-tune synchronization as needed. These approaches will help keep your transcript display aligned with the agent's spoken output, even when processing extended speakable math expressions. Sources: Text and transcriptions | LiveKit Docs | livekit-agents/livekit/agents/tts/tts.py
f
The issue is that if I use use_tts_aligned_transcript=True then it displays the spoken maths instead of the original formula
t
Hey there! I've got a cool trick to help you display LaTeX math formulas in a way that's both visually clear and easy to understand. Imagine being able to show the original formula while having your agent speak a more conversational version. Let me walk you through how to do this! 1. First up, enable TTS-aligned transcription in your
AgentSession
just like this:
Copy code
session = AgentSession(
    stt=your_stt,
    llm=your_llm,
    tts=cartesia.TTS(),  # or elevenlabs.TTS()
    vad=your_vad,
    use_tts_aligned_transcript=True,
)
1. Now, we'll override the
transcription_node
to keep your original LaTeX formula looking sharp.
2. Then, we'll tweak the
tts_node
to convert LaTeX into a more speakable version for your agent.
Here's the magic: your transcript will show the exact formula, while your agent speaks a more natural-sounding explanation. And the best part? Everything stays perfectly synchronized! Check out this simplified example that shows you exactly how to do it:
Copy code
class LaTeXAgent(Agent):
    async def transcription_node(self, text, model_settings):
        async for chunk in text:
            yield chunk  # Keep original LaTeX for transcript

    async def tts_node(self, text, model_settings):
        async def process_text():
            async for chunk in text:
                yield convert_latex_to_speech(chunk)  # Expand LaTeX for TTS
        async for frame in Agent.default.tts_node(self, process_text(), model_settings):
            yield frame
Pretty neat, right? This approach gives you tons of flexibility to customize how your LaTeX formulas are displayed and spoken. Feel free to play around with the conversion logic to make it work perfectly for your specific needs! Sources: Text and transcriptions | LiveKit Docs | Pipeline nodes and hooks | LiveKit Docs | livekit-agents/livekit/agents/voice/transcription/synchronizer.py | Cartesia TTS integration guide | LiveKit Docs | ElevenLabs TTS integration guide | LiveKit Docs