Ok, so I am having a voice agent, but. how. do i g...
# ask-ai
j
Ok, so I am having a voice agent, but. how. do i get teh stt transcript and tts transcript for each turn ?
t
Thanks for reaching out, Albert. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To get STT (Speech-to-Text) and TTS (Text-to-Speech) transcripts for each turn in your voice agent, you can follow these approaches: 1. TTS Transcripts with Timing Information: • Enable timed transcripts and override the
transcription_node
method in your agent. This allows you to capture word-level or sentence-level timestamps depending on the TTS system used (e.g., Cartesia or ElevenLabs).
Copy code
class MyAgent(Agent):
    def __init__(self):
        super().__init__(instructions="You are a helpful assistant.")

    async def transcription_node(
        self, text: AsyncIterable[str | TimedString], model_settings: ModelSettings
    ) -> AsyncGenerator[str | TimedString, None]:
        async for chunk in text:
            if isinstance(chunk, TimedString):
                <http://logger.info|logger.info>(f"TimedString: '{chunk}' ({chunk.start_time} - {chunk.end_time})")
            yield chunk

async def entrypoint(ctx: JobContext):
    session = AgentSession(
        stt=deepgram.STT(),
        llm=openai.LLM(),
        tts=cartesia.TTS(),
        vad=silero.VAD.load(),
        use_tts_aligned_transcript=True,
    )
    await session.start(agent=MyAgent(), room=ctx.room)
2. STT Transcripts: • Access the current transcript and listen to STT events to capture user speech.
Copy code
# Access current transcript (including interim if available)
current_transcript = recognition_instance.current_transcript

# Listen to STT events
async def _on_stt_event(self, ev: stt.SpeechEvent) -> None:
    if ev.type == stt.SpeechEventType.FINAL_TRANSCRIPT:
        transcript = ev.alternatives[0].text
        language = ev.alternatives[0].language
        confidence = ev.alternatives[0].confidence
        # Process the final transcript
3. Turn-Based Transcript Capture: • Listen for agent events such as
user_speech_committed
(when user speech is committed) and
agent_stopped_speaking
(when the agent stops speaking) to capture transcripts for each turn. 4. Real-Time Transcription Forwarding: • Use the
STTSegmentsForwarder
class to forward real-time transcriptions for both the agent and user.
Copy code
from livekit.agents import stt, transcription

async def _forward_transcription(
    stt_stream: stt.SpeechStream,
    stt_forwarder: transcription.STTSegmentsForwarder,
):
    async for ev in stt_stream:
        stt_forwarder.update(ev)
        if ev.type == stt.SpeechEventType.INTERIM_TRANSCRIPT:
            print(ev.alternatives[0].text, end="")
        elif ev.type == stt.SpeechEventType.FINAL_TRANSCRIPT:
            print(" -> ", ev.alternatives[0].text)
These methods allow you to capture both STT and TTS transcripts with timing information for each conversational turn in your voice agent. Sources: livekit-agents/livekit/agents/voice/agent.py | Recording agent sessions | LiveKit Docs | AI voice agents | LiveKit Docs | Transcriptions | LiveKit Docs