jolly-eve-3916
07/28/2025, 8:22 AMtall-airline-16756
07/28/2025, 8:23 AMtall-airline-16756
07/28/2025, 8:24 AMtranscription_node method in your agent. This allows you to capture word-level or sentence-level timestamps depending on the TTS system used (e.g., Cartesia or ElevenLabs).
class MyAgent(Agent):
def __init__(self):
super().__init__(instructions="You are a helpful assistant.")
async def transcription_node(
self, text: AsyncIterable[str | TimedString], model_settings: ModelSettings
) -> AsyncGenerator[str | TimedString, None]:
async for chunk in text:
if isinstance(chunk, TimedString):
<http://logger.info|logger.info>(f"TimedString: '{chunk}' ({chunk.start_time} - {chunk.end_time})")
yield chunk
async def entrypoint(ctx: JobContext):
session = AgentSession(
stt=deepgram.STT(),
llm=openai.LLM(),
tts=cartesia.TTS(),
vad=silero.VAD.load(),
use_tts_aligned_transcript=True,
)
await session.start(agent=MyAgent(), room=ctx.room)
2. STT Transcripts:
• Access the current transcript and listen to STT events to capture user speech.
# Access current transcript (including interim if available)
current_transcript = recognition_instance.current_transcript
# Listen to STT events
async def _on_stt_event(self, ev: stt.SpeechEvent) -> None:
if ev.type == stt.SpeechEventType.FINAL_TRANSCRIPT:
transcript = ev.alternatives[0].text
language = ev.alternatives[0].language
confidence = ev.alternatives[0].confidence
# Process the final transcript
3. Turn-Based Transcript Capture:
• Listen for agent events such as user_speech_committed (when user speech is committed) and agent_stopped_speaking (when the agent stops speaking) to capture transcripts for each turn.
4. Real-Time Transcription Forwarding:
• Use the STTSegmentsForwarder class to forward real-time transcriptions for both the agent and user.
from livekit.agents import stt, transcription
async def _forward_transcription(
stt_stream: stt.SpeechStream,
stt_forwarder: transcription.STTSegmentsForwarder,
):
async for ev in stt_stream:
stt_forwarder.update(ev)
if ev.type == stt.SpeechEventType.INTERIM_TRANSCRIPT:
print(ev.alternatives[0].text, end="")
elif ev.type == stt.SpeechEventType.FINAL_TRANSCRIPT:
print(" -> ", ev.alternatives[0].text)
These methods allow you to capture both STT and TTS transcripts with timing information for each conversational turn in your voice agent.
Sources: livekit-agents/livekit/agents/voice/agent.py | Recording agent sessions | LiveKit Docs | AI voice agents | LiveKit Docs | Transcriptions | LiveKit Docs