transcript doesn't have timestamps
# ask-ai
c
transcript doesn't have timestamps
t
Thanks for reaching out, Arjun. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Could you please share a bit more about how you're generating or accessing the transcript? For instance, are you using a specific API, SDK, or a feature within LiveKit? Understanding this will really help me guide you better on how to include timestamps. I'm here to help!
c
history_data = session.history.to_dict()
t
I'm happy to help clarify this for you! The method
session.history.to_dict()
does include timestamps. Each message in your conversation history has a
timestamp
field that shows when the message was created. If you’re not seeing these timestamps, it might be helpful to double-check the structure of the data returned by
to_dict()
to make sure the
timestamp
field is included. Please feel free to reach out if you need any more help with this! Sources: livekit-ffi/protocol/stats.proto | livekit/livekit_models.pb.go | livekit-protocol/livekit/protocol/analytics.py
c
share the code to fetch the transcript with sentence level timestamps
t
To help you fetch transcripts with sentence-level timestamps, here’s a friendly code example you can follow:
Copy code
import asyncio
import logging
from collections.abc import AsyncGenerator, AsyncIterable

from livekit.agents import Agent, AgentSession, JobContext, WorkerOptions, cli
from livekit.agents.voice.agent import ModelSettings
from <http://livekit.agents.voice.io|livekit.agents.voice.io> import TimedString
from livekit.plugins import deepgram, openai, cartesia, silero

logger = logging.getLogger("my-worker")
logger.setLevel(<http://logging.INFO|logging.INFO>)

class MyAgent(Agent):
    def __init__(self):
        super().__init__(instructions="You are a helpful assistant.")
        self._closing_task: asyncio.Task[None] | None = None

    async def transcription_node(
        self, text: AsyncIterable[str | TimedString], model_settings: ModelSettings
    ) -> AsyncGenerator[str | TimedString, None]:
        async for chunk in text:
            if isinstance(chunk, TimedString):
                # Here, you’ll get sentence-level timestamps for TTS providers without word-level timing
                <http://logger.info|logger.info>(f"TimedString: '{chunk}' ({chunk.start_time} - {chunk.end_time})")
            yield chunk

async def entrypoint(ctx: JobContext):
    session = AgentSession(
        stt=deepgram.STT(),
        llm=openai.LLM(),
        tts=cartesia.TTS(),  # Feel free to use any TTS provider for sentence-level timing
        vad=silero.VAD.load(),
        use_tts_aligned_transcript=True,  # This enables TTS-aligned transcripts
    )

    await session.start(agent=MyAgent(), room=ctx.room)
    await session.generate_reply(instructions="say hello to the user")

if __name__ == "__main__":
    cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
Here are some key points to keep in mind: • Make sure to set
use_tts_aligned_transcript=True
in your
AgentSession
to get those aligned transcripts. • The
transcription_node
method is where you’ll receive
TimedString
objects that include
start_time
and
end_time
. • If your TTS provider doesn’t support word-level timing, don’t worry—the alignment will be done at the sentence level instead. If you have any questions or need a hand with anything else, just let me know. I’m here to help! Sources: Text and transcriptions | LiveKit Docs | Text and transcriptions | LiveKit Docs | Docs::Agents
c
we are using realtime models
t
Hi there! If you're looking to fetch sentence-level timestamps for real-time models, you can use the TTS-aligned transcription feature in LiveKit. Let me guide you through it step-by-step: 1. Enable TTS-Aligned Transcription: Just set
use_tts_aligned_transcript=True
in your
AgentSession
configuration. This will make sure the transcription is forwarded in sync.
Copy code
session = AgentSession(
    stt=deepgram.STT(),
    llm=openai.LLM(),
    tts=cartesia.TTS(),  # or another TTS provider
    vad=silero.VAD.load(),
    use_tts_aligned_transcript=True,
)
1. Implement the
transcription_node
Method:
This method receives a stream of
TimedString
objects, which include
start_time
and
end_time
fields.
Copy code
async def transcription_node(
    self, 
    text: AsyncIterable[str | TimedString], 
    model_settings: ModelSettings
) -> AsyncGenerator[str | TimedString, None]:
    async for chunk in text:
        if isinstance(chunk, TimedString):
            # This contains start_time and end_time in seconds
            <http://logger.info|logger.info>(f"TimedString: '{chunk}' ({chunk.start_time} - {chunk.end_time})")
        yield chunk
A few important things to keep in mind: • Only some TTS providers like Cartesia and ElevenLabs support word-level timestamps; for others, alignment happens at the sentence level. • The system processes and streams text by sentence to keep the audio output smooth. • This feature is still experimental and currently available only in Python. If you'd like to see a full working example, check out timed_agent_transcript.py. And please don’t hesitate to reach out if you have any questions or need a hand—I’m here to help! Sources: Live Conversations with AI using ChatGPT and WebRTC | livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime... | Text and transcriptions | LiveKit Docs | Docs::Agents
c
this is my config: session = AgentSession( llm=openai.realtime.RealtimeModel( model="gpt-4o-realtime-preview", # TODO: use the latest model and checkout the new voice as well voice="coral", temperature=0.8, turn_detection=TurnDetection( type="semantic_vad", eagerness="auto", create_response=True, interrupt_response=True, ) ) )
t
Thanks so much for reaching out! I wanted to let you know that, at the moment, sentence-level timestamps aren’t supported for the
openai.realtime.RealtimeModel
in LiveKit. The OpenAI Realtime API integration doesn’t provide timing metadata for audio output or transcription just yet. However, LiveKit does support TTS-aligned transcriptions with precise timing when you use external TTS providers like Cartesia or ElevenLabs—this feature isn’t available with OpenAI’s native speech output, though. If you’re looking to get sentence-level timing information for speech output, a great approach is to use a separate TTS provider alongside LiveKit’s TTS-aligned transcript feature, while still using OpenAI’s Realtime API for text generation. I’d be more than happy to help you set this up or answer any questions you have—just let me know! Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime... | OpenAI Realtime API integration guide | LiveKit Docs | plugins/openai/src/realtime/realtime_model.ts
c
can we strict the language to english in realtime models config
t
Could you please let me know if you'd like to restrict the input language, the output language, or both in the openai.realtime.RealtimeModel configuration? Understanding this will really help me give you the best possible guidance.
c
both