in livekit agent is there any setting that we can ...
# ask-ai
c
in livekit agent is there any setting that we can use to make the agent record the seesion and so that we can see it on the livekit cloud,, like transcript / recording and everything.. and if not on livekit cloud then is there any option or is it exposed in some way so that we can send it or save it somewhere else atleadt
t
Thanks for reaching out, Divyanshu. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I'm glad you're interested in learning about LiveKit's recording and transcript options! Here's a friendly overview to help you out: • Audio/Video Recording: You can use the Egress feature to easily record audio and/or video during your sessions. These recordings can be safely stored in places like Amazon S3, Google Cloud Storage, Azure Blob Storage, or any S3-compatible storage you prefer. • Text Transcripts: For real-time transcripts, LiveKit offers the
llm_node
or
transcription_node
. You can access the full conversation history anytime through the
session.history
property. Plus, you can save these transcripts to files using
add_shutdown_callback
, or get instant updates via events like
conversation_item_added
and
user_input_transcribed
. Just so you know, recording starts as soon as an agent joins the room and automatically captures all audio and video shared there. It will stop once everyone has left the session. If you have any questions or need further help with this, I’m here for you! Sources: Session recording and transcripts | LiveKit Docs | Recording agent sessions | LiveKit Docs
c
so, on livekit cloud itself there is no way or monitoring and dashboard for the recording and the calls.. Like vapi provides
t
We're excited to share that LiveKit Cloud offers a user-friendly dashboard designed to help you keep a close eye on your calls and recordings. Here’s a quick overview of what you can explore: • Session and Participant Analytics: Easily check connection minutes, quality, latency, and network connection times to understand how your sessions are performing. • Track-Level Statistics: Dive into details like bitrate, frame rate, packet loss, and simulcast layer changes for each individual stream. • Call Logs and Telephony Observability: If you’re working on SIP projects, you’ll find detailed call logs and telephony analytics right at your fingertips. • Real-Time Metrics and Error Tracking: Stay informed with live updates on session counts, agent status, errors, and system load. • Usage, Billing, and Limits: Keep track of your usage against plan limits and billing information effortlessly. • Agent Observability: Get insights into agent-specific metrics like uptime, session start latency, and resource consumption to better support your team. Whenever you’re ready, you can access the dashboard anytime at cloud.livekit.io. We’re here to support you every step of the way! Sources: Deploy and scale agents on LiveKit Cloud | Operational Guidance | livekit/livekit_composite | DeepWiki | Announcing LiveKit Cloud | Say hello to improved Cloud analytics | Deploying to production | LiveKit Docs
c
• Call Logs and Telephony Observability: If you’re working on SIP projects, you’ll find detailed call logs and telephony analytics right at your fingertips. will this contain the call recording and transcript too?
t
Thanks so much for reaching out! I wanted to share that LiveKit Cloud's call logs and telephony analytics for SIP projects don’t automatically include call recordings or transcripts. But no worries—these features can be enabled and set up with a little guidance: • Call Recordings: You can use the Egress feature to programmatically record audio or video from your sessions. These recordings can then be saved to cloud storage options like S3, GCP, or Azure. • Transcripts: For real-time transcription, the Agents framework has you covered. Transcriptions come through as events, which you can easily forward or save. If you’d like, I’m here to help you get started with setting up recording or transcription for your SIP calls. Just let me know, and we’ll walk through it together! Sources: Say hello to improved Cloud analytics | packages/javascript/CHANGELOG.md | Session recording and transcripts | LiveKit Docs | Transcriptions | LiveKit Docs | agents/CHANGELOG.md
c
yeah so these are my current implementation of agent.. update it so that it saves the recording at a s3 whose creds i will give using a boto client... and as for the transcript it should be saved locaaly with all relevant metadata for each call or session my agent does.. HERE IS MY CURRENT CODE::
Copy code
import asyncio
import logging

from dotenv import load_dotenv
from livekit import rtc
from livekit.agents import JobContext, WorkerOptions, cli, get_job_context
from livekit.agents.llm import ChatContext, ChatMessage, ImageContent
from livekit.agents.voice import Agent, AgentSession
from livekit.plugins import google, openai, silero, deepgram, elevenlabs

logger = logging.getLogger("vision-agent")
logger.setLevel(<http://logging.INFO|logging.INFO>)

load_dotenv(".env.local")


class VisionAgent(Agent):
    def __init__(self) -> None:
        self._latest_frame = None
        self._video_stream = None
        self._tasks = []
        super().__init__(
            instructions="""
                You are an assistant communicating through voice with vision capabilities.
                You can see what the user is showing you through their camera.
                Don't use any unpronouncable characters.
            """,

            #  Try with realtime models of google. But not able see the screen with this.
            # llm=google.beta.realtime.RealtimeModel(model = 'gemini-2.0-flash-live-001'),

            
            # Using gemini flash llm and then STT and TTS models from other Elevnlabs and Deepgram. Screenshare working fine here.
            stt=deepgram.STT(model="nova-2"),
            llm=google.LLM(model="gemini-2.5-flash", temperature=0.25),
            tts=elevenlabs.TTS(),


            vad=silero.VAD.load(),
        )

    async def on_enter(self):
        room = get_job_context().room

        # Find the first video track (if any) from the remote participant
        if room.remote_participants:
            remote_participant = list(room.remote_participants.values())[0]
            video_tracks = [
                publication.track
                for publication in list(remote_participant.track_publications.values())
                if publication.track
                and publication.track.kind == rtc.TrackKind.KIND_VIDEO
            ]
            if video_tracks:
                self._create_video_stream(video_tracks[0])

        # Watch for new video tracks not yet published
        @room.on("track_subscribed")
        def on_track_subscribed(
            track: rtc.Track,
            publication: rtc.RemoteTrackPublication,
            participant: rtc.RemoteParticipant,
        ):
            if track.kind == rtc.TrackKind.KIND_VIDEO:
                self._create_video_stream(track)

    async def on_user_turn_completed(
        self, turn_ctx: ChatContext, new_message: ChatMessage
    ) -> None:
        # Add the latest video frame, if any, to the new message
        if self._latest_frame:
            new_message.content.append(ImageContent(image=self._latest_frame))
            self._latest_frame = None

    # Helper method to buffer the latest video frame from the user's track
    def _create_video_stream(self, track: rtc.Track):
        # Close any existing stream (we only want one at a time)
        if self._video_stream is not None:
            self._video_stream.close()

        # Create a new stream to receive frames
        self._video_stream = rtc.VideoStream(track)

        async def read_stream():
            async for event in self._video_stream:
                # Store the latest frame for use later
                self._latest_frame = event.frame

        # Store the async task
        task = asyncio.create_task(read_stream())
        task.add_done_callback(lambda t: self._tasks.remove(t))
        self._tasks.append(task)


async def entrypoint(ctx: JobContext):
    session = AgentSession()

    await session.start(agent=VisionAgent(), room=ctx.room)


if __name__ == "__main__":
    cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!