job_context.room.participants
# ask-ai
c
job_context.room.participants
t
Thanks for reaching out, faiq. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Could you clarify your question or provide more details about what you need help with regarding
job_context.room.participants
? For example, are you asking about its usage, troubleshooting, or something else?
c
Transfer failed at API: Failed to initiate transfer: TwirpError(code=permission_denied, message=twirp error unknown: sip status 403: FORBIDDEN, status=403, metadata={'error_details': 'CAcSGXNpcCBzdGF0dXMgNDAzOiBGT1JCSURERU4aLAoldHlwZS5nb29nbGVhcGlzLmNvbS9saXZla2l0LlNJUFN0YXR1cxIDCJMD', 'sip_status': 'SIP_STATUS_FORBIDDEN', 'sip_status_code': '403'})
t
Hey there! I see you're running into a tricky authentication error with your API request. No worries, I'll help you work through this step by step, just like a friend walking you through a technical hiccup. 1. Verify Token Claims: Let's double-check your access token permissions. For SIP operations, you'll want to make sure your token includes the right access. Here's what that might look like:
Copy code
{
  "sip": {
    "admin": true,
    "call": true
  }
}
2. Check Authentication Headers: Make sure your authorization header looks exactly like this:
Authorization: Bearer <your-token>
. It's an easy thing to overlook! 3. Validate API Key Configuration: Take a quick peek to confirm your API key is active and that the secret matches up correctly. 4. Review Error Metadata: The extra details in your error (like
'sip_status': 'SIP_STATUS_FORBIDDEN'
) can give us a clue about what's going on behind the scenes. 5. For SIP-Specific Issues: Just do a quick scan to ensure your token has the SIP permissions you need for this specific operation. I know technical errors can be frustrating, but we'll get this sorted out together! If you're still stuck after trying these steps, just send over a bit more info about your request, and I'll be happy to help you troubleshoot further. Sources: livekit-protocol/src/livekit.rs | pkg/service/auth.go | livekit/livekit_sip.pb.go | livekit-api/livekit/api/twirp_client.py
c
TransferSIPParticipantRequest does this have dial tone parameter ?
t
Hey there! Just wanted to let you know that the
TransferSIPParticipantRequest
in the LiveKit API does have a handy parameter for playing a dial tone. It’s called
play_dialtone
, and it’s a simple boolean that, when enabled, plays a dial tone to the SIP participant as a little audible heads-up that they’re being transferred. You’ll find this parameter available across all LiveKit SDKs, including: • Protocol definition:
bool play_dialtone = 4;
• Go SDK:
PlayDialtone bool
• Python SDK:
play_dialtone: bool
• Node.js SDK:
playDialtone?: boolean
• Ruby SDK:
play_dialtone: nil
• Kotlin SDK:
playDialtone
• Rust SDK: Comments indicate the same parameter exists. So, whenever you’re kicking off a transfer, you can use this parameter to gently let the user know with a dial tone. Hope that helps! Sources: packages/livekit-server-sdk/src/SipClient.ts | lib/livekit/sip_service_client.rb | src/main/kotlin/io/livekit/server/SipServiceClient.kt | livekit-protocol/src/livekit.rs | Transferring calls | LiveKit Docs
c
participant_identity = await livekit_manager.create_sip_participant( _phone_number_=transfer_to_phone_number, _room_name_=room_name, _user_id_=user_id_for_trunk, _is_outbound_=True # Transfer is always outbound ) is it correct ? can I add dial tone in it ?
t
Hey! The code snippet you shared for creating a SIP participant looks pretty good. Just a quick tip: if you want to add a dial tone, you can include the
play_dialtone
parameter in your request. It’s a simple boolean that lets you play a dial tone in the room, so everyone knows there’s an incoming call until the phone is picked up. Here’s a little tweak to your code to add that dial tone:
Copy code
participant_identity = await livekit_manager.create_sip_participant(
    phone_number=transfer_to_phone_number,
    room_name=room_name,
    user_id=user_id_for_trunk,
    is_outbound=True,  # Transfer is always outbound
    play_dialtone=True  # Optional: Play dial tone
)
With this, a dial tone will play during the transfer, making things clearer for everyone involved. Sources: Make outbound calls | LiveKit Docs | SIP APIs | LiveKit Docs | src/main/kotlin/io/livekit/server/SipServiceClient.kt | packages/livekit-server-sdk/src/SipClient.ts | lib/livekit/sip_service_client.rb
c
can you give me code for transfer service for outbound for agents 1.0
t
Here’s a handy example code snippet to help you set up a transfer service for outbound calls in LiveKit 1.0:
Copy code
participant_identity = await livekit_manager.create_sip_participant(
    phone_number=transfer_to_phone_number,
    room_name=room_name,
    user_id=user_id_for_trunk,
    is_outbound=True,  # Transfer is always outbound
    play_dialtone=True  # Optional: Play dial tone
)
This snippet shows how to create a SIP participant for an outbound transfer, and you can use the
play_dialtone
option if you want to play a dial tone during the transfer. If you have any questions or want me to walk you through anything else, just let me know—I’m here to help! Sources: README.md | Making calls using SIP | LiveKit Docs
c
cold transfer function
t
Hey there! If you want to set up a cold transfer function in LiveKit, you can use the
TransferSIPParticipant
API. A cold transfer basically moves a caller to another number or SIP endpoint without any handoff, ending the original session. Here’s a simple way to get it going: 1. Enable Call Transfers: First, make sure your SIP trunk provider (like Twilio) is set up to allow call transfers. For Twilio, you can turn on transfer mode with this command:
Copy code
twilio api trunking v1 trunks update --sid <twilio-trunk-sid> \
--transfer-mode enable-all \
--transfer-caller-id from-transferee
2. Use the
TransferSIPParticipant
API:
You’ll need to provide these details: ◦ `transfer_to`: The phone number (like
tel:+15105550100
) or SIP URI (like
sip:+15105550100@host
). ◦ `participant_identity`: The identity of the SIP participant you want to transfer. ◦ `room_name`: The room where the call is happening. ◦ `play_dialtone`: Optionally, you can play a dial tone during the transfer (default is
False
). 1. Example Code:
Copy code
import { SipClient } from 'livekit-server-sdk';

async function transferParticipant(participant) {
  const sipTransferOptions = { playDialtone: false };
  const sipClient = new SipClient(process.env.LIVEKIT_URL, process.env.LIVEKIT_API_KEY, process.env.LIVEKIT_API_SECRET);
  const transferTo = "tel:+15105550100";
  await sipClient.transferSipParticipant('open-room', participant.identity, transferTo, sipTransferOptions);
}
2. Handle Transfers in Your Application: Make sure to use the API client within your app’s logic, and don’t forget to set your environment variables with LiveKit credentials and server URL. 3. Exception Handling: It’s a good idea to wrap your transfer calls in try-catch blocks to handle any hiccups and give your users helpful feedback. If you want me to walk you through any part of this or need more code examples, just let me know—I’m here to help! Sources: Transferring calls | LiveKit Docs | Company directory phone assistant | LiveKit Docs
Hey there! If you want to set or grab a participant's identity in LiveKit, you’ll want to use the
participant_identity
field. Here’s a quick rundown: • Setting Participant Identity: When you create a participant, just pop in the
participant_identity
field like this:
Copy code
participant_identity = await livekit_manager.create_sip_participant(
    phone_number=transfer_to_phone_number,
    room_name=room_name,
    user_id=user_id_for_trunk,
    is_outbound=True,
    participant_identity="unique_identity"
)
• Retrieving Participant Identity: To get a participant’s identity, you can use the SDK or API you’re working with. For example, in Swift:
let identity = yourParticipant.identity.stringValue
Just a heads-up: each participant’s identity needs to be unique within a room. If two participants join with the same identity, only the latest one will stay connected. If you want me to walk you through more examples or help out with a specific language, just let me know—I’m here to help! Sources: livekit-protocol/src/livekit.rs | Sources/LiveKit/Types/Participant+Types.swift | Runtime/Scripts/Proto/E2Ee.cs | lib/src/proto/livekit_models.pb.dart
If you want to find a participant's identity in a LiveKit room, you can use the
getParticipantByIdentity
method. It first checks the local participant and then looks through the remote participants—pretty handy! Here’s a quick example of how you can do it: • JavaScript/TypeScript SDK:
Copy code
const participant = room.getParticipantByIdentity(identity);
if (participant) {
    // Feel free to use the participant object here
}
• Swift SDK:
Copy code
func getParticipantByIdentity(identity: String) -> Participant? {
    if self.localParticipant?.identity == identity {
        return self.localParticipant
    }
    return self.remoteParticipants[identity]
}
Just a heads up—make sure your app has the right permissions to access participant info in the room. If the participant isn’t found, this method will return
undefined
or
nil
, so it’s a good idea to handle that scenario. If you have any questions or want me to walk you through it, just let me know—I’m here to help!
c
from python
t
Hey there! In the LiveKit Python SDK, the main differences between
LocalParticipant
and
RemoteParticipant
come down to their roles and what they can do in a room: • LocalParticipant: ◦ This is you—the current user in the room. ◦ You can publish tracks like audio, video, or data to the room. ◦ You can also kick off actions like sending data, DTMF tones, and RPC calls. ◦ There are handy methods for publishing and managing your tracks. • RemoteParticipant: ◦ This represents other folks or processes connected to the room. ◦ You can subscribe to tracks they’ve published. ◦ You’ll get updates about their track changes and connection status. Both participant types share some common attributes like
sid
,
identity
, and
metadata
, but their capabilities differ based on their roles. If you want me to walk you through any part of this or need some examples, just let me know—I’m here to help!
Hey there! The issue happens because the local participant identity (
ctx.room.local_participant.identity
) isn’t actually a SIP participant. The
TransferSIPParticipant
API needs the identity of a SIP participant (which is a remote participant) to make the transfer work. Here’s a quick rundown of why this matters: • Local vs. Remote Participants: The local participant is basically you, while SIP participants are remote folks in the room. • Transfer Requirements: The API specifically wants the identity of a SIP participant, not the local one. To fix this, you’ll want to grab the identity of the SIP participant you want to transfer. You can do this by going through the remote participants in the room and finding the right SIP participant. If you’d like, I’m here to help you figure out how to get that SIP participant identity—just let me know! Sources: Rooms, participants, and tracks | LiveKit Docs | Transferring calls | LiveKit Docs | Company directory phone assistant | LiveKit Docs
Hey there! To grab the SIP participant identity for a cold transfer in LiveKit, you’ll want to make sure you’re looking at the right remote participant (the SIP one) instead of the local participant. Here’s a quick rundown of what might be going sideways and how to fix it: • Issue: Using
ctx.room.local_participant.identity
actually gives you the identity of the local participant (like "agent-AJ_gLWU2qz4f8Bi"), which isn’t tied to a SIP session. • Solution: Loop through the remote participants in the room to find the SIP participant. For example, in Python:
Copy code
for participant in ctx.room.remote_participants:
    if participant.is_sip:
        sip_identity = participant.identity
        break
• Verify the SIP Participant: Double-check that the participant you’re transferring has an active SIP session. The
participant.identity
should match the SIP participant you want to transfer. Once you’ve got the right SIP participant identity, just plug it into the
TransferSIPParticipant
API like this:
Copy code
await livekit_manager.transfer_sip_participant(
    room_name=room_name,
    participant_identity=sip_identity,
    transfer_to="tel:+15105550100"
)
If anything’s unclear or you want to chat more about this, just let me know—I’m here to help!
It looks like your code is trying to get the participant identity for a transfer, but there might be a few things causing it to trip up. Let’s walk through them together: 1. Remote Participants: Make sure the
ctx.room.remote_participants
attribute actually has participants in it. If there aren’t any remote participants, your code will default to the local participant, which won’t work for a SIP transfer. 2. Participant Type: Double-check that the participants inside
remote_participants
are SIP participants. You can add a quick check like this to pick out the right one:
Copy code
for participant in remote_participants.values():
    if participant.is_sip:
        participant_identity = participant.identity
        <http://logger.info|logger.info>(f"Found SIP participant for transfer: {participant_identity}")
        break
3. Fallback to Local Participant: Just a heads up—the fallback to
ctx.room.local_participant.identity
won’t work for SIP transfers. If you don’t find any remote participants, the transfer can’t go through. 4. Error Logging: Your error logging is on point! You might also want to log the type and attributes of each participant in
remote_participants
to get a clearer picture of why the right participant isn’t showing up. Here’s a little tweak to your code with some extra checks to help out:
Copy code
# Try to find the remote SIP participant
if hasattr(ctx.room, 'remote_participants'):
    remote_participants = ctx.room.remote_participants

    if isinstance(remote_participants, dict) and remote_participants:
        for participant in remote_participants.values():
            if participant.is_sip:  # Make sure it’s a SIP participant
                participant_identity = participant.identity
                <http://logger.info|logger.info>(f"Found SIP participant for transfer: {participant_identity}")
                break

if not participant_identity:
    logger.error("No SIP participant found for transfer. Ensure the room has a valid SIP participant.")
If you want to chat more about this or need a hand with anything else, just let me know—I’m here to help!
c
how does livekit support numbers like in what format?
t
LiveKit supports several phone number formats for SIP operations, with a focus on the E.164 international format: • E.164 Format (Recommended): This is the international standard for phone numbers, includingExample
+15105550100
◦ A leading
+
sign ◦ Country code ◦ National number without spaces or special characters • US Phone Numbers: LiveKit makes it easy by automatically formatting 10-digit US phone numbers: ◦ Input:
1234567890
◦ Output:
+11234567890
• Human-Readable Formats: LiveKit is pretty flexible and normalizes various formats, like: ◦
(123) 456 7890
→
+1234567890
◦
123
→
+123
◦ If your number already has a
+
prefix, it stays just as you entered it. • Validation: If something’s off—like the wrong number of digits—you’ll get an error. For US numbers, make sure it’s exactly 10 digits. • Non-Numeric Identifiers: And if you’re using usernames (like
user
) for SIP URIs, LiveKit supports those too! For the smoothest experience, we recommend using the E.164 format and double-checking that your SIP provider supports it. If you have any questions or want to chat more about this, just let me know—I’m here to help!
c
unexpected status from INVITE response: sip status: 403 (FORBIDDEN
t
Hey there! The SIP INVITE response with status 403 FORBIDDEN usually means there’s an issue with authentication or permissions. Let’s walk through some common reasons and how you can fix them: 1. Incorrect Credentials: Double-check that the username and password you’ve set up for your SIP trunk exactly match what your SIP provider gave you. 2. Regional Compliance: Sometimes, calls need to stay within certain national borders. If your call is going to another country, this error might pop up. To fix this, try using region pinning to keep calls within a specific area, or set the
destination_country
parameter for your outbound trunks. 3. Trunk Configuration: Make sure your SIP trunk is set up correctly and has the right permissions for the calls you want to make. If you want, I’m here to help you through any of these steps—just let me know! Sources: SIP troubleshooting guide | LiveKit Docs | livekit/livekit_sip.twirp.go
c
how livekit charges the customer ?
t
I’m happy to help explain how LiveKit’s pricing works! LiveKit uses a usage-based model, which means you’re charged based on the actual resources your applications use. Here’s a quick overview of the main parts of their pricing: • Connection Minutes: This is based on how long participants stay connected to LiveKit’s servers. The rate starts at $0.0005 per minute and gets lower as usage increases. Just so you know, ingress, egress, and SIP participants don’t get charged for connection minutes. • Bandwidth Usage: You’ll see charges for downstream bandwidth (that’s data sent from LiveKit servers to participants) starting at $0.12 per GB, with discounts if you use more. The good news is upstream bandwidth (data sent from clients to LiveKit servers) is free. • Transcoding Minutes: If your media needs transcoding (like for room composites or certain workflows), there’s a per-minute charge for that. • SIP Connection Minutes: For telephony features, SIP participants are billed based on how long they’re connected. LiveKit also offers self-serve pricing plans—Build, Ship, and Scale—with free resource allotments and extra features depending on the plan you choose. You can find all the detailed pricing info on the LiveKit Pricing page. If anything isn’t clear or you want to dive deeper into any part, just let me know—I’m here to help! Sources: The end of Participant Minute pricing | Towards a future-aligned pricing model | Understanding LiveKit Cloud Pricing | A Tale of Two Protocols: WebRTC vs. HLS for Live Streaming | LiveKit Pricing
Thanks for reaching out! I’d be happy to walk you through how LiveKit’s pricing works. It’s based on a usage model, which means you’re charged for the actual resources your applications use. Here’s a quick overview of the main parts: • Connection Minutes: You’re charged based on how long participants stay connected to LiveKit’s servers. The rate starts at $0.0005 per minute and gets cheaper the more you use. Just so you know, ingress, egress, and SIP participants don’t get billed for connection minutes. • Bandwidth Usage: There’s a charge for downstream bandwidth (that’s data sent from LiveKit servers to participants) starting at $0.12 per GB, with discounts if you use a lot. The good news is, upstream bandwidth (data sent from clients to LiveKit servers) is free. • Transcoding Minutes: If your media needs transcoding (like for room composites or certain workflows), you’ll be charged per minute of transcoding. • SIP Connection Minutes: For telephony features, SIP participants are billed based on how long they’re connected. LiveKit also offers self-serve pricing plans called Build, Ship, and Scale. Each plan comes with free resource allotments and extra features depending on which tier you choose. For all the details, you can check out the LiveKit Pricing page. If anything’s unclear or you’d like me to explain more, please don’t hesitate to ask—I’m here to help!
c
we want to get a real-time (or near-realtime) sense of how many rooms are active per user so that we can send them reminders as to when they will likely be reaching 80% 90% etc on usage and ultimalte turn it off once they have reached the threshold -- is there a way to get this as a webhook or alike?
t
LiveKit offers several helpful ways to keep an eye on active rooms and usage thresholds in real-time or close to it: • Webhooks for Event-Based Monitoring: You can easily set up webhooks to get notifications for important room lifecycle events, like:To get started with webhooks, just configure the
webhook
section in your LiveKit server configuration YAML with your API key and the URLs you want to target. ◦ `room_started`: This happens when a room is created and a participant joins. ◦ `room_finished`: This triggers when a room ends and everyone leaves. ◦
participant_joined
and `participant_left`: These inform your backend whenever participants come or go. • RoomService API for Direct Monitoring: You can use the
listRooms
method to programmatically fetch all active rooms. For example, in Node.js: `const rooms = await svc.listRooms(); console.log('existing rooms', rooms);`This way, you get direct access to room details and participant numbers. • Analytics API for Usage Monitoring: If you’re on LiveKit Cloud (Scale plan or higher), the Analytics API offers detailed session data — like bandwidth usage, participant counts, and session activity times. This is super useful for keeping track of usage limits and sending reminders before you hit them. If you’d like, I’m here to walk you through setting up webhooks, using the RoomService API, or accessing the Analytics API — just let me know! Sources: packages/livekit-server-sdk/README.md | Analytics API | LiveKit Docs | Webhooks | LiveKit Docs | pkg/telemetry/events.go
c
how to use analytics API
t
I'd be happy to help you get started with using the Analytics API in LiveKit to monitor your usage and thresholds. Here’s a simple step-by-step guide: 1. Generate an Access Token: You’ll want to create a token using a server-side SDK or the LiveKit CLI that includes the
roomList
grant. For example:
Copy code
lk token create \
--api-key $LIVEKIT_API_KEY \
--api-secret $LIVEKIT_SECRET_KEY \
--list \
--valid-for 24h
2. Make API Requests: With your token ready, you can query the Analytics API endpoint like this: `https://cloud-api.livekit.io/api/project/${PROJECT_ID}/sessions/`Feel free to filter your results using query parameters such as
limit
,
page
,
start
(for example,
YYYY-MM-DD
), and
end
. 3. Retrieve Session Data: The API will provide useful details including: ◦ Creation time ◦ Last active time ◦ Bandwidth in and out (in bytes) ◦ Participant counts 1. Monitor Thresholds: By fetching this data regularly, you can compare it against your thresholds and set up alerts or actions if anything goes beyond what you expect. Just a quick note: access to the Analytics API is available for LiveKit Cloud customers on the Scale plan or higher. If you have any questions or would like me to walk you through some examples, please don’t hesitate to ask. I’m here to help! Sources: Analytics API | LiveKit Docs | protobufs/livekit_analytics.proto
c
how it give real time logs?
t
LiveKit offers real-time logging features to help you monitor and troubleshoot smoothly. Let me walk you through how you can access and make the most of these logs: • Agent Logs: You can use the LiveKit CLI to stream logs live from the latest running agent instance: `lk agent logs`This gives you a continuous stream of runtime logs, including stdout, stderr, and custom logs. If you're interested in build logs, just run:
lk agent logs --log-type=build
• Client SDK Logs: The JavaScript SDK lets you extend and forward logs to third-party services like Datadog using the
setLogExtension()
function. Other SDKs, such as React Native and Android, also support customizable logging levels and WebRTC diagnostics to fit your needs. • External Integrations: You can forward logs to platforms like Datadog, CloudWatch, Sentry, or New Relic by setting up secrets (like API keys) through the LiveKit CLI. This is great for long-term storage and deeper analytics. • Realtime Room Monitor: LiveKit provides an open-source Realtime Room Monitor example that shows real-time logs of all LiveKit events, lets you see room and participant states, and even perform server-side actions. Just a quick note: LiveKit Cloud doesn’t store runtime logs itself, so if you want persistent monitoring, forwarding your logs to an external system is the way to go. If you'd like, I’m here to help you set up any of these options—just let me know! Sources: Log collection | LiveKit Docs | How to Send Logs from the JavaScript SDK to Providers like Datadog | docs/variables/log.html | livekit-android-sdk/src/main/java/io/livekit/android/LiveKit.kt | agent-console/app/_components/console-container.tsx
c
I want to monitor the call timer .
t
To help you keep an eye on call duration or set up a call timer in LiveKit, here are some friendly tips you can use: • Prometheus Metrics: LiveKit offers built-in metrics that make tracking call and session durations straightforward:You can have a Prometheus server gather these metrics to visualize or get alerts easily. ◦ `dur_call_sec`: This tracks the length of a SIP call from when it connects to when it ends. ◦ `dur_session_sec`: This one monitors the entire SIP session duration, covering setup and termination phases. • Maximum Call Duration: If you want to set a limit on how long a call can last, just use the
max_call_duration
parameter during SIP call setup to keep sessions within the boundaries you prefer. • General Session Monitoring: For sessions that aren’t SIP, LiveKit tracks their duration with the
session_duration_ms
metric, which measures the time from when the session starts until it ends. Make sure Prometheus monitoring is turned on in your LiveKit setup, and you can check out real-time data anytime via the
/metrics
endpoint on your LiveKit server. If you’d like, I’m here to help you get this up and running smoothly! Sources: pkg/sip/participant.go | SIP APIs | LiveKit Docs | pkg/telemetry/prometheus/rooms.go
c
participant vs subscriber in livekit.
t
Thanks for reaching out! Let me help clarify the terms "participant" and "subscriber" in LiveKit for you: • Participant: Think of a participant as anyone who joins a room, whether it’s a user or a process. Participants have some cool capabilities: ◦ They can publish media tracks like audio, video, or data. ◦ They can subscribe to tracks that others have published. ◦ Each participant is uniquely identified by an
identity
and has a connection state within the room. • Subscriber: This term describes the role a participant takes when they are receiving and playing media tracks published by others. So, subscribing means the participant is actively getting the data stream to enjoy or use. Usually, participants automatically subscribe to all available tracks if
autoSubscribe
is turned on. But don’t worry—you can also choose to enable or disable subscriptions locally to save bandwidth, like pausing video when it’s offscreen, without fully unsubscribing. To sum it up, every participant can be a subscriber, but "subscriber" specifically refers to the action of receiving media, not a separate kind of participant. I hope this clears things up! If you have more questions, feel free to ask. Sources: Rooms, participants, and tracks | LiveKit Docs | Realtime media | LiveKit Docs | Subscribing to tracks | LiveKit Docs
c
how many disconnect_reason are there in livekit agents
t
Hi there! I wanted to share some helpful information about the different disconnect reasons LiveKit uses to explain why a participant, including agents, might have disconnected from a room. Here’s the full list to give you a clear picture: 1. UNKNOWN_REASON (0): The reason for disconnection isn’t clear. 2. CLIENT_INITIATED (1): The client chose to disconnect on purpose. 3. DUPLICATE_IDENTITY (2): Someone else joined the room using the same identity. 4. SERVER_SHUTDOWN (3): The server instance is in the process of shutting down. 5. PARTICIPANT_REMOVED (4): The participant was removed via RoomService.RemoveParticipant. 6. ROOM_DELETED (5): The room was deleted using RoomService.DeleteRoom. 7. STATE_MISMATCH (6): The client tried to resume a session, but the server didn’t recognize it. 8. JOIN_FAILURE (7): The client wasn’t able to fully connect. 9. MIGRATION (8): (Cloud-only) The server asked the participant to move the connection elsewhere. 10. SIGNAL_CLOSE (9): The signal websocket closed unexpectedly. 11. ROOM_CLOSED (10): The room closed because all Standard and Ingress participants had left. 12. USER_UNAVAILABLE (11): The SIP callee didn’t respond in time. 13. USER_REJECTED (12): The SIP callee rejected the call (busy). 14. SIP_TRUNK_FAILURE (13): There was a SIP protocol failure or an unexpected response. 15. CONNECTION_TIMEOUT (14): The server timed out the participant session. 16. MEDIA_FAILURE (15): There was a media stream failure or media timeout. These reasons are consistent across all LiveKit SDKs and form part of the core protocol definition. If you’d like me to walk you through any of these in more detail, or if you need examples, just let me know—I’m here to help! Sources: agents/src/voice/room_io/room_io.ts | livekit-ffi/protocol/participant.proto | lib/src/types/internal.dart | livekit/livekit_models.pb.go | livekit-protocol/livekit/protocol/models.pyi
c
needs to build web ai agent from livekit.
t
I'm excited to help you build a web AI agent using LiveKit! Here’s a friendly step-by-step guide to get you started: 1. Set Up the AI Agent (Backend): You’ll want to use LiveKit Agents to create a smart participant that joins rooms and interacts with users in real time. You can choose Node.js or Python depending on what you’re comfortable with:
Copy code
// Node.js Example
import { defineAgent, voice } from '@livekit/agents';
import * as deepgram from '@livekit/agents-plugin-deepgram';
import * as openai from '@livekit/agents-plugin-openai';

export default defineAgent({
  entry: async (ctx) => {
    await ctx.connect();
    const session = new voice.AgentSession({
      stt: new deepgram.STT(),
      llm: new openai.LLM(),
    });
    await session.start({
      agent: new voice.Agent({ instructions: "You are a helpful assistant." }),
      room: ctx.room,
    });
  },
});
2. Create the Web Frontend: To make a smooth and interactive web interface, try using the Next.js Voice Assistant Template: `lk app create --template voice-assistant-frontend cd voice-assistant-frontend pnpm install pnpm dev`This template is great because it supports voice input/output, transcription, and even virtual avatars to make the experience more engaging. 3. Connect Frontend and Agent: Make sure both your agent and frontend share the same LiveKit project credentials. This way, the agent can join the room as a participant, and your users can interact seamlessly through the web frontend. 4. Deploy: When you’re ready, run your agent in production mode and deploy the frontend using platforms like Vercel. You can host the agent on LiveKit Cloud or your own server depending on what works best for you. For more detailed guidance, I warmly recommend checking out the Voice AI Quickstart and Web Frontend Guide. And please, don’t hesitate to reach out if you want help tailoring the agent to your specific needs—I’m here to support you every step of the way! Sources: LiveKit Docs | README.md | README.md | complex-agents/vision/README.md | Web and mobile frontends | LiveKit Docs
c
a backend codebase that , will have api and that api when hit from the frontend will return a sessionID and websocket from frontend can be connected with the backend socket on that session ID. is it possible using livekitweb agent ?
t
Thank you for your question! I’d be happy to help clarify how LiveKit handles session IDs and WebSocket connections. • Session Tracking: Although LiveKit doesn’t provide explicit session IDs, you can still track sessions effectively by using the
participant identity
or
room SID
. Your backend can generate scoped access tokens tied to both a room and participant, which helps manage sessions smoothly. • WebSocket Connections: The frontend connects to the LiveKit server through a WebSocket-based WebRTC connection for real-time media and data. Meanwhile, agents connect to backend services (like STT or LLMs) via separate WebSocket or HTTP connections, but these aren’t directly accessible from the frontend. • Frontend-Agent-Backend Communication: LiveKit offers some great built-in features to facilitate communication: ◦ Remote Procedure Calls (RPC): This lets the frontend call methods on the agent, making data exchange structured and straightforward. ◦ State Synchronization: You can share real-time state using participant attributes and data messages, keeping everything in sync. ◦ Text Streams: Real-time text, such as transcriptions, can be sent from the agent directly to the frontend. If your project needs explicit session IDs or direct WebSocket tunneling between frontend and backend, you might need to build a custom solution outside of LiveKit agents. Please feel free to reach out if you want to explore this further or need any additional guidance—I’m here to help! Sources: livekit-agents/livekit/agents/utils/connection_pool.py | README.md | LiveKit Docs | Web and mobile frontends | LiveKit Docs | Generating tokens | LiveKit Docs
c
can we use fastAPI on top to initiate it.
t
Absolutely, you can use FastAPI with LiveKit to initiate sessions and manage agents. FastAPI works wonderfully as a backend service to: • Host AI Models: You can use FastAPI to create endpoints for TTS, STT, or LLM models. For instance, the Kokoro TTS integration uses a FastAPI server to host a text-to-speech model, which the LiveKit agent accesses through an API endpoint. • Orchestrate Session Start: With FastAPI, you can build endpoints to: ◦ Generate LiveKit access tokens with the right permissions. ◦ Start agent sessions by calling external APIs or dispatching agent workers. ◦ Send session details (like LiveKit token and room name) back to the frontend. • Handle Webhooks/Callbacks: FastAPI can also receive event notifications from LiveKit or other services to help manage session lifecycles smoothly. Here’s a simple example of a FastAPI endpoint to kick off an agent session:
Copy code
from fastapi import FastAPI
from livekit.api import AccessToken, VideoGrants

app = FastAPI()

@app.post("/start-agent")
async def start_agent():
    token = (
        AccessToken(api_key="YOUR_API_KEY", api_secret="YOUR_API_SECRET")
        .with_grants(VideoGrants(room_join=True, room="my-room"))
        .to_jwt()
    )
    # Call external avatar API or dispatch agent worker
    return {"livekit_token": token, "room": "my-room"}
FastAPI can also smoothly integrate with LiveKit’s webhooks and other external services to keep track of session states and AI interactions. If you’d like, I’m here to help with a full working example or any further guidance you might need! Sources: README.md
c
can we use tts from deepgram?
t
Absolutely! LiveKit does support text-to-speech (TTS) using Deepgram through its dedicated Deepgram TTS plugin. I’m here to help you get it set up smoothly. Here’s a simple step-by-step guide: 1. Install the Plugin:
pip install livekit-plugins-deepgram
2. Set Up Authentication: You’ll want to grab your Deepgram API key from the Deepgram Console and add it to your environment variables like this:
DEEPGRAM_API_KEY=your_deepgram_api_key
3. Use in Code: Here’s a friendly example of how to integrate Deepgram TTS in a LiveKit agent:
Copy code
from livekit.agents import tts
from livekit.plugins import deepgram

# Initialize Deepgram TTS
tts_plugin = deepgram.TTS(
    model="aura-2-andromeda-en",
    encoding="linear16",
    sample_rate=24000
)

# Use in an AgentSession
session = voice.AgentSession(
    stt=deepgram.STT(),  # optional, if you want to use STT as well
    llm=openai.LLM(),    # your LLM
    tts=tts_plugin
)
4. Streaming TTS: If you’d like to stream TTS, here’s how you can do it:
Copy code
stream = tts_plugin.stream()
await stream.input_text("Hello, this is a test.")
await stream.flush()  # ends the segment
Just a quick note: the Deepgram TTS plugin is currently available only for Python. If you’d like, I’m more than happy to provide a complete working example or help with any other questions you might have! Sources: Models | LiveKit Docs | examples/src/comprehensive_test.ts | Deepgram TTS plugin guide | LiveKit Docs | livekit-plugins/livekit-plugins-upliftai/livekit/plugins/upliftai/tts.py | Deepgram STT plugin guide | LiveKit Docs
c
is flux model available in STT of livekit?
t
Thanks for reaching out! LiveKit does support a variety of STT (speech-to-text) models, including popular providers like Deepgram. However, I didn’t come across any specific mention of a "Flux" model in the documentation. If you happen to have more details about the Flux model or its provider, I’d be happy to help look into compatibility or suggest some great alternatives for you. Just let me know how you’d like to move forward!
c
how to change the port of livekit
t
To help you change the port configuration for your LiveKit server, you’ll want to update the
config.yaml
file that’s used when the server starts up. Here are the important settings to look at: 1. Main TCP Port (API and WebSocket): You can update the
port
field like this:
port: 7880  # Change this to your desired port
2. WebRTC UDP Port Range: This controls the range for WebRTC media traffic. Make sure to adjust it and keep this range open in your firewall:
Copy code
rtc:
  port_range_start: 50000
  port_range_end: 60000
3. WebRTC ICE/TCP Port: If you want to change the fallback TCP port for ICE connections, update it here:
Copy code
rtc:
  tcp_port: 7881  # Change to desired port
4. Using UDP Mux (Optional): To route all UDP traffic through a single port, you can set:
Copy code
rtc:
  udp_port: 7882
5. Helm Deployments (Kubernetes): If you’re using Helm, override these port settings in the
values.yaml
file like this:
Copy code
livekit:
  port: 7880
  rtc:
    tcp_port: 7881
    port_range_start: 50000
    port_range_end: 60000
Once you’ve made the changes, just restart your LiveKit server and double-check that your firewall rules allow traffic on the new ports. If you run into any issues or have questions along the way, I’m here to help—just let me know! Sources: livekit-server/values.yaml | Deploy to Kubernetes | LiveKit Docs | config-sample.yaml | Ports and firewall | LiveKit Docs
c
how to change the port of agent from 8081 to 9091
t
To change the port of a LiveKit agent from 8081 to 9091, you can simply update the
port
parameter in the
WorkerOptions
configuration when you initialize the agent. Here’s a friendly example to guide you:
Copy code
from livekit.agents import WorkerOptions, cli

def entrypoint():
    # Your agent logic here
    pass

if __name__ == "__main__":
    worker_options = WorkerOptions(
        entrypoint_fnc=entrypoint,
        port=9091  # Set custom port here
    )
    cli.run_app(worker_options)
This change will bind your agent's local HTTP health check server to port 9091 instead of the default 8081. If you’re deploying in Docker or similar environments, just make sure that port 9091 is exposed and mapped correctly. If you have any questions or need a hand with this, feel free to reach out—I’m here to help! Sources: python-agent-example-app/Dockerfile | livekit-agents/livekit/agents/worker.py | config-sample.yaml
c
Oct 14 091106 voiceai-box python[861275]: {"message": "received job request", "level": "INFO", "name": "livekit.agents", "job_id": "AJ_rq3aCCxXiNgQ", "dispatch_id": "", "room_name": "voice_assistant_room_39cd3bfa", "agent_name": "", "resuming": false, "timestamp": "2025-10-14T091106.105314+00:00"} Oct 14 091106 voiceai-box python[861275]: {"message": "initializing process", "level": "INFO", "name": "livekit.agents", "pid": 991648, "timestamp": "2025-10-14T091106.118429+00:00"} Oct 14 091106 voiceai-box python[861275]: {"message": "process initialized", "level": "INFO", "name": "livekit.agents", "pid": 991648, "elapsed_time": 0.14, "timestamp": "2025-10-14T091106.255376+00:00"} Oct 14 091155 voiceai-box python[861275]: {"message": "closing agent session due to participant disconnect (disable via
RoomInputOptions.close_on_disconnect=False
)", "level": "INFO", "name": "livekit.agents", "participant": "voice_assistant_user_ffd71286", "reason": "CLIENT_INITIATED", "pid": 861306, "job_id": "AJ_rq3aCCxXiNgQ", "timestamp": "2025-10-14T091155.824786+00:00"} Oct 14 091157 voiceai-box python[861275]: {"message": "process exiting", "level": "INFO", "name": "livekit.agents", "reason": "", "pid": 861306, "job_id": "AJ_rq3aCCxXiNgQ", "timestamp": "2025-10-14T091157.962698+00:00"} Oct 14 091157 voiceai-box python[861306]: [agent] participant voice_assistant_user_ffd71286 metadata: {'agent_name': 'steve 213', 'business_name': 'business', 'abn': '12345678901', 'industry_type': 'retail', 'initial_greeting': 'hello world', 'system_prompt': 'hello world'} why this is happeninng?
t
It looks like the LiveKit agent session is closing because the participant disconnected with the reason
CLIENT_INITIATED
. This usually happens when someone intentionally leaves the room, maybe by calling
Room.disconnect()
or using the app’s interface to exit. By default, LiveKit agents are set up to close their sessions as soon as the participant disconnects. This is managed by the
RoomInputOptions.close_on_disconnect
flag. When it’s set to
True
(which is the default), the agent session will automatically end when the participant leaves. But if you’d prefer the session to stay open, you can change this by setting
close_on_disconnect=False
in your agent’s configuration. Here are some common reasons for
CLIENT_INITIATED
disconnects: • The user manually leaves the room through the app’s UI. • The client SDK calls
disconnect()
on the room instance. • The agent session is linked to the participant’s lifecycle, and it’s set to close on disconnect. To help with this, you might want to check your application’s logic to make sure participants aren’t disconnecting by accident, or adjust the
close_on_disconnect
setting if you want the agent session to keep running after someone leaves. If you have any questions or need a hand with this, just let me know—I’m here to help! Sources: livekit-agents/livekit/agents/voice/room_io/room_io.py | lib/src/proto/livekit_models.pbenum.dart | livekit-rtc/livekit/rtc/_proto/participant_pb2.pyi | LiveKit Server Components | livekit/livekit_composite | DeepWiki | Connecting to LiveKit | LiveKit Docs
c
but the thing is it just sits on process Initialized ? and doesnt work??
t
It sounds like your LiveKit agent is stuck on 'process initialized' and not moving forward. This often happens when the
ctx.connect()
call is missing. After the job entry is triggered, it's important to explicitly call
ctx.connect()
to set up the connection. Without this step, the agent can’t move past initialization. Just a quick note: if you’re using
AgentSession
, it takes care of connecting automatically when started. But if you’re not using
AgentSession
or need to control exactly when the connection happens, make sure to include that
ctx.connect()
call in your code. If you'd like, I’m here to help with a code example or any other questions you have—just let me know! Sources: Job lifecycle | LiveKit Docs
c
from dotenv import load_dotenv import os from livekit import agents from livekit.agents import AgentSession, Agent, RoomInputOptions from livekit.plugins import deepgram, openai, silero, noise_cancellation # Removed heavy multilingual turn detection to improve startup time load_dotenv(".env.local") azure_endpoint = os.getenv("AZURE_OPENAI_ENDPOINT") azure_deployment = os.getenv("AZURE_OPENAI_DEPLOYMENT") openai_api_key = os.getenv("AZURE_OPENAI_API_KEY") class Assistant(Agent): def __init__(self) -> None: super().__init__( instructions="""You are a helpful voice AI assistant. You eagerly assist users with their questions by providing information from your extensive knowledge. Your responses are concise, to the point, and without any complex formatting or punctuation including emojis, asterisks, or other symbols. You are curious, friendly, and have a sense of humor.""", ) async def entrypoint(ctx: agents.JobContext): # Connect to the LiveKit room before interacting with participants await ctx.connect() # Wait for at least one participant to join participant = await ctx.wait_for_participant() try: import json meta = json.loads(participant.metadata) if participant.metadata else {} except Exception: meta = participant.metadata print("[agent] participant", participant.identity, "metadata:", meta) session = AgentSession( stt=deepgram.STT(model="nova-2"), llm=openai.LLM( api_key=openai_api_key, ).with_azure( model="gpt-4o-mini", azure_endpoint=azure_endpoint, azure_deployment=azure_deployment, api_version="2024-08-01-preview" ), tts=deepgram.TTS(model="aura-asteria-en"), vad=silero.VAD.load(), # turn_detection=MultilingualModel(), # Removed turn_detection to speed up initialization ) await session.start( room=ctx.room, agent=Assistant(), room_input_options=RoomInputOptions( # For telephony applications, use
BVCTelephony
instead for best results noise_cancellation=noise_cancellation.BVC(), ), ) await session.generate_reply( instructions="Greet the user and offer your assistance." ) if name == "__main__": agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint, port=9091)) this is my code?
t
Thank you for sharing your LiveKit agent code! It looks well-structured, but I understand how frustrating it can be when it gets stuck on "process initialized" and doesn’t move forward. Let’s walk through some common causes and solutions together: 1. Missing or Incorrect Environment Variables: It’s important to double-check that all required environment variables are set correctly. This includes: ◦
LIVEKIT_URL
◦
LIVEKIT_API_KEY
◦
LIVEKIT_API_SECRET
◦
AZURE_OPENAI_ENDPOINT
,
AZURE_OPENAI_DEPLOYMENT
, and
AZURE_OPENAI_API_KEY
if you’re using Azure OpenAI integration. 1. Initialization Timeout: The agent needs to respond to the initialization request within a certain timeframe. Make sure the
ctx.connect()
call happens promptly, and there aren’t any blocking operations delaying this step. 2. Error in Prewarm or Model Loading: Sometimes, if
silero.VAD.load()
or other model loading doesn’t succeed, the agent can get stuck. Checking your logs for any errors during this prewarm phase can be really helpful. 3. Blocking Code: Try to avoid any synchronous or long-running tasks that might block the initialization. For example, heavy computations or I/O operations inside the
entrypoint
function could cause delays. 4. Check Logs: Taking a close look at the logs for any unhandled exceptions or errors during initialization often reveals useful clues about what’s going on. If you’d like, I’m here to help you dig deeper or provide a tailored troubleshooting guide for your setup. Just let me know! Sources: README.md | Getting Help &amp; Troubleshooting | livekit/livekit_composite | DeepWiki | livekit-agents/livekit/agents/ipc/proto.py
c
Oct 14 095318 voiceai-box python[995544]: {"message": "_SegmentSynchronizerImpl.resume called after close", "level": "WARNING", "name": "livekit.agents", "pid": 995566, "job_id": "AJ_dqWrkqrwwnoj", "timestamp": "2025-10-14T095318.579698+00:00"}
t
I understand that seeing the warning
_SegmentSynchronizerImpl.resume called after close
can be a bit confusing. This message means that the
resume()
method was called on a synchronizer that has already been closed. Usually, this happens when there’s an attempt to resume transcription or audio synchronization after the session or segment has ended. While this warning isn’t harmful, it does hint at a possible hiccup in the control flow. It often pops up during session cleanup, interruptions, or when a segment rotation closes the current synchronizer and starts a new one. If some external code tries to resume playback after this point, the closed synchronizer will kindly remind you with this warning. To help prevent this, it’s best to make sure that
resume()
is only called when the synchronizer is still active and hasn’t been closed yet. Keeping state changes aligned with session lifecycle events like
close
or
drain
usually does the trick. If you’d like, I’m here to help you work through this or provide more tailored advice for your setup. Just let me know! Sources: livekit-agents/livekit/agents/voice/agent_session.py | agents/src/voice/transcription/synchronizer.ts | Agent Session &amp; Activity Management | livekit/livekit_composite | DeepWiki | Events and error handling | LiveKit Docs
c
is this a problem in livekit agent for me ?
t
Hi there! I wanted to reassure you that the warning
_SegmentSynchronizerImpl.resume called after close
isn’t something critical and doesn’t mean there’s a failure in the LiveKit agent itself. It does, however, hint that there might be a little hiccup in how the agent’s session or segment lifecycle is being handled in your code. Even though this warning won’t cause a crash, it might lead to some unexpected behavior if the agent tries to resume operations on a synchronizer that’s already closed. To keep everything running smoothly, it’s a good idea to double-check your code to make sure
resume()
is only called when the synchronizer is active and hasn’t been closed yet. Taking care with session cleanup and segment transitions can really help avoid these warnings. If you’d like, I’m here to help you debug or fine-tune your agent’s lifecycle management—just let me know!
c
from dotenv import load_dotenv import os from livekit import agents from livekit.agents import AgentSession, Agent, RoomInputOptions from livekit.plugins import deepgram, openai, silero, noise_cancellation # Removed heavy multilingual turn detection to improve startup time load_dotenv(".env.local") azure_endpoint = os.getenv("AZURE_OPENAI_ENDPOINT") azure_deployment = os.getenv("AZURE_OPENAI_DEPLOYMENT") openai_api_key = os.getenv("AZURE_OPENAI_API_KEY") class Assistant(Agent): def __init__(self) -> None: super().__init__( instructions="""You are a helpful voice AI assistant. You eagerly assist users with their questions by providing information from your extensive knowledge. Your responses are concise, to the point, and without any complex formatting or punctuation including emojis, asterisks, or other symbols. You are curious, friendly, and have a sense of humor.""", ) async def entrypoint(ctx: agents.JobContext): # Connect to the LiveKit room before interacting with participants await ctx.connect() # Wait for at least one participant to join participant = await ctx.wait_for_participant() try: import json meta = json.loads(participant.metadata) if participant.metadata else {} except Exception: meta = participant.metadata print("[agent] participant", participant.identity, "metadata:", meta) session = AgentSession( stt=deepgram.STT(model="nova-2"), llm=openai.LLM( api_key=openai_api_key, ).with_azure( model="gpt-4o-mini", azure_endpoint=azure_endpoint, azure_deployment=azure_deployment, api_version="2024-08-01-preview" ), tts=deepgram.TTS(model="aura-asteria-en"), vad=silero.VAD.load(), # turn_detection=MultilingualModel(), # Removed turn_detection to speed up initialization ) await session.start( room=ctx.room, agent=Assistant(), room_input_options=RoomInputOptions( # For telephony applications, use
BVCTelephony
instead for best results noise_cancellation=noise_cancellation.BVC(), ), ) await session.generate_reply( instructions="Greet the user and offer your assistance." ) if name == "__main__": agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint, port=9091)) this is my code
t
Your code for the LiveKit agent looks well-structured and should work as expected. If you’re running into any issues like warnings or the agent not behaving as it should, no worries—let’s walk through a few things together: 1. Environment Variables: Double-check that all the necessary environment variables are correctly set in your
.env.local
file: ◦
AZURE_OPENAI_ENDPOINT
◦
AZURE_OPENAI_DEPLOYMENT
◦
AZURE_OPENAI_API_KEY
◦
LIVEKIT_URL
,
LIVEKIT_API_KEY
, and
LIVEKIT_API_SECRET
1. Model Loading: Make sure that
silero.VAD.load()
and other model initializations (like Deepgram STT and TTS) finish without errors. It’s a good idea to check the logs for any hiccups during this step. 2. Session Lifecycle: Confirm that the
ctx.connect()
call happens promptly and that the agent is properly waiting for participants using
ctx.wait_for_participant()
. If no one joins, the agent will naturally stay idle. 3. Warnings: If you notice warnings such as
_SegmentSynchronizerImpl.resume called after close
, this might point to a logic hiccup in how session or segment transitions are managed. Just make sure
resume()
is only called on synchronizers that are still active. 4. Port Configuration: Your agent is set to run on port 9091. It’s worth checking that this port is open and not being blocked by a firewall or any network restrictions. If things still aren’t working as you expect, please feel free to share the specific issue or error message you’re seeing. I’m here to help you get it all sorted out!
c
do I need to modify my code ?
t
It looks like your code is nicely structured and should work well as long as all dependencies and environment variables are set up correctly. To help you along, here are a few things you might want to double-check to keep everything running smoothly: 1. Environment Variables: Take a moment to verify that all the necessary environment variables (like
LIVEKIT_URL
,
LIVEKIT_API_KEY
,
LIVEKIT_API_SECRET
,
AZURE_OPENAI_ENDPOINT
, and others) are properly configured in your
.env.local
file. 2. Model Initialization: Make sure your models (such as
silero.VAD.load()
, Deepgram STT, and TTS) are loading without any hiccups. It’s a good idea to check the logs for any errors during this phase. 3. Session Lifecycle: Your use of
ctx.connect()
and
ctx.wait_for_participant()
looks spot on. Just ensure that participants are actually joining the room, since otherwise the agent might stay idle. 4. Warnings: If you come across warnings like
_SegmentSynchronizerImpl.resume called after close
, it’s worth reviewing how session and segment transitions are handled to keep things running smoothly. If everything is set up correctly and you’re not seeing any specific errors, your code should be good to go! But if you run into any particular issues or want a bit more help, please don’t hesitate to reach out—I’m here to support you.
c
how to emit metrics in my code live
t
To help you emit custom metrics in your LiveKit agent code, here’s a friendly step-by-step guide you can follow: 1. Define Custom Metric Classes: Start by creating a Pydantic model for your custom metrics and wrap it in an event class:
Copy code
from pydantic import BaseModel, Field
from typing import Literal
import time

class CustomMetrics(BaseModel):
    type: Literal["custom_metrics"] = "custom_metrics"
    label: str
    timestamp: float
    custom_value: float

class CustomMetricsCollectedEvent(BaseModel):
    type: Literal["custom_metrics_collected"] = "custom_metrics_collected"
    metrics: CustomMetrics
    created_at: float = Field(default_factory=time.time)
2. Emit Custom Metrics: You can then use the
session.emit()
method to publish your custom metric event like this:
Copy code
custom_metric = CustomMetrics(
    label="api_response_time",
    timestamp=time.time(),
    custom_value=0.45  # e.g., response time in seconds
)

session.emit("custom_metrics_collected", CustomMetricsCollectedEvent(metrics=custom_metric))
3. Handle Emitted Metrics: Next, register an event listener to process the metrics:
Copy code
@session.on("custom_metrics_collected")
def _on_custom_metrics_collected(ev: CustomMetricsCollectedEvent):
    print(f"Received custom metric: {ev.metrics.label} = {ev.metrics.custom_value}")
4. Export to Monitoring Tools: Finally, you can integrate with third-party systems like Prometheus or OpenTelemetry for more advanced observability. LiveKit’s OpenTelemetry integration allows you to export metrics asynchronously. This method helps everything work smoothly with LiveKit’s logging and telemetry systems. If you have any questions or want more examples, just let me know—I’m here to help! Sources: Logs, metrics, and telemetry | LiveKit Docs | chain-of-thought-tts/agent.py | docs/index.yaml | examples/voice_agents/langfuse_trace.py | Creating and Publishing Custom Metrics
c
from dotenv import load_dotenv import os from livekit import agents from livekit.agents import AgentSession, Agent, RoomInputOptions from livekit.plugins import deepgram, openai, silero, noise_cancellation # Removed heavy multilingual turn detection to improve startup time load_dotenv(".env.local") azure_endpoint = os.getenv("AZURE_OPENAI_ENDPOINT") azure_deployment = os.getenv("AZURE_OPENAI_DEPLOYMENT") openai_api_key = os.getenv("AZURE_OPENAI_API_KEY") class Assistant(Agent): def __init__(self) -> None: super().__init__( instructions="""You are a helpful voice AI assistant. You eagerly assist users with their questions by providing information from your extensive knowledge. Your responses are concise, to the point, and without any complex formatting or punctuation including emojis, asterisks, or other symbols. You are curious, friendly, and have a sense of humor.""", ) async def entrypoint(ctx: agents.JobContext): # Connect to the LiveKit room before interacting with participants await ctx.connect() # Wait for at least one participant to join participant = await ctx.wait_for_participant() try: import json meta = json.loads(participant.metadata) if participant.metadata else {} except Exception: meta = participant.metadata print("[agent] participant", participant.identity, "metadata:", meta) session = AgentSession( stt=deepgram.STT(model="nova-2"), llm=openai.LLM( api_key=openai_api_key, ).with_azure( model="gpt-4o-mini", azure_endpoint=azure_endpoint, azure_deployment=azure_deployment, api_version="2024-08-01-preview" ), tts=deepgram.TTS(model="aura-asteria-en"), vad=silero.VAD.load(), # turn_detection=MultilingualModel(), # Removed turn_detection to speed up initialization ) await session.start( room=ctx.room, agent=Assistant(), room_input_options=RoomInputOptions( # For telephony applications, use
BVCTelephony
instead for best results noise_cancellation=noise_cancellation.BVC(), ), ) await session.generate_reply( instructions="Greet the user and offer your assistance." ) if name == "__main__": agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint, port=9091)) add livekit tts stt and metrics in this code
t
Here’s an updated version of your code with LiveKit TTS, STT, and custom metrics integrated:
Copy code
from dotenv import load_dotenv
import os
import time
from livekit import agents
from livekit.agents import AgentSession, Agent, RoomInputOptions
from livekit.plugins import deepgram, openai, silero, noise_cancellation
from pydantic import BaseModel, Field
from typing import Literal

# Load environment variables
load_dotenv(".env.local")

azure_endpoint = os.getenv("AZURE_OPENAI_ENDPOINT")
azure_deployment = os.getenv("AZURE_OPENAI_DEPLOYMENT")
openai_api_key = os.getenv("AZURE_OPENAI_API_KEY")

# Define custom metrics
class CustomMetrics(BaseModel):
    type: Literal["custom_metrics"] = "custom_metrics"
    label: str
    timestamp: float
    custom_value: float

class CustomMetricsCollectedEvent(BaseModel):
    type: Literal["custom_metrics_collected"] = "custom_metrics_collected"
    metrics: CustomMetrics
    created_at: float = Field(default_factory=time.time)

# Define the Assistant class
class Assistant(Agent):
    def __init__(self) -> None:
        super().__init__(
            instructions="""You are a helpful voice AI assistant.
            You eagerly assist users with their questions by providing information from your extensive knowledge.
            Your responses are concise, to the point, and without any complex formatting or punctuation including emojis, asterisks, or other symbols.
            You are curious, friendly, and have a sense of humor.""",
        )

# Define the entrypoint
async def entrypoint(ctx: agents.JobContext):
    # Connect to the LiveKit room before interacting with participants
    await ctx.connect()

    # Wait for at least one participant to join
    participant = await ctx.wait_for_participant()
    try:
        import json
        meta = json.loads(participant.metadata) if participant.metadata else {}
    except Exception:
        meta = participant.metadata
    print("[agent] participant", participant.identity, "metadata:", meta)

    # Initialize the session
    session = AgentSession(
        stt=deepgram.STT(model="nova-2"),
        llm=openai.LLM(
            api_key=openai_api_key,
        ).with_azure(
            model="gpt-4o-mini",
            azure_endpoint=azure_endpoint,
            azure_deployment=azure_deployment,
            api_version="2024-08-01-preview"
        ),
        tts=deepgram.TTS(model="aura-asteria-en"),
        vad=silero.VAD.load(),
    )

    # Emit custom metrics
    custom_metric = CustomMetrics(
        label="session_start",
        timestamp=time.time(),
        custom_value=1.0  # Example metric value
    )
    session.emit("custom_metrics_collected", CustomMetricsCollectedEvent(metrics=custom_metric))

    # Register a listener for custom metrics
    @session.on("custom_metrics_collected")
    def _on_custom_metrics_collected(ev: CustomMetricsCollectedEvent):
        print(f"Received custom metric: {ev.metrics.label} = {ev.metrics.custom_value}")

    # Start the session
    await session.start(
        room=ctx.room,
        agent=Assistant(),
        room_input_options=RoomInputOptions(
            noise_cancellation=noise_cancellation.BVC(),
        ),
    )

    # Generate a reply
    await session.generate_reply(
        instructions="Greet the user and offer your assistance."
    )

if __name__ == "__main__":
    agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint, port=9091))
This updated code includes: • Integration of LiveKit TTS and STT using Deepgram plugins. • Custom metrics emission using
session.emit()
. • A listener to handle and log custom metrics. If you have any questions or would like to explore additional features, just let me know—I'm here to help!
c
livekit.agents.llm
t
livekit.agents.llm
is a key part of the LiveKit Agents framework that beautifully brings large language models (LLMs) into voice and conversational agents. It offers a seamless interface to work with different LLM providers and supports both real-time and non-real-time interactions. Key Features: • LLM Class: This represents a language model with the main method
chat()
, which helps generate responses based on the conversation context. • LLMStream: Allows streaming access to LLM completions, delivering response chunks step-by-step. • ChatContext: Keeps track of the conversation history between the user and the agent, making interactions feel natural. • Tool System: Lets LLMs call external functions as part of their responses, adding extra capabilities. Integration: Here’s how it typically works together with STT and TTS components: 1. The user’s speech is transcribed by an STT model. 2. This transcription is added to the
ChatContext
. 3. The
LLM.chat()
method then processes this context to create a thoughtful response. 4. The response is turned into audio by a TTS model, ready to be heard. Example Usage:
Copy code
# Basic Agent Setup
agent = Agent(
    instructions="You are a friendly voice assistant",
    llm="openai/gpt-4o-mini",
    tools=[lookup_weather],
)
Supported Providers: LiveKit supports a variety of hosted models like OpenAI GPT-4o, plugins for providers such as OpenAI and Anthropic, and even local models through Ollama. If you’d like, I’m here to help with more details or examples tailored to your specific needs—just let me know! Sources: livekit-agents/livekit/agents/__init__.py | Large language models (LLM) | LiveKit Docs | Ollama LLM plugin guide | LiveKit Docs
c
I have a custom frontend and agent.py backend , the voice from web livekit frontend is getting to the agent but agent doesnt speak , the tts is giving issue, it speaks after whole 10 minutes and gives every reply or questions we have asked together.
t
Thank you for reaching out! I understand how frustrating it can be when your LiveKit agent's TTS delays responses and then speaks all replies at once after a long pause. Let’s walk through some common causes and solutions together to help smooth things out: 1. End-of-Utterance Detection: The system waits to make sure you’ve finished speaking before it generates a response. This is controlled by settings like
max_endpointing_delay
, which is usually set to 6 seconds. You might want to try lowering this value to reduce waiting time. 2. Text Pacing and Stream Batching: Sometimes, the TTS buffers text before speaking to keep audio quality high. You can adjust parameters like
min_remaining_audio
and
max_text_length
to help reduce these batching delays:
Copy code
tts=cartesia.TTS(
    text_pacing=tts.SentenceStreamPacer(
        min_remaining_audio=3.0,  # Reduce from default 5.0
        max_text_length=200,      # Reduce from default 300
    )
)
3. Preemptive Speech Generation: Enabling this lets the system start generating responses before you’ve fully finished speaking, which can really help cut down on perceived delays. 4. Proactive Engagement: To keep the conversation feeling lively, you can use short messages or audio cues while the system processes your input. For example:
await agent.say("Let me look that up for you...")
5. Improve Transcription Synchronization: Turning on
use_tts_aligned_transcript=True
in the
AgentSession
config can help keep everything in sync better. 6. Control Response Length: If responses get too long and cause batching, you can override the TTS node to interrupt lengthy replies gracefully, like this:
Copy code
async def tts_node(self, text: AsyncIterable[str], model_settings: ModelSettings):
    MAX_CHUNKS = 20
    chunk_count = 0
    async def process_text():
        nonlocal chunk_count
        async for chunk in text:
            chunk_count += 1
            if chunk_count > MAX_CHUNKS:
                self.session.interrupt()
                self.session.say("I’m sorry, that will take too long to say.")
                break
            yield chunk
    return Agent.default.tts_node(self, process_text(), model_settings)
I hope these tips help make your agent’s responses quicker and more natural! If you have any questions or want to dive deeper into any of these steps, please don’t hesitate to ask—I’m here to help. Sources: Live Conversations with AI using ChatGPT and WebRTC | Turn detection and interruptions | LiveKit Docs | examples/voice_agents/tts_text_pacing.py | Agent speech and audio | LiveKit Docs | delay-handling/README.mdoc
c
await session.say( instructions="Greet the user and offer your assistance." )
t
The
session.say()
function in LiveKit agents is a wonderful tool that lets the agent speak a predefined message to the user. It works by using the configured Text-to-Speech (TTS) engine to turn text into speech, which is then played back to the user in real time. Key Features: • Text-to-Speech: You can provide text to be synthesized into speech. This text will be added to the transcript and chat context unless you specify
add_to_chat_ctx=False
. • Pre-Synthesized Audio: If you have pre-synthesized audio, you can play it by providing an audio stream. In this case, nothing gets added to the transcript or chat context. • Combined Mode: You can also use both text and audio together, where the audio is played, and the text is used for the transcript and chat context. Optional Parameters: • `allow_interruptions`: This lets you decide if the user can interrupt the agent’s speech (default is
True
). • `add_to_chat_ctx`: Controls whether the spoken text is included in the chat context. Example Usage:
Copy code
# Greet the user
await session.say("Hello, how can I help you today?")
Just a quick note: make sure a TTS plugin is set up for the agent unless you’re providing pre-synthesized audio. If you have any questions or want more examples, feel free to ask—I’m here to help! Sources: agents/src/voice/agent_session.ts | basics/exit_message.py | Agent speech and audio | LiveKit Docs | livekit-agents/livekit/agents/voice/agent_session.py | Node.js Agents Architecture | livekit/livekit_composite | DeepWiki
c
from dotenv import load_dotenv import os import logging from livekit import agents from livekit.agents import AgentSession, Agent, RoomInputOptions, metrics, MetricsCollectedEvent from livekit.plugins import deepgram, openai, silero, noise_cancellation # Removed heavy multilingual turn detection to improve startup time load_dotenv(".env.local") azure_endpoint = os.getenv("AZURE_OPENAI_ENDPOINT") azure_deployment = os.getenv("AZURE_OPENAI_DEPLOYMENT") openai_api_key = os.getenv("AZURE_OPENAI_API_KEY") # Preload models for faster response times stt = deepgram.STT(model="nova-2") llm = openai.LLM( api_key=openai_api_key, ).with_azure( model="gpt-4o-mini", azure_endpoint=azure_endpoint, azure_deployment=azure_deployment, api_version="2024-08-01-preview" ) tts = deepgram.TTS(model="aura-asteria-en") vad = silero.VAD.load() noise_cancellation_filter = noise_cancellation.BVC() # Set up logging and a global metrics collector logging.basicConfig(level=logging.INFO) logger = logging.getLogger(name) usage_collector = metrics.UsageCollector() class Assistant(Agent): def __init__(self) -> None: super().__init__( instructions="""You are a helpful voice AI assistant. You eagerly assist users with their questions by providing information from your extensive knowledge. Your responses are concise, to the point, and without any complex formatting or punctuation including emojis, asterisks, or other symbols. You are curious, friendly, and have a sense of humor.""", ) async def entrypoint(ctx: agents.JobContext): # Connect to the LiveKit room before interacting with participants await ctx.connect() # Wait for at least one participant to join participant = await ctx.wait_for_participant() try: import json meta = json.loads(participant.metadata) if participant.metadata else {} except Exception: meta = participant.metadata print("[agent] participant", participant.identity, "metadata:", meta) session = AgentSession( stt=stt, llm=llm, tts=tts, vad=vad, # turn_detection=MultilingualModel(), # Removed turn_detection to speed up initialization ) await session.start( room=ctx.room, agent=Assistant(), room_input_options=RoomInputOptions( # For telephony applications, use
BVCTelephony
instead for best results noise_cancellation=noise_cancellation_filter, ), ) # Collect metrics @session.on("metrics_collected") def on_metrics_collected(event: MetricsCollectedEvent): usage_collector.collect(event.metrics) await session.say( instructions="Hello, how can I help you today?" ) # Log aggregated usage metrics when the session ends summary = usage_collector.get_summary() logger.info(f"Usage: {summary}") # Define a shutdown callback to log usage when the job is finalized async def log_usage(): summary = usage_collector.get_summary() logger.info(f"Usage summary on shutdown: {summary}") ctx.add_shutdown_callback(log_usage) if name == "__main__": agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint, port=9091)) have to do logging of stt , tts and llm , like when they are called , give us log when they are called .
t
To help you keep track of when STT, TTS, and LLM are called in your LiveKit agent, you can simply add some logging statements right where these components are used in your code. I’ve taken the liberty to update your code with logging included, so you can see exactly how it fits together:
Copy code
from dotenv import load_dotenv
import os
import logging
from livekit import agents
from livekit.agents import AgentSession, Agent, RoomInputOptions, metrics, MetricsCollectedEvent
from livekit.plugins import deepgram, openai, silero, noise_cancellation

# Load environment variables
load_dotenv(".env.local")

azure_endpoint = os.getenv("AZURE_OPENAI_ENDPOINT")
azure_deployment = os.getenv("AZURE_OPENAI_DEPLOYMENT")
openai_api_key = os.getenv("AZURE_OPENAI_API_KEY")

# Preload models for faster response times
stt = deepgram.STT(model="nova-2")
llm = openai.LLM(
    api_key=openai_api_key,
).with_azure(
    model="gpt-4o-mini",
    azure_endpoint=azure_endpoint,
    azure_deployment=azure_deployment,
    api_version="2024-08-01-preview"
)
tts = deepgram.TTS(model="aura-asteria-en")
vad = silero.VAD.load()
noise_cancellation_filter = noise_cancellation.BVC()

# Set up logging and a global metrics collector
logging.basicConfig(level=<http://logging.INFO|logging.INFO>)
logger = logging.getLogger(__name__)
usage_collector = metrics.UsageCollector()

class Assistant(Agent):
    def __init__(self) -> None:
        super().__init__(
            instructions="""You are a helpful voice AI assistant.
            You eagerly assist users with their questions by providing information from your extensive knowledge.
            Your responses are concise, to the point, and without any complex formatting or punctuation including emojis, asterisks, or other symbols.
            You are curious, friendly, and have a sense of humor.""",
        )

async def entrypoint(ctx: agents.JobContext):
    # Connect to the LiveKit room before interacting with participants
    await ctx.connect()

    # Wait for at least one participant to join
    participant = await ctx.wait_for_participant()
    try:
        import json
        meta = json.loads(participant.metadata) if participant.metadata else {}
    except Exception:
        meta = participant.metadata
    <http://logger.info|logger.info>(f"[agent] participant {participant.identity} metadata: {meta}")

    # Initialize the session
    session = AgentSession(
        stt=stt,
        llm=llm,
        tts=tts,
        vad=vad,
    )

    # Add logging for STT, TTS, and LLM calls
    <http://logger.info|logger.info>("STT initialized with model nova-2.")
    <http://logger.info|logger.info>("LLM initialized with model gpt-4o-mini.")
    <http://logger.info|logger.info>("TTS initialized with model aura-asteria-en.")

    # Start the session
    await session.start(
        room=ctx.room,
        agent=Assistant(),
        room_input_options=RoomInputOptions(
            noise_cancellation=noise_cancellation_filter,
        ),
    )

    # Collect metrics
    @session.on("metrics_collected")
    def on_metrics_collected(event: MetricsCollectedEvent):
        usage_collector.collect(event.metrics)

    await session.say("Hello, how can I help you today?")

    # Log aggregated usage metrics when the session ends
    summary = usage_collector.get_summary()
    <http://logger.info|logger.info>(f"Usage: {summary}")

    # Define a shutdown callback to log usage when the job is finalized
    async def log_usage():
        summary = usage_collector.get_summary()
        <http://logger.info|logger.info>(f"Usage summary on shutdown: {summary}")

    ctx.add_shutdown_callback(log_usage)

if __name__ == "__main__":
    agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint, port=9091))
Key Additions: • Added logging statements to indicate when STT, TTS, and LLM are initialized. • Logs participant metadata and usage metrics for better observability. If you’d like, I’m here to help you customize this further or add any extra features you might need. Just let me know!
c
max function calls nested depth reached, dropping function context. increase max_nested_fnc_calls to enable additional nesting.
t
It looks like you're encountering the error "max function calls nested depth reached," which happens when your LiveKit agent goes beyond the set limit for sequential function or tool calls during a single conversation turn. This is actually a helpful safety feature designed to prevent infinite loops or using too many resources. Default Configuration: By default, the
max_tool_steps
parameter is set to 3. This means the agent can make up to 3 sequential function calls before hitting this limit. How to Resolve: 1. Increase the Limit: You can easily increase the
max_tool_steps
parameter when creating your
AgentSession
like this:
Copy code
session = AgentSession(
    max_tool_steps=5,  # Increase from default of 3
    # ... other parameters
)
2. Review Function Logic: Take a moment to check that your functions aren’t calling each other in loops. Sometimes combining smaller functions or adding clear exit conditions can really help avoid too much nesting. 3. Implement Error Handling: Using LiveKit’s error handling features can help manage these situations smoothly and provide fallback responses for your users. 4. Monitor and Debug: Don’t forget to look at the logs for the
speech_id
related to the error—that way you can pinpoint which part of the conversation caused it. If you’d like, I’m here to help with more examples or walk you through any of these steps. Just let me know! Sources: livekit-agents/livekit/agents/voice/agent_session.py | Events and error handling | LiveKit Docs
c
agent = VoicePipelineAgent( vad=vad, stt=stt, llm=llm_component, tts=tts, chat_ctx=initial_ctx, turn_detector=turn_detector.EOUModel(), fnc_ctx=function_context # Add function context to enable LLM tools ) how to ?
t
To use
VoicePipelineAgent
with function context and LLM tools in LiveKit, you can define tools in a function context and pass it to the agent. Here’s how you can set it up: Python Implementation:
Copy code
from livekit.agents import llm
from livekit.agents.pipeline import VoicePipelineAgent

# Define function context with tools
function_context = llm.FunctionContext()

@function_context.ai_callable()
def get_weather(location: str) -> str:
    """Get the weather for a specific location."""
    return f"The weather in {location} is sunny."

# Create the VoicePipelineAgent
agent = VoicePipelineAgent(
    chat_ctx=llm.ChatContext().append(
        role="system",
        text="You are a helpful voice assistant."
    ),
    vad=silero.VAD.load(),
    stt=deepgram.STT(),
    llm=openai.LLM(),
    tts=openai.TTS(),
    fnc_ctx=function_context  # Pass function context to enable tools
)
Key Steps: 1. Define Tools: Use the
FunctionContext
to define callable tools, such as
get_weather
. 2. Pass Function Context: Include the
fnc_ctx
parameter when creating the
VoicePipelineAgent
. 3. Use Tools in Conversations: The agent can now call these tools during conversations to provide dynamic responses. If you have any questions or would like me to walk you through more examples for specific tools, please feel free to ask. I’m here to help you every step of the way! Sources: src/agent.ts | examples/voice_agents/README.md | Agents v0.x migration guide - Node.js | LiveKit Docs | Agents v0.x migration guide - Python | LiveKit Docs | Tool definition and use | LiveKit Docs
c
max function calls nested depth reached how to do this above ?
t
It looks like you're encountering the error "max function calls nested depth reached" because the agent has gone beyond the set limit for sequential function or tool calls in a single conversation turn. Don’t worry, I’m here to help you work through this! To get things running smoothly and implement the functionality you described, here are some friendly steps you can follow: 1. Increase the Function Call Limit: One way to tackle this is by increasing the
max_tool_steps
parameter when you create your
VoicePipelineAgent
. This allows more nested function calls. Here’s how you can do it:
Copy code
agent = VoicePipelineAgent(
    chat_ctx=llm.ChatContext().append(
        role="system",
        text="You are a helpful voice assistant."
    ),
    vad=silero.VAD.load(),
    stt=deepgram.STT(),
    llm=openai.LLM(),
    tts=openai.TTS(),
    fnc_ctx=function_context,  # Pass function context to enable tools
    max_tool_steps=5  # Increase from the default of 3
)
2. Optimize Function Logic: It’s a good idea to review your function definitions to make sure they aren’t calling each other in loops or nesting calls more than needed. Sometimes combining smaller functions into one can simplify things. 3. Add Proper Exit Conditions: Double-check that each function has clear stopping points. This helps prevent going too deep into nested calls. 4. Debug and Monitor: Adding logging can be really helpful to see which functions are being called and where the nesting limit is being hit. For example:
Copy code
import logging
logging.basicConfig(level=<http://logging.INFO|logging.INFO>)
logger = logging.getLogger(__name__)

@function_context.ai_callable()
def get_weather(location: str) -> str:
    <http://logger.info|logger.info>(f"get_weather called with location: {location}")
    return f"The weather in {location} is sunny."
By increasing the limit thoughtfully and tidying up your function logic, you’ll be able to avoid that nesting depth error while keeping your calls efficient and clear. If you’d like, I’m here to guide you through any of these steps or answer any questions you have. Just let me know!
c
what is max_nested_calls ? because max_tool_steps doesnt work.
t
Thanks for your question! The parameter
max_tool_steps
in LiveKit agents is what controls the maximum number of consecutive tool calls per LLM turn. By default, it’s set to 3, but you’re absolutely welcome to increase it if you need more flexibility:
Copy code
session = AgentSession(
    max_tool_steps=5,  # Allow up to 5 consecutive tool calls
    # other parameters...
)
Regarding
max_nested_calls
, it seems that this parameter isn’t part of the LiveKit agents framework. If you’re running into an issue or have a specific scenario in mind, please feel free to share a bit more detail. I’m here to help and want to make sure you get the best support possible! Sources: README.md
c
how much is the context of chat_context in livekit ?
t
Thanks for reaching out! I’d love to help you understand the
chat_context
limits in LiveKit agents. Here’s a quick overview: • Maximum of 40 messages: The chat context keeps things manageable by automatically trimming to the most recent 40 messages. • Maximum size of 1 kilobyte (1024 bytes) per message: If a message gets too long, it starts a new one instead of adding on to the previous message. These limits help keep everything running smoothly and compatible with the model’s requirements. If anything’s unclear or you want to dive deeper, just let me know—I’m here to help! Sources: livekit-plugins/livekit-plugins-aws/livekit/plugins/aws/experimental/realti...
c
explain chat_context in detail please.
t
The
chat_context
in LiveKit agents plays a really important role in managing the conversational history during AI agent interactions. Think of it as a well-organized, chronological record of all the messages and function execution events that happen in an agent session. This helps the agent provide responses that truly understand the context and handle more complex workflows smoothly. Purpose: The main goal of
chat_context
is to keep track of the entire interaction history in a LiveKit agent session. This includes messages from both the user and assistant, function calls made by the LLM, and their results. By doing this, agents can easily refer back to earlier parts of the conversation, keep track of the session’s state, and support helpful features like agent handoffs and real-time updates. Structure: The
chat_context
consists of a collection of
ChatItem
objects, which can be: • ChatMessage: A message from a participant, with details like
role
(such as 'user' or 'assistant'),
content
(which could be text, images, or audio), and other metadata. • FunctionCall: A function call initiated by the LLM, including info like
callId
,
name
, and
arguments
. • FunctionCallOutput: The outcome of a function execution, showing
output
and whether there was an error with
isError
. • AgentHandoff: When control is passed from one agent to another, making sure the conversation context continues seamlessly. Usage: Throughout the agent’s lifecycle,
chat_context
helps build and maintain the conversation history. Developers can update it directly using methods like
addMessage
or
update_chat_ctx
. It also supports real-time syncing between client and server, so the state stays consistent and up-to-date. Examples: • Adding a Message:
Copy code
chat_context.addMessage({
  role: 'user',
  content: 'Hello, how can you help me today?'
})
• Tracking Function Calls: Function calls and their results are automatically added to the context, making it easy to reference later. • Context Truncation: For longer conversations, only the most recent 40 messages are kept, with each message limited to 1 kilobyte to keep things efficient. If you’d like, I’m here to help with more details or specific examples tailored to your needs—just let me know! Sources: Agents v0.x migration guide - Python | LiveKit Docs | Agents v0.x migration guide - Node.js | LiveKit Docs | .cursor/rules/agent-core.mdc | Agent Session &amp; Client Integration | livekit/livekit_composite | DeepWiki | livekit/agent/livekit_agent_session.pb.go
c
from dotenv import load_dotenv import os import logging import time from livekit import agents from livekit.agents import AgentSession, Agent, RoomInputOptions, metrics, MetricsCollectedEvent from livekit.plugins import deepgram, openai, silero, noise_cancellation # Removed heavy multilingual turn detection to improve startup time # Import RAG service from rag_service import rag_service load_dotenv(".env.local") azure_endpoint = os.getenv("AZURE_OPENAI_ENDPOINT") azure_deployment = os.getenv("AZURE_OPENAI_DEPLOYMENT") openai_api_key = os.getenv("AZURE_OPENAI_API_KEY") # Set up logging and a global metrics collector logging.basicConfig(level=logging.INFO) logger = logging.getLogger(name) usage_collector = metrics.UsageCollector() # Increase verbosity for this module and RAG service try: logger.setLevel(logging.DEBUG) logging.getLogger("rag_service").setLevel(logging.DEBUG) except Exception: # Fallback silently if logger configuration fails pass class Assistant(Agent): def __init__(self) -> None: super().__init__( instructions="""You are a helpful voice AI assistant. You eagerly assist users with their questions by providing information from your extensive knowledge. Your responses are concise, to the point, and without any complex formatting or punctuation including emojis, asterisks, or other symbols. You are curious, friendly, and have a sense of humor. When answering questions, you will be provided with relevant information from a knowledge base. Use this information to enhance your responses, but maintain a natural conversational tone. If the retrieved information doesn't fully answer the question, use your general knowledge to provide a complete response.""", ) async def on_message(self, message: str) -> str: """ Override the on_message method to integrate RAG capabilities. Args: message: The user's message Returns: The assistant's response """ try: logger.info("on_message invoked. message_len=%d", len(message) if message is not None else -1) # Quick environment checks for RAG qdrant_url_set = bool(os.getenv("QDRANT_URL")) qdrant_key_set = bool(os.getenv("QDRANT_API_KEY")) embed_ep_set = bool(os.getenv("AZURE_OPENAI_EMBEDDING_ENDPOINT")) embed_key_set = bool(os.getenv("AZURE_OPENAI_EMBEDDING_API_KEY")) embed_deploy = os.getenv("AZURE_OPENAI_EMBEDDING_DEPLOYMENT", "") logger.info( "RAG env check: QDRANT_URL=%s, QDRANT_API_KEY=%s, AZURE_EMBED_ENDPOINT=%s, AZURE_EMBED_KEY=%s, DEPLOYMENT=%s", qdrant_url_set, qdrant_key_set, embed_ep_set, embed_key_set, embed_deploy, ) # Get the current participant from the session context #participant = self.session.context.participant # Extract user_id from participant metadata #user_id = "default_user" #if participant and participant.metadata: # try: # import json # meta = json.loads(participant.metadata) if participant.metadata else {} # # Use user_id from metadata if available # user_id = meta.get("user_id", participant.identity) # except Exception as e: # logger.warning(f"Failed to parse metadata: {e}") # # Fallback to participant identity if metadata parsing fails # user_id = participant.identity #logger.info(f"Using user_id for RAG: {user_id}") # Hardcode user_id to "1" as requested user_id = "1122" logger.info("Using hardcoded user_id for RAG: %s", user_id) # Retrieve relevant information from the RAG service try: # Log RAG service readiness rag_ready_qdrant = rag_service.qdrant_client is not None rag_ready_embed = rag_service.embedding_client is not None logger.info( "RAG service state before search: qdrant_ready=%s, embed_ready=%s", rag_ready_qdrant, rag_ready_embed, ) t0 = time.time() search_results = rag_service.search(query=message, user_id=user_id, limit=3) dt = time.time() - t0 logger.info( "RAG search returned %d results in %.3fs", len(search_results), dt, ) if not search_results: logger.warning( "RAG returned no results. Check tenant filter user_id=%s and collection contents.", user_id, ) else: # Log a sample of sources for quick visibility sample_sources = [r.get("source") for r in search_results[:3]] logger.info("RAG sample sources: %s", sample_sources) except Exception as e: logger.exception("RAG search raised an exception: %s", str(e)) search_results = [] # Format the search results as context for the LLM context = rag_service.format_for_context(search_results) logger.info("RAG context length (chars): %d", len(context) if context else 0) # Prepare the prompt with the retrieved context if context: logger.info("Using augmented prompt with RAG context") prompt = f"The user asked: {message}\n\n{context}\n\nBased on this information, provide a helpful response:" else: logger.info("No RAG context available; using original message as prompt") prompt = message # Let the parent class handle the LLM call return await super().on_message(prompt) except Exception as e: logger.exception(f"Error in RAG processing: {str(e)}") # Fallback to standard processing if RAG fails return await super().on_message(message) async def entrypoint(ctx: agents.JobContext): # Connect to the LiveKit room before interacting with participants await ctx.connect() # Wait for at least one participant to join participant = await ctx.wait_for_participant() try: import json meta = json.loads(participant.metadata) if participant.metadata else {} except Exception: meta = participant.metadata print("[agent] participant", participant.identity, "metadata:", meta) # Preload models for faster response times stt = deepgram.STT(model="nova-2") llm = openai.LLM( api_key=openai_api_key, ).with_azure( model="gpt-4o-mini", azure_endpoint=azure_endpoint, azure_deployment=azure_deployment, api_version="2024-08-01-preview" ) tts = deepgram.TTS(model="aura-asteria-en") vad = silero.VAD.load() noise_cancellation_filter = noise_cancellation.BVC() session = AgentSession( stt=stt, llm=llm, tts=tts, vad=vad, # turn_detection=MultilingualModel(), # Removed turn_detection to speed up initialization ) await session.start( room=ctx.room, agent=Assistant(), room_input_options=RoomInputOptions( # For telephony applications, use
BVCTelephony
instead for best results noise_cancellation=noise_cancellation_filter, ), ) # Collect metrics @session.on("metrics_collected") def on_metrics_collected(event: MetricsCollectedEvent): usage_collector.collect(event.metrics) await session.say( "Hello, I am checking RAG Today, how can I help you today?" ) # Log aggregated usage metrics when the session ends summary = usage_collector.get_summary() logger.info(f"Usage: {summary}") # Define a shutdown callback to log usage when the job is finalized async def log_usage(): summary = usage_collector.get_summary() logger.info(f"Usage summary on shutdown: {summary}") ctx.add_shutdown_callback(log_usage) if name == "__main__": agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint, port=9091)) Analyze whats wrong in this code.
on_message in livekit ?\
where does livekit store chat_context messages ? basically data?