i have two people speaking from two different devi...
# ask-ai
s
i have two people speaking from two different devices but they are in the same livekit room, and i want the AI to listen to both and trasncript that, how that can happen
t
Thanks for reaching out, Mostafa. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Hi there! Just checking in—are you using a particular AI transcription service or tool for your transcription needs? Also, would you like some help with capturing audio streams from the LiveKit room or maybe integrating with a transcription API? I'm here to help you through it!
s
i'm not using specfic thing, you recomend what should i do
t
I'm happy to help you transcribe audio from multiple participants in a LiveKit room! Here’s a friendly step-by-step guide you can follow: 1. Subscribe to Audio Tracks: When participants join and publish their audio, you can use the
track_subscribed
event to subscribe to their audio tracks. 2. Stream Audio to a Speech-to-Text (STT) Service: For each participant’s audio track, create an
AudioStream
and send the audio frames to an STT engine like Deepgram or OpenAI Whisper. 3. Handle Transcription Output: The STT engine will process the audio and generate transcription results. You can then forward these results back to the room using LiveKit’s
STTSegmentsForwarder
. 4. Enable Transcription Forwarding: Using the
RoomIO
class in the LiveKit Agents framework, you can automatically publish transcriptions to the room so everyone can access them. If you’d like, I’m here to walk you through any of these steps in more detail or provide some helpful code snippets. Just let me know how I can support you! Sources: livekit-agents/livekit/agents/voice/room_io/room_io.py | Transcriptions | LiveKit Docs | main.py
s
so agent for each user, and will transcript the audio, but how can i mix the transcript back and send it to llm
t
Thanks for reaching out! Could you please share a bit more about how you'd like the transcriptions combined? For instance, are you looking to have them merged into one continuous text, or would you prefer a structured format that includes speaker identifiers and timestamps? Also, it would be great to know which LLM you're planning to use—I'm here to help if you need any guidance on working with its API. Looking forward to assisting you!
s
i want to send coninous to the llm, the transcriptions, identified by user id, so the llm can read the transcript and alert the user x or y if there is a violance of the policy
t
I'm glad you're looking into sending continuous transcriptions with user IDs to an LLM for policy violation detection. Here's a friendly guide to help you through the process: 1. Enable Transcription with User Identification: It’s important to use a speech-to-text (STT) service that supports diarization, so you can clearly associate each transcription with the right user ID. 2. Capture Transcriptions in Real Time: On the frontend, you can tap into the
RoomEvent.DataReceived
event to grab transcriptions as they come in. Alternatively, on the agent side, the
conversation_item_added
event is great for processing user messages the moment they’re transcribed. 3. Handle Partial and Final Transcripts: For real-time policy checks, it’s helpful to work with both interim and final transcriptions. Events like
input_audio_transcription_completed
let you catch updates from the STT engine smoothly. 4. Send Transcriptions to LLM: Once you have the transcription data paired with user IDs, format it nicely and send it over to your LLM using its API. Don’t forget to include a clear prompt that guides the LLM to look out for policy violations and provide the results. If you’d like, I’m here to help with any specific implementation details—whether it’s integrating with a particular LLM or setting up your transcription workflow. Just let me know, and we’ll work through it together! Sources: Text and transcriptions | LiveKit Docs | How to capture real-time agent transcriptions during conversations | agent/main.py | livekit-plugins/livekit-plugins-aws/livekit/plugins/aws/experimental/realti...