i have two people speaking and i want the AI to li...
# ask-ai
s
i have two people speaking and i want the AI to listen to both and trasncript that, how that can happen
t
Thanks for reaching out, Mostafa. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To transcribe audio from two speakers using LiveKit AI, here’s a friendly guide to help you through the process: 1. Choose an STT Provider with Diarization Support: It’s great to use providers like Speechmatics or Deepgram that support speaker diarization. For instance, you can enable diarization by setting
enable_diarization=True
when you initialize the STT instance. 2. Use the MultiSpeakerAdapter: Wrapping your STT engine with the
MultiSpeakerAdapter
really helps improve speaker detection. You can configure it to format transcripts with speaker IDs and even choose to suppress background speakers if you like. Here’s an example configuration:
Copy code
stt = MultiSpeakerAdapter(
    stt=speechmatics.STT(enable_diarization=True),
    detect_primary_speaker=True,
    suppress_background_speaker=False,
    primary_format="<{speaker_id}>{text}",
    background_format="<{speaker_id}>{text}"
)
3. Integrate into an Agent Session: Once set up, use this STT instance within an
AgentSession
to process real-time audio and get transcripts that clearly show who’s speaking. 4. Process Transcripts: Your transcriptions will include speaker labels like
<s1>
and
<s2>
, making it easier to follow the conversation between speakers. 5. Optional: For even better clarity, you can map speaker IDs to participant names using tools like
set_speaker_name
. Just a quick reminder to ensure your STT service supports diarization and that your audio input is a single mixed track. If you have any questions or need a hand with the setup, please don’t hesitate to reach out—I’m here to help! Sources: examples/other/translation/multi-user-translator.py | pipeline-stt/diarization.py