steep-shampoo-86436
09/22/2025, 3:36 PMtall-airline-16756
09/22/2025, 3:36 PMtall-airline-16756
09/22/2025, 3:37 PMenable_diarization=True when you initialize the STT instance.
2. Use the MultiSpeakerAdapter: Wrapping your STT engine with the MultiSpeakerAdapter really helps improve speaker detection. You can configure it to format transcripts with speaker IDs and even choose to suppress background speakers if you like. Here’s an example configuration:
stt = MultiSpeakerAdapter(
stt=speechmatics.STT(enable_diarization=True),
detect_primary_speaker=True,
suppress_background_speaker=False,
primary_format="<{speaker_id}>{text}",
background_format="<{speaker_id}>{text}"
)
3. Integrate into an Agent Session: Once set up, use this STT instance within an AgentSession to process real-time audio and get transcripts that clearly show who’s speaking.
4. Process Transcripts: Your transcriptions will include speaker labels like <s1> and <s2>, making it easier to follow the conversation between speakers.
5. Optional: For even better clarity, you can map speaker IDs to participant names using tools like set_speaker_name.
Just a quick reminder to ensure your STT service supports diarization and that your audio input is a single mixed track. If you have any questions or need a hand with the setup, please don’t hesitate to reach out—I’m here to help!
Sources: examples/other/translation/multi-user-translator.py | pipeline-stt/diarization.py