chilly-soccer-432
07/24/2025, 7:17 PMtall-airline-16756
07/24/2025, 7:17 PMtall-airline-16756
07/24/2025, 7:18 PMuseVoiceAssistant and `useTrackTranscription`This method combines the useVoiceAssistant hook for agent transcriptions with useTrackTranscription for user transcriptions:
import { useTrackTranscription, useVoiceAssistant } from "@livekit/components-react";
import { useMemo } from "react";
import useLocalMicTrack from "./useLocalMicTrack";
export default function useCombinedTranscriptions() {
const { agentTranscriptions } = useVoiceAssistant();
const micTrackRef = useLocalMicTrack();
const { segments: userTranscriptions } = useTrackTranscription(micTrackRef);
const combinedTranscriptions = useMemo(() => {
return [
...agentTranscriptions.map((val) => {
return { ...val, role: "assistant" };
}),
...userTranscriptions.map((val) => {
return { ...val, role: "user" };
}),
].sort((a, b) => a.firstReceivedTime - b.firstReceivedTime);
}, [agentTranscriptions, userTranscriptions]);
return combinedTranscriptions;
}
2. Manual Track AttributionFor more control, you can manually track transcriptions for specific participants:
export function TranscriptionTile({
agentAudioTrack,
accentColor,
}: {
agentAudioTrack?: TrackReferenceOrPlaceholder;
accentColor: string;
}) {
const agentMessages = useTrackTranscription(agentAudioTrack || undefined);
const localParticipant = useLocalParticipant();
const localMessages = useTrackTranscription({
publication: localParticipant.microphoneTrack,
source: Track.Source.Microphone,
participant: localParticipant.localParticipant,
});
// Attribution logic
function segmentToChatMessage(
s: TranscriptionSegment,
existingMessage: ChatMessageType | undefined,
participant: Participant,
): ChatMessageType {
const msg: ChatMessageType = {
message: s.final ? s.text : `${s.text} ...`,
name: participant instanceof LocalParticipant ? "You" : "Agent",
isSelf: participant instanceof LocalParticipant,
timestamp: existingMessage?.timestamp ?? Date.now(),
};
return msg;
}
}
3. Room Event-Based AttributionFor comprehensive transcription handling, listen to room events and attribute based on participant properties:
useEffect(() => {
if (!room) {
return;
}
const updateRawSegments = (
segments: TranscriptionSegment[],
participant?: Participant,
publication?: TrackPublication,
) => {
setRawSegments((prev) => {
const newSegments = { ...prev };
for (const segment of segments) {
newSegments[segment.id] = { segment, participant, publication };
}
return newSegments;
});
};
room.on(RoomEvent.TranscriptionReceived, updateRawSegments);
return () => {
room.off(RoomEvent.TranscriptionReceived, updateRawSegments);
};
}, [room]);
Key Attribution Principles:
⢠Transcriptions are delivered in segments, each associated with a particular Participant and Track.
⢠Use useVoiceAssistant for agent transcriptions and useTrackTranscription for user transcriptions.
⢠Sort combined results by firstReceivedTime to maintain chronological order.
Let me know if you need further clarification or assistance!
Sources: src/transcriptions/TranscriptionTile.tsx | web/src/hooks/use-agent.tsx | AI voice agents | LiveKit Docs | Transcriptions | LiveKit Docs | packages/react/src/hooks/useVoiceAssistant.tschilly-soccer-432
07/24/2025, 8:01 PMtall-airline-16756
07/24/2025, 8:02 PMstt = speechmatics.STT(
transcription_config=speechmatics.types.TranscriptionConfig(
language="en",
operating_point="enhanced",
enable_partials=True,
max_delay=0.7,
diarization="speaker",
speaker_diarization_config={"max_speakers": 2}, # Adjust as needed
)
)
Cool part? The results will include a speaker_id for each transcription event. 2. Map Speaker IDs to Names Let's make those speaker IDs more personal! You can create a simple system to give them friendly names:
@function_tool
async def set_speaker_name(self, context: RunContext[SpeechmaticsUserData], speaker_id: str, name: str):
context.userdata.speaker_names[speaker_id] = name
return f"I'll now refer to speaker {speaker_id} as {name}."
3. Process Speaker Attribution Now, let's add some magic to track who's speaking:
async def process_stream():
current_speaker_id = None
async for event in parent_stream:
if hasattr(event, 'type') and str(event.type) == "SpeechEventType.FINAL_TRANSCRIPT" and event.alternatives:
transcript = event.alternatives[0].text
speaker_id = event.alternatives[0].speaker_id
if current_speaker_id != speaker_id:
current_speaker_id = speaker_id
speaker_name = context.userdata.speaker_names.get(speaker_id, speaker_id)
modified_transcript = f"{speaker_name}: {transcript}"
event.alternatives[0].text = modified_transcript
4. Fixed Participant Attribution Want to focus on a specific participant? Here's a neat trick:
session = AgentSession(llm=openai.realtime.RealtimeModel())
room_io = RoomIO(session, room=ctx.room, participant="participant-identity")
await room_io.start()
await session.start(agent=MyAgent())
5. Frontend Integration On the frontend, transcriptions are super smart and automatically link to participants. Just use hooks like useVoiceAssistant for agents and useTrackTranscription for users. The system will handle all the name mapping for you! Hope this helps make your speaker tracking smooth and easy. Got questions? I'm always here to help! š Sources: Speechmatics STT integration guide | LiveKit Docs | pipeline-stt/diarization.py | How to set a fixed participant for agent audio processing | src/transcriptions/TranscriptionTile.tsx | livekit-agents/livekit/agents/voice/room_io/room_io.pychilly-soccer-432
07/24/2025, 8:03 PMtall-airline-16756
07/24/2025, 8:06 PMconst deepgram = new Deepgram(DEEPGRAM_API_KEY);
const options = {
punctuate: true,
diarize: true, // Enable speaker diarization
model: 'general',
};
const response = await deepgram.transcription.preRecorded({
url: 'audio-file-url',
}, options);
response.results.channels[0].alternatives.forEach((alternative) => {
alternative.words.forEach((word) => {
console.log(`Speaker ${word.speaker}: ${word.word}`);
});
});
2. Process Speaker Attribution Deepgram makes it super simple by providing a speaker property in the results. You can easily track who's saying what:
response.results.channels[0].alternatives.forEach((alternative) => {
alternative.words.forEach((word) => {
console.log(`Speaker ${word.speaker}: ${word.word}`);
});
});
3. Map Speaker IDs to Names Want to make things even more user-friendly? Map those speaker IDs to actual names! Here's a quick way to do that:
const speakerNames = {
0: 'Agent',
1: 'User',
};
response.results.channels[0].alternatives.forEach((alternative) => {
alternative.words.forEach((word) => {
const speakerName = speakerNames[word.speaker] || `Speaker ${word.speaker}`;
console.log(`${speakerName}: ${word.word}`);
});
});
4. Frontend Integration On the frontend, the useTrackTranscription hook is your best friend for handling user transcriptions and matching them with the right speaker. Note: If you're using Deepgram's Python SDK, just double-check the diarization support ā make sure you're running the latest version. Need more help or want to dive deeper into a specific setup? I'm all ears! Just shoot me a message, and we'll figure it out together. š Sources: plugins/deepgram/src/stt.ts | livekit-agents/livekit/agents/voice/room_io/room_io.py | pipeline-stt/diarization.py | Speechmatics STT integration guide | LiveKit Docs | Speechmatics integration guide | LiveKit Docs