how can I detect the agents transcription and the ...
# ask-ai
c
how can I detect the agents transcription and the users transcriptions properly and attribute them to the correct owner in the react frontend
t
Thanks for reaching out, Aaron. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To detect and attribute agent and user transcriptions correctly in a React frontend using LiveKit, you can use the following approaches: 1. Using
useVoiceAssistant
and `useTrackTranscription`This method combines the
useVoiceAssistant
hook for agent transcriptions with
useTrackTranscription
for user transcriptions:
Copy code
import { useTrackTranscription, useVoiceAssistant } from "@livekit/components-react";
import { useMemo } from "react";
import useLocalMicTrack from "./useLocalMicTrack";

export default function useCombinedTranscriptions() {
  const { agentTranscriptions } = useVoiceAssistant();

  const micTrackRef = useLocalMicTrack();
  const { segments: userTranscriptions } = useTrackTranscription(micTrackRef);

  const combinedTranscriptions = useMemo(() => {
    return [
      ...agentTranscriptions.map((val) => {
        return { ...val, role: "assistant" };
      }),
      ...userTranscriptions.map((val) => {
        return { ...val, role: "user" };
      }),
    ].sort((a, b) => a.firstReceivedTime - b.firstReceivedTime);
  }, [agentTranscriptions, userTranscriptions]);

  return combinedTranscriptions;
}
2. Manual Track AttributionFor more control, you can manually track transcriptions for specific participants:
Copy code
export function TranscriptionTile({
  agentAudioTrack,
  accentColor,
}: {
  agentAudioTrack?: TrackReferenceOrPlaceholder;
  accentColor: string;
}) {
  const agentMessages = useTrackTranscription(agentAudioTrack || undefined);
  const localParticipant = useLocalParticipant();
  const localMessages = useTrackTranscription({
    publication: localParticipant.microphoneTrack,
    source: Track.Source.Microphone,
    participant: localParticipant.localParticipant,
  });

  // Attribution logic
  function segmentToChatMessage(
    s: TranscriptionSegment,
    existingMessage: ChatMessageType | undefined,
    participant: Participant,
  ): ChatMessageType {
    const msg: ChatMessageType = {
      message: s.final ? s.text : `${s.text} ...`,
      name: participant instanceof LocalParticipant ? "You" : "Agent",
      isSelf: participant instanceof LocalParticipant,
      timestamp: existingMessage?.timestamp ?? Date.now(),
    };
    return msg;
  }
}
3. Room Event-Based AttributionFor comprehensive transcription handling, listen to room events and attribute based on participant properties:
Copy code
useEffect(() => {
  if (!room) {
    return;
  }
  const updateRawSegments = (
    segments: TranscriptionSegment[],
    participant?: Participant,
    publication?: TrackPublication,
  ) => {
    setRawSegments((prev) => {
      const newSegments = { ...prev };
      for (const segment of segments) {
        newSegments[segment.id] = { segment, participant, publication };
      }
      return newSegments;
    });
  };
  room.on(RoomEvent.TranscriptionReceived, updateRawSegments);

  return () => {
    room.off(RoomEvent.TranscriptionReceived, updateRawSegments);
  };
}, [room]);
Key Attribution Principles: • Transcriptions are delivered in segments, each associated with a particular
Participant
and
Track
. • Use
useVoiceAssistant
for agent transcriptions and
useTrackTranscription
for user transcriptions. • Sort combined results by
firstReceivedTime
to maintain chronological order. Let me know if you need further clarification or assistance! Sources: src/transcriptions/TranscriptionTile.tsx | web/src/hooks/use-agent.tsx | AI voice agents | LiveKit Docs | Transcriptions | LiveKit Docs | packages/react/src/hooks/useVoiceAssistant.ts
c
transcripts dont properly attribute to the correct speaker. How can I ensure this happens?
t
To make sure your transcripts are correctly linked to the right speaker in LiveKit, I've got some friendly tips to help you out! 1. Enable Speaker Diarization If you're working with an awesome STT provider like Speechmatics, you can easily distinguish between speakers. Here's a quick way to do it:
Copy code
stt = speechmatics.STT(
    transcription_config=speechmatics.types.TranscriptionConfig(
        language="en",
        operating_point="enhanced",
        enable_partials=True,
        max_delay=0.7,
        diarization="speaker",
        speaker_diarization_config={"max_speakers": 2}, # Adjust as needed
    )
)
Cool part? The results will include a speaker_id for each transcription event. 2. Map Speaker IDs to Names Let's make those speaker IDs more personal! You can create a simple system to give them friendly names:
Copy code
@function_tool
async def set_speaker_name(self, context: RunContext[SpeechmaticsUserData], speaker_id: str, name: str):
    context.userdata.speaker_names[speaker_id] = name
    return f"I'll now refer to speaker {speaker_id} as {name}."
3. Process Speaker Attribution Now, let's add some magic to track who's speaking:
Copy code
async def process_stream():
    current_speaker_id = None
    async for event in parent_stream:
        if hasattr(event, 'type') and str(event.type) == "SpeechEventType.FINAL_TRANSCRIPT" and event.alternatives:
            transcript = event.alternatives[0].text
            speaker_id = event.alternatives[0].speaker_id

            if current_speaker_id != speaker_id:
                current_speaker_id = speaker_id
                speaker_name = context.userdata.speaker_names.get(speaker_id, speaker_id)
                modified_transcript = f"{speaker_name}: {transcript}"
                event.alternatives[0].text = modified_transcript
4. Fixed Participant Attribution Want to focus on a specific participant? Here's a neat trick:
Copy code
session = AgentSession(llm=openai.realtime.RealtimeModel())
room_io = RoomIO(session, room=ctx.room, participant="participant-identity")
await room_io.start()
await session.start(agent=MyAgent())
5. Frontend Integration On the frontend, transcriptions are super smart and automatically link to participants. Just use hooks like useVoiceAssistant for agents and useTrackTranscription for users. The system will handle all the name mapping for you! Hope this helps make your speaker tracking smooth and easy. Got questions? I'm always here to help! 😊 Sources: Speechmatics STT integration guide | LiveKit Docs | pipeline-stt/diarization.py | How to set a fixed participant for agent audio processing | src/transcriptions/TranscriptionTile.tsx | livekit-agents/livekit/agents/voice/room_io/room_io.py
c
I'm using deepgram, do you have examples with deepgram
t
Hey there! šŸ‘‹ Let me walk you through how to make sure your transcripts are perfectly attributed to the right speakers using Deepgram in LiveKit. It's actually pretty straightforward! 1. Enable Diarization in Deepgram Deepgram has this awesome feature called speaker diarization that makes tracking who's saying what super easy. Just set the diarize parameter to true. Check out this simple example:
Copy code
const deepgram = new Deepgram(DEEPGRAM_API_KEY);
const options = {
  punctuate: true,
  diarize: true, // Enable speaker diarization
  model: 'general',
};

const response = await deepgram.transcription.preRecorded({
  url: 'audio-file-url',
}, options);

response.results.channels[0].alternatives.forEach((alternative) => {
  alternative.words.forEach((word) => {
    console.log(`Speaker ${word.speaker}: ${word.word}`);
  });
});
2. Process Speaker Attribution Deepgram makes it super simple by providing a speaker property in the results. You can easily track who's saying what:
Copy code
response.results.channels[0].alternatives.forEach((alternative) => {
  alternative.words.forEach((word) => {
    console.log(`Speaker ${word.speaker}: ${word.word}`);
  });
});
3. Map Speaker IDs to Names Want to make things even more user-friendly? Map those speaker IDs to actual names! Here's a quick way to do that:
Copy code
const speakerNames = {
  0: 'Agent',
  1: 'User',
};

response.results.channels[0].alternatives.forEach((alternative) => {
  alternative.words.forEach((word) => {
    const speakerName = speakerNames[word.speaker] || `Speaker ${word.speaker}`;
    console.log(`${speakerName}: ${word.word}`);
  });
});
4. Frontend Integration On the frontend, the useTrackTranscription hook is your best friend for handling user transcriptions and matching them with the right speaker. Note: If you're using Deepgram's Python SDK, just double-check the diarization support – make sure you're running the latest version. Need more help or want to dive deeper into a specific setup? I'm all ears! Just shoot me a message, and we'll figure it out together. 😊 Sources: plugins/deepgram/src/stt.ts | livekit-agents/livekit/agents/voice/room_io/room_io.py | pipeline-stt/diarization.py | Speechmatics STT integration guide | LiveKit Docs | Speechmatics integration guide | LiveKit Docs