This message was deleted.
# helpdesk
s
This message was deleted.
a
Depends on what you want to do. If you want to look at all the tracks separately, then, yes, you’ll need a separate TrackEgress per participant. But you can also use Room Composite Egress to get all the tracks rendered at once. When I’ve worked on captioning/transcription in the past, one difficulty was separating out voice tracks to identify the speaker — the algorithms were terrible at processing crosstalk. But that was several years ago, and I understand the state of the art has progressed considerably. To zoom out a bit, are you doing real-time captioning, or producing captions after the fact? I ask because egress may not be ideal for real-time captioning; for that use case it might be better to run the captioning process as a LiveKit client that connects to the room as a Participant.
n
Ideally, we aim to cover both use-cases: real-time captioning (with some reasonable delay of course) and transcripts by a single approach. TrackEgress seems to us a good candidate at first glance: it can be relatively easy integrated with speech-to-text services via WebSockets, we already familiar with the concept by using RoomCompositeEgress for recordings, also Livekit team recommended it as to-go approach for captioning in the documentation. Documentation also suggests the way how to distinguish speakers by a query parameter :
<wss://your-server.com/egress?trackID=><trackID>&participant=<participantIdentity>
(so we really didn’t thought it can be an issue) And at this point I got my question: we potentially have very big room with most of the participants muted (but can be unmuted and say something at any moment). We should transfer TrackId in TrackEgressRequest, so this setting implies having TrackEgress per participant. That’s how I came up with this raw idea of having ActiveSpeakerTrackEgress. We do not mind captioning a single active participant -- the loudest 🙂. Do you see pitfalls in case we decide to chose this path? The problems you are implying specifically are… delay? RoomComposite seems nice and familiar, just the chain of technologies looks bulkier: Livekit -> Headless Chrome + Gstreamer -> RTMP -> Speech-to-text.
a
RoomComposite is bulkier than a single track egress, but track egress on 100 participants is going to be much higher overhead if people are simply muting their track after speaking (the track egress will then sit there waiting in case they unmute). In the future we plan on adding something like a RoomAudio track where all the audio will be mixed into one, which could then be used for track egress, however we do not have a timeline set for that yet
n
thank you for your answers!