@microscopic-mechanic-41538 are you trying to replicate KITT’s functionality or do something slightly different?
What I’m understanding from the above is you want to capture a user’s speech, transform their voice on the backend, then transmit that out to other users?
m
microscopic-mechanic-41538
09/10/2023, 5:23 PM
@magnificent-art-43333 It’s a little different.
We have a default audio stream (we read audio from a microphone and encode it in Opus)
The task is as follows:
1 - Subscribe to remoteParticipant track
2 - Turn speech to text (STT) from rp track
3 - Synthesize speech from text (TTS)
4 - Create a second audio stream where we send the synthesized speech of the user.
Total:
We have an audio stream with audio from the microphone
We have an audio stream with synthesized speech (if we close the stream, the main stream will be the one from the microphone).
This will allow the conversation partner to hear the original / synthesized voice
m
magnificent-art-43333
09/10/2023, 5:29 PM
I think I’m still not understanding what you’re trying to do. From a product perspective, what are you trying to build? Feel free to DM me if you don’t want to share it publicly. Is what I described in the previous message not accurate? What do you do with the original audio stream from the user?