Hey - I'm running a voice AI agent on LiveKit with...
# ask-ai
q
Hey - I'm running a voice AI agent on LiveKit with SIP telephony integration and Plivo as my SIP trunk provider. However, the transcription quality is very poor - any idea on how this can be improved? Using Plivo in India - I do not think there is any method to improve the incoming audio quality. What i have tried so far - 1. Tried adding a gain function in the STT node before transcription happens. No benefit 2. Tried different STT providers - Deepgram Nova 3 and Nova 2 phonecall provide the best transcription, but the transcription quality is still unreasonable for single word responses.
t
Thanks for reaching out, Shreyas. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Could you clarify a few details to help address the transcription quality issue? • Are you using any audio preprocessing techniques (e.g., noise cancellation, echo suppression) before sending audio to the STT provider? • What is the audio codec being used in your SIP telephony setup? • Are you using LiveKit's enhanced noise cancellation audio filters, and if not, would you like guidance on enabling them?
q
1. Not using noise cancellation or echo suppression currently. 2. G.711 U-law (PCMU) 3. no - noise is not an issue currently.
t
Hey there! I wanted to share some friendly tips to help boost the transcription quality for your LiveKit voice AI agent with SIP telephony. I think these suggestions could make a real difference in how clearly your system captures conversations. • Enable Noise Cancellation: Even if background noise isn't driving you crazy right now, turning on noise cancellation can be a game-changer. LiveKit has this awesome enhanced noise cancellation for SIP-based telephony that can really clean up the audio. I'd be happy to walk you through setting it up if you'd like! • Upgrade from G.711 Codec: Think of the G.711 codec like an old radio – it gets the job done, but the sound isn't super crisp. Switching to a wideband codec like G.722 is like upgrading to a high-definition speaker. It'll give you much clearer audio, often called HD Voice. Just check with your SIP provider to see if they support this upgrade. • Use Enhanced STT Models: For phone conversations, there are special speech-to-text models that are total pros at handling telephony audio. The BVCTelephony model is a great example. If you're using OpenAI's service, you can also tweak noise reduction settings to make things even clearer. Right now, your G.711 setup with no noise cancellation is a bit like listening to a distant radio. But with these changes, I'm confident you'll hear a huge improvement in your transcription quality. Want to give it a try? Sources: Noise & echo cancellation | LiveKit Docs | Enhanced noise cancellation | LiveKit Docs | SIP overview | LiveKit Docs | Audio and video | LiveKit Docs | HD voice for SIP | LiveKit Docs
q
can you walk me through the 1st and 3rd suggestion
t
Hey there! Let me walk you through a couple of awesome ways to boost your transcription quality that I think you'll find super helpful. 1. Enable Noise Cancellation LiveKit has this cool noise cancellation feature for SIP-based calls that can really clean up the audio. Here's how you can set it up: 1. First, grab the noise cancellation plugin with this quick command:
Copy code
pip install livekit-plugins-noise-cancellation
1. Then, in your code, you'll want to configure the noise cancellation when starting your session. Check out this example:
Copy code
await session.start(
    agent=MyAgent(),
    room=ctx.room,
    room_input_options=RoomInputOptions(
        noise_cancellation=noise_cancellation.BVC(),
    ),
    room_output_options=RoomOutputOptions(transcription_enabled=True),
)
And just like that, you've got noise cancellation working for your SIP calls! 3. Use Enhanced STT Models Want to take your transcription to the next level? Using models specifically designed for telephony can make a huge difference. If you're using OpenAI's STT service, you can easily adjust the noise reduction: 1. Just set the
noise_reduction_type
parameter in your STT setup. Here's a quick example:
Copy code
stt = STT(
    noise_reduction_type="far_field"  # Options: "near_field" or "far_field"
)
Pro tip: If you're using LiveKit's BVCTelephony model, make sure it's selected - it's specifically optimized for phone conversations! Need any more help working through these steps? I'm totally here for you - just give me a shout!
q
I tried using the Telephony Noise Cancellation, but I still don't have better results. I'll give you an example. When I say "sales", the transcription model either hears "save" or the word "six". What can be done? Can you give me some innovative ideas that I might not be trying already?
t
Hey there! I totally get how frustrating it can be when transcription accuracy goes a bit wonky, especially with short words like "sales". Let me share some friendly tips that might just save the day! 1. Post-Processing Text Replacement: Think of this like a little translation helper. We can set up a smart system that catches those tricky misheard words and swaps them out. Imagine it's like having a buddy who gently corrects you. Here's a cool way to do that:
Copy code
async def transcription_node(self, text: AsyncIterable[str], model_settings: ModelSettings):
    """Modify the transcription output by replacing certain words."""
    replacements = {
        "save": "sales",
        "six": "sales",
    }
    for word in text:
        yield replacements.get(word, word)
2. Customizable Speech Models: How awesome would it be to have a speech model that totally gets your lingo? Providers like Deepgram let you train models to understand your specific context. It's like teaching a friend your unique communication style! 3. Optimized STT Providers: Some speech-to-text providers are real pros at catching those quick, conversational words. Check out models like Cartesia's Ink-Whisper - they're designed to catch every word in real-time. 4. Frame Size Optimization: This might sound technical, but it's basically about finding the perfect audio "snapshot" size. A little tweaking here can make a big difference in catching those single words accurately. Hope these suggestions help! Feel free to reach out if you want to dive deeper into any of these approaches. I'm always happy to help! Sources: Live Conversations with AI using ChatGPT and WebRTC | pipeline-llm/transcription_node.py | Deepgram STT integration guide | LiveKit Docs | Cartesia STT integration guide | LiveKit Docs
q
Can we talk about frame size optimization? How would that work?
t
Let's chat about frame size optimization - it's like finding the perfect recipe for crystal-clear audio transcription! Think of it as adjusting the "slice size" of your audio data to get the best possible results. How It Works: • Imagine breaking down your audio into tiny segments, kind of like cutting a cake into different sized pieces. Each segment (or frame) is a small chunk of sound, usually measured in milliseconds. • Smaller frames (like 20ms) mean faster processing - it's like getting quick updates. But just like a short conversation can miss nuances, these tiny frames might not catch all the audio details. • Larger frames (around 100ms) are like having a more in-depth chat. They give the transcription service more context, which can help it understand your words more accurately. Steps to Optimize Frame Size: 1. First, figure out what frame size you're currently using. It's usually tucked away in your audio settings. 2. Time to experiment! If you're using 20ms frames, try bumping it up to 50ms or 100ms and see how it changes things. 3. Test these changes with your specific needs. Pay special attention to short responses like "sales" to see how the accuracy shifts. 4. Once you find your sweet spot, update your audio settings to match. Need a hand figuring this out? I'm totally here to help you navigate this audio optimization adventure!
q
Even this doesn't help. What else could be the issue here? Is there any way to increase the gain of the audio so that the input audio just sounds louder?
t
Hey there! I totally understand your challenge with audio gain in SIP telephony. While LiveKit doesn't have a direct solution, I've got some friendly tips that might help you out: • Use Auto Gain Control (AGC): Think of this like a helpful audio buddy that automatically adjusts your volume levels. If your audio processing setup supports it, this could be a game-changer! • Preprocess Audio Before Transcription: No worries - you can totally boost your audio signal before sending it off. Tools like PyDub or FFmpeg are super handy for giving your audio that extra volume lift. • Switch to Wideband Codecs: This won't directly increase gain, but codecs like G.722 or Opus can make your audio sound clearer and crisper, which might make a big difference. Want to dive deeper into any of these? I'm here to help and walk you through the details. Just let me know what sounds good to you! Sources: Configuring HD Voice for SIP | LiveKit Docs | livekit-android-sdk/src/main/java/io/livekit/android/audio/ScreenAudioCaptu... | livekit-rtc/livekit/rtc/apm.py
q
How can pre-processing be done? I am using Python, can you give me some code examples?
t
Hey there! Want to boost your audio's volume using Python? I've got a super handy solution with PyDub that'll make your sound crystal clear. Let me walk you through it like we're chatting over coffee! 😊 1. Install PyDub:
Copy code
pip install pydub
2. Example Code to Increase Gain:
Copy code
from pydub import AudioSegment
from pydub.playback import play

# Load your audio file (e.g., from a WAV file or raw audio stream)
audio = AudioSegment.from_file("input_audio.wav")

# Increase the gain by 10 dB (adjust the value as needed)
amplified_audio = audio + 10

# Export the amplified audio to a new file
amplified_audio.export("amplified_audio.wav", format="wav")

# Optionally, play the amplified audio
play(amplified_audio)
3. Integrate with Your Pipeline: Working with real-time audio streams? No worries! Here's a neat trick to process audio chunks on the fly:
Copy code
from pydub import AudioSegment

def process_audio_chunk(audio_chunk):
    # Convert raw audio chunk to PyDub AudioSegment
    audio = AudioSegment(
        data=audio_chunk,
        sample_width=2,  # Adjust based on your audio format
        frame_rate=8000,  # Adjust based on your codec (e.g., G.711 is 8kHz)
        channels=1  # Mono audio
    )

    # Increase the gain
    amplified_audio = audio + 10

    # Return the processed audio as raw data
    return amplified_audio.raw_data
This approach is perfect for amplifying audio in real-time before sending it off to transcription. Pretty cool, right? Need any help setting this up? I'm totally here for you – just give me a shout!