quiet-airplane-65783
07/29/2025, 4:09 PMtall-airline-16756
07/29/2025, 4:09 PMtall-airline-16756
07/29/2025, 4:10 PMquiet-airplane-65783
07/29/2025, 4:13 PMtall-airline-16756
07/29/2025, 4:14 PMquiet-airplane-65783
07/29/2025, 4:15 PMtall-airline-16756
07/29/2025, 4:15 PMpip install livekit-plugins-noise-cancellation
1. Then, in your code, you'll want to configure the noise cancellation when starting your session. Check out this example:
await session.start(
agent=MyAgent(),
room=ctx.room,
room_input_options=RoomInputOptions(
noise_cancellation=noise_cancellation.BVC(),
),
room_output_options=RoomOutputOptions(transcription_enabled=True),
)
And just like that, you've got noise cancellation working for your SIP calls!
3. Use Enhanced STT Models
Want to take your transcription to the next level? Using models specifically designed for telephony can make a huge difference. If you're using OpenAI's STT service, you can easily adjust the noise reduction:
1. Just set the noise_reduction_type parameter in your STT setup. Here's a quick example:
stt = STT(
noise_reduction_type="far_field" # Options: "near_field" or "far_field"
)
Pro tip: If you're using LiveKit's BVCTelephony model, make sure it's selected - it's specifically optimized for phone conversations!
Need any more help working through these steps? I'm totally here for you - just give me a shout!quiet-airplane-65783
07/29/2025, 4:40 PMtall-airline-16756
07/29/2025, 4:41 PMasync def transcription_node(self, text: AsyncIterable[str], model_settings: ModelSettings):
"""Modify the transcription output by replacing certain words."""
replacements = {
"save": "sales",
"six": "sales",
}
for word in text:
yield replacements.get(word, word)
2. Customizable Speech Models:
How awesome would it be to have a speech model that totally gets your lingo? Providers like Deepgram let you train models to understand your specific context. It's like teaching a friend your unique communication style!
3. Optimized STT Providers:
Some speech-to-text providers are real pros at catching those quick, conversational words. Check out models like Cartesia's Ink-Whisper - they're designed to catch every word in real-time.
4. Frame Size Optimization:
This might sound technical, but it's basically about finding the perfect audio "snapshot" size. A little tweaking here can make a big difference in catching those single words accurately.
Hope these suggestions help! Feel free to reach out if you want to dive deeper into any of these approaches. I'm always happy to help!
Sources: Live Conversations with AI using ChatGPT and WebRTC | pipeline-llm/transcription_node.py | Deepgram STT integration guide | LiveKit Docs | Cartesia STT integration guide | LiveKit Docsquiet-airplane-65783
07/29/2025, 4:55 PMtall-airline-16756
07/29/2025, 4:55 PMquiet-airplane-65783
07/29/2025, 5:52 PMtall-airline-16756
07/29/2025, 5:53 PMquiet-airplane-65783
07/29/2025, 5:55 PMtall-airline-16756
07/29/2025, 5:55 PMpip install pydub
2. Example Code to Increase Gain:
from pydub import AudioSegment
from pydub.playback import play
# Load your audio file (e.g., from a WAV file or raw audio stream)
audio = AudioSegment.from_file("input_audio.wav")
# Increase the gain by 10 dB (adjust the value as needed)
amplified_audio = audio + 10
# Export the amplified audio to a new file
amplified_audio.export("amplified_audio.wav", format="wav")
# Optionally, play the amplified audio
play(amplified_audio)
3. Integrate with Your Pipeline:
Working with real-time audio streams? No worries! Here's a neat trick to process audio chunks on the fly:
from pydub import AudioSegment
def process_audio_chunk(audio_chunk):
# Convert raw audio chunk to PyDub AudioSegment
audio = AudioSegment(
data=audio_chunk,
sample_width=2, # Adjust based on your audio format
frame_rate=8000, # Adjust based on your codec (e.g., G.711 is 8kHz)
channels=1 # Mono audio
)
# Increase the gain
amplified_audio = audio + 10
# Return the processed audio as raw data
return amplified_audio.raw_data
This approach is perfect for amplifying audio in real-time before sending it off to transcription. Pretty cool, right?
Need any help setting this up? I'm totally here for you – just give me a shout!