아웃바운드 agent call 통화에서 사용자의 음성퀄리티가 너무 낮은데 통신사 이슈로 인...
# ask-ai
l
아웃바운드 agent call 통화에서 사용자의 음성퀄리티가 너무 낮은데 통신사 이슈로 인해 통화음질을 HD퀄리티로 올릴수가 없는데 input audio transription 을 할때 text로 처리되는 내용이 이상하게 나올때가 많아! 이문제를 해결할수 있는 방법이 있을까? 사용자의 음성을 크게 증폭시킨다거나 하는 방법이 없을까?
t
Thanks for reaching out, Dan(송희웅). I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I understand how frustrating low audio quality can be when it affects transcription accuracy. Let’s explore some ways you can improve this in LiveKit: • Enable Enhanced Noise Cancellation: You can use LiveKit Cloud's Background Voice Cancellation (BVC) model to boost transcription quality. This is easy to enable in the agent's
room_input_options
when starting your session. • Client-Side Noise Filtering: Consider applying advanced noise filtering with the
KrispNoiseFilter
from the
@livekit/krisp-noise-filter
package. It helps reduce background noise before the audio is transmitted. • Configure STT Noise Reduction: If you're using OpenAI's transcription API, setting the
noise_reduction_type
parameter to "near_field" or "far_field" depending on your microphone setup can make a big difference. • Default WebRTC Noise Suppression: You can also adjust WebRTC's
noiseSuppression
and
echoCancellation
settings through
AudioCaptureOptions
to improve audio clarity. • Audio Processing Module (APM): Setting up an APM with features like echo cancellation, noise suppression, high-pass filtering, and auto gain control can really enhance your audio quality. These approaches can help make your audio clearer and improve transcription results. If you'd like, I’m here to guide you through detailed steps for any of these methods—just let me know! Sources: Noise & echo cancellation | LiveKit Docs | Turn detection and interruptions | LiveKit Docs | components/room.tsx | livekit-plugins/livekit-plugins-gladia/livekit/plugins/gladia/stt.py
l
샘플코드를 보여줘
t
아래는 요청하신 샘플 코드입니다. 각 방법에 대한 예제를 포함했으니 편하게 참고해 주세요: • Client-Side Noise Filtering: `KrispNoiseFilter`를 사용하는 방법:
Copy code
import { KrispNoiseFilter, isKrispNoiseFilterSupported } from '@livekit/krisp-noise-filter';

if (isKrispNoiseFilterSupported()) {
  trackPublication.track?.setProcessor(KrispNoiseFilter());
}
• Audio Processing Module (APM): Python에서 APM을 설정하는 방법:
Copy code
from rtc import AudioProcessingModule

apm = AudioProcessingModule(
    echo_cancellation=True,
    noise_suppression=True,
    high_pass_filter=True,
    auto_gain_control=True,
)
• STT Noise Reduction: OpenAI의 실시간 전사 API를 사용할 때:
Copy code
{
  "noise_reduction_type": "near_field"  # 또는 "far_field"
}
이 샘플 코드들을 참고해서 설정해 보시고, 혹시 진행하시면서 궁금한 점이나 도움이 필요하시면 언제든지 편하게 말씀해 주세요. 함께 해결해 나가요!
l
Openai realtime api에서 input audio transcription을 사용하고 있는데 여기에서 할수 있는 방법은 없어?
t
OpenAI Realtime API를 사용하여 input audio transcription 품질을 개선하는 데 도움이 될 수 있는 몇 가지 방법을 함께 살펴볼게요: • Noise Cancellation 활성화: LiveKit의 내장 노이즈 캔슬링 기능을 활용하면 배경 소음을 효과적으로 줄여 전사 정확도를 높일 수 있어요. • 오디오 입력 설정 최적화: OpenAI Realtime API는 PCM16 오디오 형식과 24kHz 샘플 속도를 권장합니다. 입력 오디오가 이 사양에 맞게 설정되어 있는지 한번 확인해 보시면 좋을 것 같아요. • 적절한 전사 모델 사용: OpenAI에서는 `gpt-4o-mini-transcribe`와
whisper-1
같은 다양한 전사 모델을 제공하고 있어요. 특히
gpt-4o-mini-transcribe
모델은 LiveKit 통합에 최적화되어 있으니 참고해 주세요. • 언어 설정 구성: 가능하다면 입력 오디오의 언어를 명확히 지정해 주시면 전사 품질 향상에 큰 도움이 됩니다. STT 클래스에서는 언어 매개변수 설정과 자동 언어 감지 기능도 지원하고 있어요. • 적절한 오디오 스트리밍 구현: OpenAI로 오디오를 50ms 단위로 나누어 스트리밍하면 실시간 성능을 유지하면서도 충분한 오디오 컨텍스트를 제공해 정확한 전사가 가능해집니다. • Multimodal Agent API 사용: LiveKit의 Multimodal Agent API는 OpenAI Realtime API를 래핑하여 오디오와 텍스트 모달리티를 함께 처리하고, 텍스트 전사를 오디오 재생과 자연스럽게 동기화해 줍니다. • 실시간 기능 활용: Realtime API 사용 시 적절한 턴 감지 설정을 구성하면 대화 흐름이 훨씬 자연스러워져요. 서버 기반 음성 활동 감지(VAD) 설정도 가능하니 참고해 주세요. 이 방법들을 적용해 보시면 전사 품질이 훨씬 더 좋아질 거예요. 혹시 구현하시면서 궁금한 점이나 도움이 필요하시면 언제든지 편하게 말씀해 주세요! Sources: OpenAI and LiveKit | LiveKit Docs | livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/real... | OpenAI STT integration guide | LiveKit Docs
l
Server Vad 와 Semantic Vad 의 차이점에 대해서 알려줘.
t
Server VAD와 Semantic VAD의 차이점을 함께 살펴볼까요? 😊 • Server VAD: 기본 모드로, 사용자의 발화가 끝났는지 판단할 때 *침묵 시간*을 기준으로 해요. 오디오를 침묵 구간에 따라 청크로 나누는데요, 주요 설정은 다음과 같습니다: ◦ `threshold`: 오디오 볼륨에 대한 민감도 ◦ `prefix_padding_ms`: 감지된 발화 이전에 포함할 오디오 양 ◦ `silence_duration_ms`: 발화 종료를 감지하기 위한 침묵 지속 시간 • Semantic VAD: *의미 기반 분류기*를 사용해서 발화 내용을 이해하고, 사용자가 말을 끝냈는지 더 똑똑하게 판단해요. 단순히 침묵에만 의존하지 않아서, 사용자가 문장을 끝내기 전에 중단되거나 전사가 조각나는 걸 방지해 줍니다. 주요 설정은 다음과 같습니다: ◦ `eagerness`: 응답 타이밍을 조절해요. "low"는 좀 더 여유 있게 기다리고, "high"는 빠르게 응답하며, "medium"이나 "auto"는 균형 잡힌 접근을 제공합니다. 참고로, OpenAI Realtime API에서는 두 모드 모두 사용 가능하지만, Azure OpenAI는 현재 Server VAD만 지원하고 있답니다. 궁금한 점이나 도움이 필요하시면 언제든지 편하게 말씀해 주세요! Sources: Azure OpenAI Realtime API and LiveKit | LiveKit Docs | OpenAI Realtime API integration guide | LiveKit Docs | OpenAI Realtime API integration guide | LiveKit Docs
l
속도와 관련해서 더 유리한건 뭐야?
t
속도 측면에서 Server VAD가 Semantic VAD보다 일반적으로 더 빠릅니다. Server VAD는 단순히 *침묵 구간*을 감지하여 발화 종료를 판단하기 때문에, 의미 분석 없이도 빠르게 반응할 수 있어요.
silence_duration_ms
값을 짧게 설정하면 더욱 빠른 턴 감지가 가능하답니다. 반면에, Semantic VAD는 발화의 *의미와 구조*를 분석해서 사용자가 말을 끝냈는지 판단해요. 그래서 더 자연스러운 대화 흐름을 제공하지만, 그만큼 조금 더 시간이 걸릴 수 있어요.
eagerness
매개변수를 "high"로 설정하면 반응 속도를 높일 수 있지만, Server VAD만큼 빠르지는 않답니다. 그래서 속도가 가장 중요하신 경우에는 Server VAD를 사용하는 게 더 좋을 것 같아요. 혹시 더 궁금한 점 있으시면 언제든지 편하게 말씀해 주세요! Sources: OpenAI Realtime API integration guide | LiveKit Docs