lively-horse-90722
07/31/2025, 3:25 PMtall-airline-16756
07/31/2025, 3:25 PMtall-airline-16756
07/31/2025, 3:26 PMVADEventType.START_OF_SPEECH is not firing while vad.VADEventType.INFERENCE_DONE fires directly indicates that the VAD is processing audio but not detecting speech above the configured thresholds. Here are steps to troubleshoot and resolve this:
1. Check Activation Threshold:
Ensure the activation_threshold is not set too high. Lowering it can help detect speech more easily.
// Example in JavaScript
const vad = await VAD.load({
activationThreshold: 0.3, // Lower from default 0.5
minSpeechDuration: 50, // Default is 50ms
});
2. Adjust Minimum Speech Duration:
If speech segments are too short, START_OF_SPEECH won’t fire. Reduce the minSpeechDuration value.
// Example in JavaScript
vad.updateOptions({
activationThreshold: 0.3,
minSpeechDuration: 20
});
3. Monitor Speech Probability:
Check the probability values in INFERENCE_DONE events to understand what the VAD is detecting.
// Example in JavaScript
for await (const event of vadStream) {
if (event.type === VADEventType.INFERENCE_DONE) {
console.log(`Speech probability: ${event.probability}`);
console.log(`Speaking state: ${event.speaking}`);
}
}
4. Check Audio Input Quality:
Ensure the audio input has sufficient volume and quality. The sample rate should be 8kHz or 16kHz, and the audio should not be too quiet or heavily processed.
5. Debug with Test Audio:
Use test patterns to verify the VAD is working as expected.
// Example in Python
frames, _ = await utils.make_test_speech(sample_rate=16000)
stream = vad.stream()
for frame in frames:
stream.push_frame(frame)
stream.end_input()
async for ev in stream:
print(f"Event: {ev.type}, Probability: {ev.probability}")
By following these steps, you should be able to identify and resolve the issue. Let me know if you need further assistance!
Sources: livekit-agents/livekit/agents/stt/stream_adapter.pylively-horse-90722
07/31/2025, 3:32 PMVAD is processing audio but not detecting speech above the configured thresholds.
Then how come INFERENCE_DONE is firedtall-airline-16756
07/31/2025, 3:34 PMINFERENCE_DONE event. It actually fires for every inference window that the VAD model processes, whether or not it detects speech above the activation threshold. This is totally on purpose and helps in a few important ways:
• Regular Heartbeat: The INFERENCE_DONE event gives you steady updates, so you know the VAD is actively working on the audio.
• Debugging and Monitoring: It lets you peek at the raw probability values, even when they’re below that speech threshold.
• Metrics Collection: The system uses these events to gather useful VAD stats, like how long each inference takes and how many have happened.
• Fine-Grained Control: You can build your own custom logic based on those probability values, even if they’re below the threshold.
Here’s how the event firing sequence works:
1. The VAD model processes a chunk of audio and calculates a speech probability.
2. The INFERENCE_DONE event fires right after, carrying the probability and some other info.
3. Only after that does the system check if the probability passes the activation_threshold to decide whether to fire START_OF_SPEECH or END_OF_SPEECH events.
So, even if the probability is below the threshold, the INFERENCE_DONE event still fires to keep you in the loop about what the VAD is doing. In those cases, the speaking field will be false, but the event’s there to make sure you get consistent updates and can debug smoothly.
Sources: tests/test_vad.py