VADEventType.START_OF_SPEECH event is not fired ev...
# ask-ai
l
VADEventType.START_OF_SPEECH event is not fired even though user has started speaking. vad.VADEventType.INFERENCE_DONE is fired directly.
t
Thanks for reaching out, Deepak. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
The issue where
VADEventType.START_OF_SPEECH
is not firing while
vad.VADEventType.INFERENCE_DONE
fires directly indicates that the VAD is processing audio but not detecting speech above the configured thresholds. Here are steps to troubleshoot and resolve this: 1. Check Activation Threshold: Ensure the
activation_threshold
is not set too high. Lowering it can help detect speech more easily.
Copy code
// Example in JavaScript
const vad = await VAD.load({
    activationThreshold: 0.3,  // Lower from default 0.5
    minSpeechDuration: 50,     // Default is 50ms
});
2. Adjust Minimum Speech Duration: If speech segments are too short,
START_OF_SPEECH
won’t fire. Reduce the
minSpeechDuration
value.
Copy code
// Example in JavaScript
vad.updateOptions({
    activationThreshold: 0.3,
    minSpeechDuration: 20
});
3. Monitor Speech Probability: Check the probability values in
INFERENCE_DONE
events to understand what the VAD is detecting.
Copy code
// Example in JavaScript
for await (const event of vadStream) {
    if (event.type === VADEventType.INFERENCE_DONE) {
        console.log(`Speech probability: ${event.probability}`);
        console.log(`Speaking state: ${event.speaking}`);
    }
}
4. Check Audio Input Quality: Ensure the audio input has sufficient volume and quality. The sample rate should be 8kHz or 16kHz, and the audio should not be too quiet or heavily processed. 5. Debug with Test Audio: Use test patterns to verify the VAD is working as expected.
Copy code
// Example in Python
frames, _ = await utils.make_test_speech(sample_rate=16000)
stream = vad.stream()

for frame in frames:
    stream.push_frame(frame)

stream.end_input()

async for ev in stream:
    print(f"Event: {ev.type}, Probability: {ev.probability}")
By following these steps, you should be able to identify and resolve the issue. Let me know if you need further assistance! Sources: livekit-agents/livekit/agents/stt/stream_adapter.py
l
VAD is processing audio but not detecting speech above the configured thresholds.
Then how come INFERENCE_DONE is fired
t
Hey there! Just wanted to explain a bit about the
INFERENCE_DONE
event. It actually fires for every inference window that the VAD model processes, whether or not it detects speech above the activation threshold. This is totally on purpose and helps in a few important ways: • Regular Heartbeat: The
INFERENCE_DONE
event gives you steady updates, so you know the VAD is actively working on the audio. • Debugging and Monitoring: It lets you peek at the raw probability values, even when they’re below that speech threshold. • Metrics Collection: The system uses these events to gather useful VAD stats, like how long each inference takes and how many have happened. • Fine-Grained Control: You can build your own custom logic based on those probability values, even if they’re below the threshold. Here’s how the event firing sequence works: 1. The VAD model processes a chunk of audio and calculates a speech probability. 2. The
INFERENCE_DONE
event fires right after, carrying the probability and some other info. 3. Only after that does the system check if the probability passes the
activation_threshold
to decide whether to fire
START_OF_SPEECH
or
END_OF_SPEECH
events. So, even if the probability is below the threshold, the
INFERENCE_DONE
event still fires to keep you in the loop about what the VAD is doing. In those cases, the
speaking
field will be
false
, but the event’s there to make sure you get consistent updates and can debug smoothly. Sources: tests/test_vad.py