Can some STT plugins (I'm thinking Azure STT or De...
# ask-ai
t
Can some STT plugins (I'm thinking Azure STT or Deepgram) help me collect the speaker's rate of speech ? I'd like to adapt my bot's voice speed to the caller's
t
Thanks for reaching out, Pierrick. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Neither Azure STT nor Deepgram plugins directly provide the speaker's rate of speech (e.g., words per minute). However, you can estimate it by combining metadata from recognized speech segments, such as timestamps for speech start and end events, and the word count from the transcription text. Deepgram offers additional accuracy by providing word-level timing metadata, allowing you to calculate the actual vocal duration more precisely. If you'd like, I’d be happy to guide you through this or even share a code example to help you implement this estimation in your LiveKit integration. Just let me know! Sources: livekit-plugins/livekit-plugins-speechmatics/livekit/plugins/speechmatics/s... | Azure integration guide | LiveKit Docs | plugins/deepgram/src/stt.ts
t
Yep, I'd like to see some code please
t
Here’s a friendly example to help you estimate the speaker's rate of speech using Deepgram STT in Python. This method uses word-level timing metadata to give you a precise calculation:
Copy code
import time

def calculate_speech_rate(transcription):
    words = transcription['words']
    if not words:
        return 0

    # Extract start and end times of the first and last words
    start_time = words[0]['start']
    end_time = words[-1]['end']

    # Calculate duration in seconds
    duration_sec = end_time - start_time

    # Count the number of words
    word_count = len(words)

    # Calculate words per minute (WPM)
    wpm = (word_count / duration_sec) * 60
    return wpm

# Example transcription data from Deepgram
transcription = {
    'words': [
        {'start': 0.5, 'end': 1.0, 'word': 'Hello'},
        {'start': 1.1, 'end': 1.5, 'word': 'world'},
        {'start': 1.6, 'end': 2.0, 'word': 'example'}
    ]
}

# Calculate and print the speech rate
speech_rate = calculate_speech_rate(transcription)
print(f"Estimated speech rate: {speech_rate:.2f} words per minute")
This example works if you’re getting transcription data with word-level timing from Deepgram. If you happen to be using Azure STT, the idea is quite similar, but you’ll use start and end timestamps for the whole speech segment. Please feel free to reach out if you’d like me to help with a customized example for Azure STT or if you have any other questions! I’m here to help.
t
How's called the stt plugin event I should hook on?
t
For Deepgram STT, I recommend connecting to the following events: • SpeechStarted: This marks when speech detection begins. • Results: You'll want to use this event with
is_final=True
and
speech_final=True
to catch the end of a spoken segment and get the final transcription. These events will give you the key information needed to calculate the speech rate. If you have any questions or need a hand with the implementation, please feel free to reach out—I’m here to help!