billions-tomato-8365
09/25/2025, 12:18 PMuse_realtime option does not change this behavior either so what might be the issue?tall-airline-16756
09/25/2025, 12:18 PMtall-airline-16756
09/25/2025, 12:19 PMuse_realtime option. This is just how the API works—it only shares the transcription once the user has finished speaking.
If having real-time interim transcriptions is really important for your project, you might want to explore other options like the Speechmatics or Gladia STT plugins, or consider an STT-LLM-TTS pipeline. I’m here to help if you want to discuss these alternatives or need any further guidance!
Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/real... | LiveKit turn detector plugin | LiveKit Docs | Live Conversations with AI using ChatGPT and WebRTC | Realtime model integrations | LiveKit Docsbillions-tomato-8365
09/25/2025, 12:23 PMtall-airline-16756
09/25/2025, 12:23 PMmax_delay_mode parameter to flexible. This helps ensure the final transcript waits until the number formatting is properly done.
• Consider using the additional_vocab parameter to add specific number sequences or phrases along with their phonetic pronunciations. For example:
transcription_config=TranscriptionConfig(
...
max_delay_mode="flexible",
additional_vocab=[
{
"content": "2255",
"sounds_like": ["double two double five"]
}
]
)
It’s also a good idea to check the end_of_utterance_silence_trigger and max_delay parameters, so there’s enough time to capture the entire utterance before the transcript is finalized. You can adjust these based on the characteristics of your audio input.
If you still encounter challenges after these tweaks, it might be due to some limitations in how the model handles these patterns, or it could mean further fine-tuning of timing parameters is needed. Please feel free to reach out if you want to explore this together—I’m here to help!billions-tomato-8365
09/25/2025, 12:31 PMtall-airline-16756
09/25/2025, 12:32 PMmax_delay_mode parameter isn’t supported in the current version of the Speechmatics STT plugin for LiveKit. But don’t worry, you can use some other parameters that work just as well to get similar results:
• `max_delay`: This controls the maximum delay in seconds for transcription results.
• `end_of_utterance_mode`: Try using EndOfUtteranceMode.ADAPTIVE for a behavior that’s like the "flexible" mode.
• `end_of_utterance_silence_trigger`: This sets how long silence lasts before it triggers the end of an utterance.
Here’s a little example configuration to help you out:
stt = speechmatics.STT(
max_delay=0.7,
end_of_utterance_mode=speechmatics.EndOfUtteranceMode.ADAPTIVE,
end_of_utterance_silence_trigger=0.5
)
So, if you remove the max_delay_mode parameter from your code and switch to these supported options, it should clear up the issue. If you need any more help with this, just let me know—I’m here for you!
Sources: livekit-plugins/livekit-plugins-gladia/livekit/plugins/gladia/stt.py | Speechmatics STT integration guide | LiveKit Docsbillions-tomato-8365
09/25/2025, 12:34 PMtall-airline-16756
09/25/2025, 12:35 PMadditional_vocab parameter is being given as a list of dictionaries instead of a list of AdditionalVocabEntry objects. The Speechmatics STT plugin expects each item in additional_vocab to be an AdditionalVocabEntry instance, which includes the content and sounds_like attributes.
To help you out, here’s the right way to define and use `additional_vocab`:
from livekit.plugins.speechmatics.types import AdditionalVocabEntry
additional_vocab = [
AdditionalVocabEntry(content="gnocchi", sounds_like=["nyohki", "nokey", "nochi"]),
AdditionalVocabEntry(content="CEO", sounds_like=["C.E.O."])
]
stt = speechmatics.STT(additional_vocab=additional_vocab)
Make sure to replace any dictionary-style entries with AdditionalVocabEntry objects, and that should clear things up. If you’re still running into issues, it might be worth double-checking your code or configuration to see if a dictionary is sneaking in somewhere. I’m here to help if you need any more guidance!
Sources: Speechmatics STT integration guide | LiveKit Docs | livekit-agents/livekit/agents/stt/stt.py | Speechmatics integration guide | LiveKit Docsbillions-tomato-8365
09/25/2025, 12:39 PMtall-airline-16756
09/25/2025, 12:39 PMadditional_vocab parameter—it’s a great tool to help improve recognition accuracy! While it doesn’t guarantee that every spoken phrase will be replaced exactly with the specified content, it does increase the chances of recognizing those terms correctly.
To help you get the best results, here are a few tips you might find useful:
• Double-check that your additional_vocab entries are properly formatted using AdditionalVocabEntry, just as you’ve done.
• Make sure the audio input aligns with the language and domain settings in your STT configuration (for example, English audio paired with language="en").
• It’s a good idea to remove any deprecated parameters like transcription_config to prevent any potential conflicts.
• Try testing with a simple setup—enable only additional_vocab without other advanced features such as diarization—to isolate the effect.
Remember, additional_vocab helps guide recognition but might not always completely override the transcription. If you continue to see issues, you might want to test using a known audio sample or enable live transcription logs to better understand what’s happening.
Sources: livekit-plugins/livekit-plugins-aws/livekit/plugins/aws/stt.py | Speechmatics integration guide | LiveKit Docsbillions-tomato-8365
09/25/2025, 12:42 PMtall-airline-16756
09/25/2025, 12:43 PMadditional_vocab feature, there are a few things you can try to help improve numeric transcription:
• Make sure your additional_vocab includes all relevant numeric sequences and their phonetic forms, just like you’ve already done—great job on that!
• Consider tweaking the max_delay and end_of_utterance_silence_trigger settings to give the system a bit more time to process and finalize what it hears.
• Also, double-check the audio quality—clearer audio can really make a difference in recognition accuracy.
If you’re still seeing some inconsistencies with numbers, it might be a current limitation of the Speechmatics model. But please don’t hesitate to reach out if you want me to help explore other options or if you have any more questions—I’m here to help!
Sources: Speechmatics integration guide | LiveKit Docs | livekit-plugins/livekit-plugins-baseten/livekit/plugins/baseten/stt.py