The OpenAI STT models is not producing real-time i...
# ask-ai
b
The OpenAI STT models is not producing real-time interim results, enabling the
use_realtime
option does not change this behavior either so what might be the issue?
t
Thanks for reaching out, Hemanth. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Thank you for reaching out! I wanted to share that the OpenAI Realtime API doesn’t provide real-time interim results, even if you enable the
use_realtime
option. This is just how the API works—it only shares the transcription once the user has finished speaking. If having real-time interim transcriptions is really important for your project, you might want to explore other options like the Speechmatics or Gladia STT plugins, or consider an STT-LLM-TTS pipeline. I’m here to help if you want to discuss these alternatives or need any further guidance! Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/real... | LiveKit turn detector plugin | LiveKit Docs | Live Conversations with AI using ChatGPT and WebRTC | Realtime model integrations | LiveKit Docs
b
i tried speechmatics but the english version is kind of bad while picking up numbers , when i say double two double five, it picks up either that way itself or w2 w5 instead of 2255. Is there anyway i can fix that?
t
I'm glad you're looking to improve Speechmatics' recognition of numbers like "double two double five" as "2255." Here are some helpful configuration options you can try: • Try enabling numeral formatting by setting the
max_delay_mode
parameter to
flexible
. This helps ensure the final transcript waits until the number formatting is properly done. • Consider using the
additional_vocab
parameter to add specific number sequences or phrases along with their phonetic pronunciations. For example:
Copy code
transcription_config=TranscriptionConfig(
    ...
    max_delay_mode="flexible",
    additional_vocab=[
        {
            "content": "2255",
            "sounds_like": ["double two double five"]
        }
    ]
)
It’s also a good idea to check the
end_of_utterance_silence_trigger
and
max_delay
parameters, so there’s enough time to capture the entire utterance before the transcript is finalized. You can adjust these based on the characteristics of your audio input. If you still encounter challenges after these tweaks, it might be due to some limitations in how the model handles these patterns, or it could mean further fine-tuning of timing parameters is needed. Please feel free to reach out if you want to explore this together—I’m here to help!
b
2025-09-25 175937,891 [INFO] main: Initialized SME onboarding session: identity-eTxu {"message": "unhandled exception while running the job task\nTraceback (most recent call last):\n File \"/home/hem/anaconda3/lib/python3.12/site-packages/opentelemetry/util/_decorator.py\", line 71, in async_wrapper\n return await func(*args, **kwargs) # type: ignore\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/home/hem/anaconda3/lib/python3.12/site-packages/livekit/agents/ipc/job_proc_lazy_main.py\", line 240, in _traceable_entrypoint\n await self._job_entrypoint_fnc(job_ctx)\n File \"/home/hem/projects_new/Livekit-final/backend_vin/backend_vin/ASB_SME_Onboarding_11-Aug-2025/AI-SME-ONBOARDING-BACKEND/main.py\", line 907, in entrypoint\n stt = speechmatics.STT(\n ^^^^^^^^^^^^^^^^^\nTypeError: STT.__init__() got an unexpected keyword argument 'max_delay_mode'", "level": "ERROR", "name": "livekit.agents", "pid": 255948, "job_id": "AJ_BygnLKZsg8fG", "timestamp": "2025-09-25T122937.896164+00:00"}" @dataclasses.dataclass class STTOptions: operating_point: OperatingPoint = OperatingPoint.ENHANCED domain: str | None = None language: str = "en" output_locale: str | None = None enable_partials: bool = True enable_diarization: bool = False max_delay: float = 0.7 end_of_utterance_silence_trigger: float = 0.3 end_of_utterance_mode: EndOfUtteranceMode = EndOfUtteranceMode.FIXED additional_vocab: list[AdditionalVocabEntry] = dataclasses.field(default_factory=list) punctuation_overrides: dict = dataclasses.field(default_factory=dict) diarization_sensitivity: float = 0.5 speaker_active_format: str = "{text}" speaker_passive_format: str = "{text}" prefer_current_speaker: bool = False focus_speakers: list[str] = dataclasses.field(default_factory=list) ignore_speakers: list[str] = dataclasses.field(default_factory=list) focus_mode: DiarizationFocusMode = DiarizationFocusMode.RETAIN known_speakers: list[DiarizationKnownSpeaker] = dataclasses.field(default_factory=list) class STT(stt.STT): def __init__( self, *, api_key: NotGivenOr[str] = NOT_GIVEN, base_url: NotGivenOr[str] = NOT_GIVEN, operating_point: OperatingPoint = OperatingPoint.ENHANCED, domain: NotGivenOr[str] = NOT_GIVEN, language: NotGivenOr[str] = NOT_GIVEN, output_locale: NotGivenOr[str] = NOT_GIVEN, enable_partials: bool = True, enable_diarization: bool = False, max_delay: float = 1.0, end_of_utterance_silence_trigger: float = 0.5, end_of_utterance_mode: EndOfUtteranceMode = EndOfUtteranceMode.FIXED, additional_vocab: NotGivenOr[list[AdditionalVocabEntry]] = NOT_GIVEN, punctuation_overrides: NotGivenOr[dict] = NOT_GIVEN, diarization_sensitivity: float = 0.5, speaker_active_format: str = "{text}", speaker_passive_format: str = "{text}", prefer_current_speaker: bool = False, focus_speakers: NotGivenOr[list[str]] = NOT_GIVEN, ignore_speakers: NotGivenOr[list[str]] = NOT_GIVEN, focus_mode: DiarizationFocusMode = DiarizationFocusMode.RETAIN, known_speakers: NotGivenOr[list[DiarizationKnownSpeaker]] = NOT_GIVEN, sample_rate: int = 16000, chunk_size: int = 160, audio_encoding: AudioEncoding = AudioEncoding.PCM_S16LE, transcription_config: NotGivenOr[TranscriptionConfig] = NOT_GIVEN, # Deprecated audio_settings: NotGivenOr[AudioSettings] = NOT_GIVEN, # Deprecated http_session: NotGivenOr[aiohttp.ClientSession] = NOT_GIVEN, ):" these are the params in plugin
t
Thanks for reaching out! The error you're seeing happens because the
max_delay_mode
parameter isn’t supported in the current version of the Speechmatics STT plugin for LiveKit. But don’t worry, you can use some other parameters that work just as well to get similar results: • `max_delay`: This controls the maximum delay in seconds for transcription results. • `end_of_utterance_mode`: Try using
EndOfUtteranceMode.ADAPTIVE
for a behavior that’s like the "flexible" mode. • `end_of_utterance_silence_trigger`: This sets how long silence lasts before it triggers the end of an utterance. Here’s a little example configuration to help you out:
Copy code
stt = speechmatics.STT(
    max_delay=0.7,
    end_of_utterance_mode=speechmatics.EndOfUtteranceMode.ADAPTIVE,
    end_of_utterance_silence_trigger=0.5
)
So, if you remove the
max_delay_mode
parameter from your code and switch to these supported options, it should clear up the issue. If you need any more help with this, just let me know—I’m here for you! Sources: livekit-plugins/livekit-plugins-gladia/livekit/plugins/gladia/stt.py | Speechmatics STT integration guide | LiveKit Docs
b
SME onboarding session: identity-TMtL {"message": "unhandled exception while running the job task\nTraceback (most recent call last):\n File \"/home/hem/anaconda3/lib/python3.12/site-packages/opentelemetry/util/_decorator.py\", line 71, in async_wrapper\n return await func(*args, **kwargs) # type: ignore\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/home/hem/anaconda3/lib/python3.12/site-packages/livekit/agents/ipc/job_proc_lazy_main.py\", line 240, in _traceable_entrypoint\n await self._job_entrypoint_fnc(job_ctx)\n File \"/home/hem/projects_new/Livekit-final/backend_vin/backend_vin/ASB_SME_Onboarding_11-Aug-2025/AI-SME-ONBOARDING-BACKEND/main.py\", line 907, in entrypoint\n stt = speechmatics.STT(\n ^^^^^^^^^^^^^^^^^\n File \"/home/hem/anaconda3/lib/python3.12/site-packages/livekit/plugins/speechmatics/stt.py\", line 307, in __init__\n self._process_config()\n File \"/home/hem/anaconda3/lib/python3.12/site-packages/livekit/plugins/speechmatics/stt.py\", line 367, in _process_config\n \"content\": e.content,\n ^^^^^^^^^\nAttributeError: 'dict' object has no attribute 'content'", "level": "ERROR", "name": "livekit.agents", "pid": 259441, "job_id": "AJ_bQPuATtxSF7W", "timestamp": "2025-09-25T123345.884540+00:00"}
t
It looks like the error is happening because the
additional_vocab
parameter is being given as a list of dictionaries instead of a list of
AdditionalVocabEntry
objects. The Speechmatics STT plugin expects each item in
additional_vocab
to be an
AdditionalVocabEntry
instance, which includes the
content
and
sounds_like
attributes. To help you out, here’s the right way to define and use `additional_vocab`:
Copy code
from livekit.plugins.speechmatics.types import AdditionalVocabEntry

additional_vocab = [
    AdditionalVocabEntry(content="gnocchi", sounds_like=["nyohki", "nokey", "nochi"]),
    AdditionalVocabEntry(content="CEO", sounds_like=["C.E.O."])
]

stt = speechmatics.STT(additional_vocab=additional_vocab)
Make sure to replace any dictionary-style entries with
AdditionalVocabEntry
objects, and that should clear things up. If you’re still running into issues, it might be worth double-checking your code or configuration to see if a dictionary is sneaking in somewhere. I’m here to help if you need any more guidance! Sources: Speechmatics STT integration guide | LiveKit Docs | livekit-agents/livekit/agents/stt/stt.py | Speechmatics integration guide | LiveKit Docs
b
additional_vocab = [ AdditionalVocabEntry(content="2255", _sounds_like_=["double two double five"]), AdditionalVocabEntry(content="CEO", _sounds_like_=["C.E.O."]) ]" even after putting an example like that , its still displayed as double two double five
t
I’m glad you’re exploring the
additional_vocab
parameter—it’s a great tool to help improve recognition accuracy! While it doesn’t guarantee that every spoken phrase will be replaced exactly with the specified content, it does increase the chances of recognizing those terms correctly. To help you get the best results, here are a few tips you might find useful: • Double-check that your
additional_vocab
entries are properly formatted using
AdditionalVocabEntry
, just as you’ve done. • Make sure the audio input aligns with the language and domain settings in your STT configuration (for example, English audio paired with
language="en"
). • It’s a good idea to remove any deprecated parameters like
transcription_config
to prevent any potential conflicts. • Try testing with a simple setup—enable only
additional_vocab
without other advanced features such as diarization—to isolate the effect. Remember,
additional_vocab
helps guide recognition but might not always completely override the transcription. If you continue to see issues, you might want to test using a known audio sample or enable live transcription logs to better understand what’s happening. Sources: livekit-plugins/livekit-plugins-aws/livekit/plugins/aws/stt.py | Speechmatics integration guide | LiveKit Docs
b
why doesnt it detect it as numbers properly anyway? are there no other params to fix this?
t
Thank you for reaching out! While Speechmatics STT doesn’t have specific parameters solely for enhancing number recognition beyond the
additional_vocab
feature, there are a few things you can try to help improve numeric transcription: • Make sure your
additional_vocab
includes all relevant numeric sequences and their phonetic forms, just like you’ve already done—great job on that! • Consider tweaking the
max_delay
and
end_of_utterance_silence_trigger
settings to give the system a bit more time to process and finalize what it hears. • Also, double-check the audio quality—clearer audio can really make a difference in recognition accuracy. If you’re still seeing some inconsistencies with numbers, it might be a current limitation of the Speechmatics model. But please don’t hesitate to reach out if you want me to help explore other options or if you have any more questions—I’m here to help! Sources: Speechmatics integration guide | LiveKit Docs | livekit-plugins/livekit-plugins-baseten/livekit/plugins/baseten/stt.py