I want to use Gemini 2.5 Flash as my STT, how can ...
# ask-ai
a
I want to use Gemini 2.5 Flash as my STT, how can i setup that in my AgentSession?
t
Thanks for reaching out, Aditya. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Hi there! I wanted to let you know that Gemini 2.5 Flash isn’t available for Speech-to-Text (STT) in AgentSession, but it works great for Text-to-Speech (TTS) or as a Large Language Model (LLM). Could you please let me know if you’re aiming to use it for speech recognition (STT), or did you mean text-to-speech (TTS)? If STT is what you need, I’d be happy to help you get started with the supported Google Cloud STT models. Just let me know—I’m here to help! Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/models.py | Gemini TTS integration guide | LiveKit Docs | Gemini TTS integration guide | LiveKit Docs | Google Gemini integration guide | LiveKit Docs | Gemini Live API integration guide | LiveKit Docs
a
I want to use gemini 2.0 flash lite for Speech-to-Text and implement it in the AgentSession
t
Hi there! I wanted to share a quick note that Gemini 2.0 Flash Lite isn’t designed to be used as a Speech-to-Text (STT) service in AgentSession. It’s really great for Large Language Model (LLM) tasks or real-time language processing, but speech recognition isn’t its main focus. If you’re looking to build a voice AI pipeline, a good approach is to use a supported STT engine like Deepgram, Silero, or Google Cloud STT, and then pair that with Gemini 2.0 Flash Lite as your LLM. Here’s a simple example to illustrate:
Copy code
from livekit.plugins import google, deepgram

session = AgentSession(
    stt=deepgram.STT(),
    llm=google.LLM(model="gemini-2.0-flash"),
    tts=google.beta.GeminiTTS(model="gemini-2.0-flash-tts")
)
If you'd like, I’d be happy to walk you through setting up a full voice AI pipeline with a compatible STT engine. Just let me know! Sources: Google AI and LiveKit | LiveKit Docs | Gemini Live API integration guide | LiveKit Docs | pipeline-llm/large_context.py | Gemini TTS integration guide | LiveKit Docs | Google Gemini integration guide | LiveKit Docs
a
@proud-battery-54912
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!
a
But if i want to try it out , how to approach the Speech-to-Text using Gemini 2.5 Flash or gemini-flash-2.0-lite
But if i want to try it out , how to approach the Speech-to-Text using Gemini 2.5 Flash or gemini-flash-2.0-lite (edited)
d
hey @adorable-petabyte-73906, did you try to implement this?
p
I don't think it's possible, if you want you can do LLM -> TTS pipeline, and have gemini flash as llm with output modality as TEXT. But had issues with tool calling
d
currently i am using STT->LLM->TTS pipeline with google, but i am facing issues with short text such as "hi", "yes", "no". Am not getting the text result for short audio. Is there any other option to improve it rather than trying to use gemini flash?
stt=google.STT(model="telephony", spoken_punctuation=False, languages=["en-IN", "hi-IN", "en-US"])
this is my stt config
a
the performance of all the models out there is bad on short phrases.