Join Slack
Powered by
How to approach the Speech-to-Text using Gemini 2....
# ask-ai
a
adorable-petabyte-73906
09/25/2025, 9:15 AM
How to approach the Speech-to-Text using Gemini 2.5 Flash or gemini-flash-2.0-lite and implement it in the AgentSession?
t
tall-airline-16756
09/25/2025, 9:15 AM
Thanks for reaching out, Aditya. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
tall-airline-16756
09/25/2025, 9:16 AM
Thank you for reaching out! I wanted to share that LiveKit doesn’t currently support Gemini 2.5 Flash or gemini-flash-2.0-lite for Speech-to-Text (STT). Gemini is fantastic for Text-to-Speech (TTS) and language modeling (LLM), but when it comes to speech recognition, we recommend using providers like Deepgram, AssemblyAI, or Google Cloud STT. If you’d like, I’m here to help you with: • Getting started with Gemini TTS for speech generation? • Setting up Google Cloud STT or another provider for speech recognition? • Building a complete voice agent pipeline with STT + LLM + TTS? Sources:
Gemini TTS integration guide | LiveKit Docs
|
livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/gemini_t...
|
Gemini Live API integration guide | LiveKit Docs
|
Google Gemini integration guide | LiveKit Docs
|
plugins/google/src/beta/gemini_tts.ts
Open in Slack
Previous
Next