does livekit support elevenlabs stt?
# ask-ai
c
does livekit support elevenlabs stt?
t
Thanks for reaching out, Faraz. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Yes, LiveKit supports integration with ElevenLabs for Speech-to-Text (STT) functionality. Here are the key details: • The ElevenLabs plugin includes an STT class that uses the "scribe_v1" model for speech-to-text processing. • It utilizes the ElevenLabs API endpoint
/speech-to-text
. • Language configuration is supported, with English ("en") as the default. • An ElevenLabs API key is required, which should be set via the
ELEVEN_API_KEY
environment variable. • The STT functionality is non-streaming but supports interim results. While the STT functionality is implemented, the current documentation primarily emphasizes Text-to-Speech (TTS) integration with ElevenLabs. Sources: livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/__ini... | ElevenLabs TTS integration guide | LiveKit Docs
b
How do i use this in my agent?
t
To use ElevenLabs with a LiveKit agent, you can integrate its Text-to-Speech (TTS) functionality, as ElevenLabs Speech-to-Text (STT) is not officially supported or documented in the LiveKit agents framework. Here’s a friendly guide to help you set up ElevenLabs TTS: 1. Install the Plugin
Copy code
pip install "livekit-agents[elevenlabs]~=1.2"
2. Set Authentication Just pop your ElevenLabs API key into the
.env
file like this:
Copy code
ELEVEN_API_KEY=your_api_key_here
3. Use in an AgentSession You can integrate ElevenLabs TTS within an
AgentSession
for real-time speech synthesis. Here’s a simple example:
Copy code
from livekit.plugins import elevenlabs, openai, deepgram, silero
from livekit.agents.voice import AgentSession, Agent

class MyAgent(Agent):
    def __init__(self):
        super().__init__(
            stt=deepgram.STT(),                         # or another supported STT
            llm=openai.LLM(),                          # generates responses
            tts=elevenlabs.TTS(model="eleven_multilingual_v2", voice_id="ODq5zmih8GrVes37Dizd"),  # ElevenLabs TTS
            vad=silero.VAD.load()                      # voice activity detection
        )

async def entrypoint(ctx):
    session = AgentSession()
    await session.start(agent=MyAgent(), room=ctx.room)
4. Real-Time TTS Streaming ElevenLabs TTS supports real-time streaming, which is great for low-latency synthesis. You can use the
stream()
method like this:
Copy code
tts = elevenlabs.TTS()
stream = tts.stream()

# Send text in chunks
for chunk in ["Hello, ", "how are you?", ""]:
    stream.push_text(chunk)  # pushes text incrementally
stream.end_input()

# Consume audio frames
async for audio_event in stream:
    await audio_source.capture_frame(audio_event.frame)
For the full details, you can always check out the example usage. It’s a handy resource! Limitations and Notes • STT Not Available: ElevenLabs STT isn’t officially supported or documented yet. If you need speech recognition, I recommend using Deepgram, OpenAI, or Azure — they work really well. • TTS Streaming Supported: The
tts.stream()
method uses WebSocket connections to provide smooth real-time audio synthesis. • Aligned Transcripts: You can enable
sync_alignment=True
(which is the default) to get synchronized transcription events — this can really enhance the user experience. If you’re interested in adding Speech-to-Text functionality, I’m here to help you get started with one of the officially supported providers like Deepgram or OpenAI. Just let me know, and we can walk through it together! Sources: Speech-to-text (STT) integrations | LiveKit Docs | translators/tts_translator.py | ElevenLabs TTS integration guide | LiveKit Docs | livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/__ini... | ElevenLabs integration guide | LiveKit Docs