crooked-tiger-31253
07/25/2025, 4:58 AMtall-airline-16756
07/25/2025, 4:58 AMtall-airline-16756
07/25/2025, 4:58 AM/speech-to-text.
• Language configuration is supported, with English ("en") as the default.
• An ElevenLabs API key is required, which should be set via the ELEVEN_API_KEY environment variable.
• The STT functionality is non-streaming but supports interim results.
While the STT functionality is implemented, the current documentation primarily emphasizes Text-to-Speech (TTS) integration with ElevenLabs.
Sources: livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/__ini... | ElevenLabs TTS integration guide | LiveKit Docsbumpy-student-61140
09/09/2025, 11:09 AMtall-airline-16756
09/09/2025, 11:11 AMpip install "livekit-agents[elevenlabs]~=1.2"
2. Set Authentication
Just pop your ElevenLabs API key into the .env file like this:
ELEVEN_API_KEY=your_api_key_here
3. Use in an AgentSession
You can integrate ElevenLabs TTS within an AgentSession for real-time speech synthesis. Here’s a simple example:
from livekit.plugins import elevenlabs, openai, deepgram, silero
from livekit.agents.voice import AgentSession, Agent
class MyAgent(Agent):
def __init__(self):
super().__init__(
stt=deepgram.STT(), # or another supported STT
llm=openai.LLM(), # generates responses
tts=elevenlabs.TTS(model="eleven_multilingual_v2", voice_id="ODq5zmih8GrVes37Dizd"), # ElevenLabs TTS
vad=silero.VAD.load() # voice activity detection
)
async def entrypoint(ctx):
session = AgentSession()
await session.start(agent=MyAgent(), room=ctx.room)
4. Real-Time TTS Streaming
ElevenLabs TTS supports real-time streaming, which is great for low-latency synthesis. You can use the stream() method like this:
tts = elevenlabs.TTS()
stream = tts.stream()
# Send text in chunks
for chunk in ["Hello, ", "how are you?", ""]:
stream.push_text(chunk) # pushes text incrementally
stream.end_input()
# Consume audio frames
async for audio_event in stream:
await audio_source.capture_frame(audio_event.frame)
For the full details, you can always check out the example usage. It’s a handy resource!
Limitations and Notes
• STT Not Available: ElevenLabs STT isn’t officially supported or documented yet. If you need speech recognition, I recommend using Deepgram, OpenAI, or Azure — they work really well.
• TTS Streaming Supported: The tts.stream() method uses WebSocket connections to provide smooth real-time audio synthesis.
• Aligned Transcripts: You can enable sync_alignment=True (which is the default) to get synchronized transcription events — this can really enhance the user experience.
If you’re interested in adding Speech-to-Text functionality, I’m here to help you get started with one of the officially supported providers like Deepgram or OpenAI. Just let me know, and we can walk through it together!
Sources: Speech-to-text (STT) integrations | LiveKit Docs | translators/tts_translator.py | ElevenLabs TTS integration guide | LiveKit Docs | livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/__ini... | ElevenLabs integration guide | LiveKit Docs