while building an end to end streaming voice agent...
# ask-ai
a
while building an end to end streaming voice agent with livekit, can i swapout LLM in the livekit starter project and pass the STT output to a different sytem which acts as the agent/brain and get text output from there and do TTS and send it to user ... are there any references to this and how to make it work with less latencies ?
t
Thanks for reaching out, pavan. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I'd love to help you out! To get a clearer picture and offer the best advice, could you share a bit more about your setup and what you're aiming for? • Are you working with a specific LiveKit starter project, or have you built a custom implementation? • Which system or platform will you be using as the "agent/brain" to process the STT output? • Do you have any particular latency goals or benchmarks you'd like to hit? Once I have these details, I’ll be able to guide you more effectively towards your goals.
a
i am using one of the example projects and want to swaoput the LLM , i am thinking to build a different agent system o langgraph , so specific goals, want it to be realtime
t
I'm glad you're interested in customizing a LiveKit example project to replace the LLM with your own agent system like LangGraph while keeping everything running smoothly in real-time for STT and TTS. Let me walk you through the steps to help you get started: 1. Install Required Dependencies: First, you'll want to install the LiveKit LangChain plugin with pip:
pip install "livekit-plugins-langchain~=1.1"
2. Define Your LangGraph Workflow: Next, create a LangGraph workflow by defining a
StateGraph
that will handle the STT output and generate the responses you need. 3. Wrap the Workflow: To make sure your LangGraph workflow fits nicely into LiveKit’s setup, use
langchain.LLMAdapter
to wrap it. 4. Set Up STT, TTS, and VAD: Then, configure low-latency STT and TTS services (like Deepgram) along with voice activity detection (such as Silero) to keep your voice processing real-time and responsive. 5. Start the Agent Session: Once everything is ready, initialize and start the agent session, connecting it to your custom LangGraph-based agent. 6. Optimize for Real-Time Performance: Finally, make sure to use streaming-capable STT/TTS providers and enable turn detection to keep latency as low as possible. If you’d like to dive deeper or see some code examples, I highly recommend checking out the LangChain integration documentation and the Voice AI quickstart guide. They’re great resources to support you along the way! Sources: LangChain integration guide | LiveKit Docs | Pipeline nodes and hooks | LiveKit Docs | Docs::Agents | examples/src/cartersia_tts.ts | livekit-plugins/livekit-plugins-langchain/livekit/plugins/langchain/langgra...
a
what about mastra instead of langgraph ?
t
I'd love to understand a bit more about how you're planning to use Mastra in your project. Could you please clarify: • Are you thinking of using Mastra as a workflow framework to replace LangGraph? • Or is Mastra meant to offer STT/TTS capabilities? Knowing this will really help me guide you on the best way to integrate it with LiveKit. I'm here to help every step of the way! Sources: livekit-plugins/livekit-plugins-sarvam/livekit/plugins/sarvam/__init__.py | Mistral STT integration guide | LiveKit Docs | Large language model (LLM) integrations | LiveKit Docs
a
mastra is to replace langgraph , it will be the main agent backend, it won't support STT or TTS, they will be part of our livekit project only
t
I'd love to help you integrate Mastra as the main agent backend in your LiveKit project while keeping STT and TTS within LiveKit. Here’s a friendly guide to get you started: 1. Customize the LLM Node: You can override the
llm_node()
method in a custom
Agent
class. This way, transcription output is routed to Mastra, and its response is returned for TTS processing. 2. Configure STT and TTS: For a smooth experience, use LiveKit plugins for STT (like Deepgram) and TTS (such as ElevenLabs) within your
AgentSession
. They handle transcription and synthesis effortlessly. 3. Pipeline Hooks: Hooks like
on_user_turn_completed()
are great for sending user input to Mastra and then feeding Mastra’s response back into LiveKit’s TTS system. 4. Deploy the Agent: Finally, run your custom agent worker using
cli.run_app()
, connecting it to your LiveKit server. Mastra can be set up as a separate service or API. This setup is designed to give you real-time performance by combining LiveKit’s low-latency STT and TTS with Mastra’s agent logic. If you'd like, I’m here to provide detailed code examples or walk you through any part of this process—just let me know! Sources: Plugin System | livekit/livekit_composite | DeepWiki | Docs::Agents | README.md
a
i am asking about Mastra, and not Mistral, can you recheck and share updated response
t
To integrate Mastra as the main agent backend in your LiveKit project while keeping STT and TTS within LiveKit, here’s a simple approach you can follow: 1. Keep STT and TTS in LiveKit: You can continue using LiveKit plugins like Deepgram for STT and ElevenLabs for TTS to manage transcription and synthesis smoothly. 2. Replace the LLM Node: Try implementing a custom
llm_node()
in your
Agent
class that routes the transcription output to Mastra through its API and then returns the response for TTS processing. 3. Configure the AgentSession: Set up the
AgentSession
with LiveKit’s STT and TTS plugins, while keeping the LLM logic within your custom implementation. 4. Optional Plugin: If you’d like, you can create a custom plugin for Mastra that fits the LLM interface LiveKit expects, making the integration even smoother. Since Mastra isn’t supported out of the box, you’ll need to handle the API calls manually. Please don’t hesitate to reach out if you’d like me to provide detailed code examples or any further help—I’m here to support you! Sources: Working with plugins | LiveKit Docs | LiveKit integration guides | LiveKit Docs | README.md | Pipeline nodes and hooks | LiveKit Docs
a
give me detailed info or reference links or examples to implement this
t
I'd love to help you integrate Mastra as the main agent backend in your LiveKit project while keeping STT and TTS within LiveKit. Here’s a friendly, step-by-step approach you can follow: 1. Initialize the Agent: You can continue using LiveKit plugins for STT (like Deepgram) and TTS (such as ElevenLabs) to handle transcription and synthesis—these parts stay just as they are. 2. Override the LLM Node: Try implementing a custom
llm_node()
method in your
Agent
class. This way, you can route the transcription output to Mastra through its API and stream the response back for TTS processing. Here’s a simple example:
Copy code
class CustomLLMAgent(Agent):
    async def llm_node(self, chat_ctx, tools, model_settings=None):
        async def process_stream():
            # Send chat_ctx.messages to Mastra backend
            response_stream = await send_to_mastra_backend(chat_ctx.messages)
            async for chunk in response_stream:
                yield chunk  # Stream response chunks to TTS
        return process_stream()
1. Customize Pipeline Nodes: You can use hooks like
on_user_turn_completed()
to catch user input, send it over to Mastra, and then feed the response into LiveKit’s TTS pipeline. It’s a neat way to keep everything connected smoothly. 2. Deploy the Agent: Finally, run your custom agent worker using
cli.run_app()
, connecting it to your LiveKit server. Mastra can operate as a separate service or API alongside this. To help you get started, here are some useful resources with detailed examples and references: • Customizing Pipeline Nodes • Example: Replacing LLM Output • LLM Integrations Overview • LiveKit Agents README If you happen to have Mastra’s API documentation handy, I’d be more than happy to help you design the integration logic in even more detail. Just let me know whenever you're ready! Sources: Large language model (LLM) integrations | LiveKit Docs | README.md | Pipeline nodes and hooks | LiveKit Docs