can i use RAG with realtime models? if yes is ther...
# ask-ai
b
can i use RAG with realtime models? if yes is there any example? also can i use realtime models for TTS and SST only?
t
Thanks for reaching out, Sudo. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I'm happy to share that you can definitely use Retrieval-Augmented Generation (RAG) with real-time models in LiveKit. There are some great examples available that show how to seamlessly integrate RAG with real-time workflows, including the use of vector databases and low-latency retrieval techniques. You might find the LiveKit Docs RAG and LlamaIndex RAG integrations especially helpful—they’re open-source and come with clear setup instructions. If you'd like, I’d be more than happy to guide you through adapting these examples to fit your specific needs. As for your second question, LiveKit’s real-time models go beyond just Text-to-Speech (TTS) and Speech-to-Speech Translation (SST). They support direct speech-to-speech generation and can also work alongside separate TTS or STT pipelines if that suits your project better. Plus, they offer some pretty neat features like tool calling, image understanding, and turn detection based on verbal cues. If you want to explore any of these capabilities further, just let me know—I’m here to help! Sources: LiveKit integration guides | LiveKit Docs | External data and RAG | LiveKit Docs | External data and RAG | LiveKit Docs | rag/README.md | delay-handling/agent.py
b
can you give me a basic example for both, lets use ultravox's realtime model
t
Let me walk you through a simple example of how to use Retrieval-Augmented Generation (RAG) with Ultravox's real-time model: 1. First, set up the Ultravox real-time model using the
RealtimeModel
class. This model is great because it supports dynamic context injection through the
update_chat_ctx
method, letting you add retrieved information as deferred instructions. 2. Next, create a RAG retriever that taps into a vector database (like Annoy) to find relevant info based on user queries. The retrieved context is then added into the conversation as a system message wrapped in
<instruction>
tags. 3. Finally, integrate this retriever into your agent as a function tool, so it can dynamically fetch and inject context during real-time chats. If you'd like, I’m happy to provide the full Python code for this setup or help you get started with components like the vector database. Just let me know how I can assist! Sources: LiveKit Docs | delay-handling/agent.py | examples/README.md | rag/README.md
b
can you give me a python example, and also given python code for the second question
t
Here’s a Python example for using RAG with Ultravox's real-time model:
Copy code
import asyncio
from livekit.agents import Agent, AgentSession, JobContext, WorkerOptions, cli, llm
from livekit.plugins.ultravox.realtime import RealtimeModel
from livekit.plugins import openai as openai_plugin
import annoy
import pickle
from pathlib import Path

class RAGRetriever:
    def __init__(self, index_path: str, data_path: str):
        self._index_path = Path(index_path)
        self._data_path = Path(data_path)
        self._annoy_index = annoy.AnnoyIndex(1536, 'angular')
        self._annoy_index.load(str(self._index_path))
        with open(self._data_path, 'rb') as f:
            self._paragraphs_by_uuid = pickle.load(f)

    async def search(self, query: str) -> str:
        query_embedding = await openai_plugin.create_embeddings(
            input=[query],
            model="text-embedding-3-small",
            dimensions=1536,
        )
        results = self._annoy_index.get_nns_by_vector(query_embedding[0].embedding, n=1)
        if results:
            paragraph_uuid = results[0]
            return self._paragraphs_by_uuid.get(paragraph_uuid, "No relevant information found.")
        return "No relevant information found."

class VoiceAssistantWithRAG(Agent):
    def __init__(self):
        super().__init__(
            instructions="You are a helpful assistant.",
            llm=RealtimeModel(model_id="fixie-ai/ultravox", voice="Mark"),
        )
        self._retriever = RAGRetriever("index.annoy", "data.pkl")
        self.add_tool(self._rag_search)

    @llm.function_tool
    async def _rag_search(self, query: str):
        context = await self._retriever.search(query)
        chat_ctx = llm.ChatContext()
        chat_ctx.add_message(role="system", content=f"<instruction>{context}</instruction>")
        await self.session.update_chat_ctx(chat_ctx)
        return f"Retrieved context for query: {query}"

async def entrypoint(ctx: JobContext):
    session = AgentSession()
    await session.start(agent=VoiceAssistantWithRAG(), room=ctx.room)

if __name__ == "__main__":
    cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
For TTS and SST, here’s an example using OpenAI's real-time model:
Copy code
from livekit.agents.voice import AgentSession
from livekit.plugins import openai

class WeatherAgent:
    def __init__(self):
        self.llm = openai.realtime.RealtimeModel(modalities=["text"])
        self.tts = openai.TTS(voice="ash")

async def entrypoint(ctx):
    session = AgentSession()
    agent = WeatherAgent()
    await session.start(agent=agent, room=ctx.room)
If you have any questions or need a hand walking through these examples, please don’t hesitate to ask—I’m here to help! Sources: Realtime model integrations | LiveKit Docs | OpenAI TTS integration guide | LiveKit Docs | examples/other/text-to-speech/elevenlabs_tts.py | Groq TTS integration guide | LiveKit Docs | LMNT TTS integration guide | LiveKit Docs
b
can i use ultravox for speech to text m=node only? if yes can i get a python code?
t
Absolutely, you can use Ultravox just for Speech-to-Text (STT). Here’s a friendly Python example to help you get started with setting up and using Ultravox’s real-time model for STT:
Copy code
from livekit import agents, rtc
from livekit.plugins import ultravox

async def run(ctx: agents.JobContext):
    # Wait for the first participant
    await ctx.wait_for_participant()

    # Initialize the Ultravox plugin
    ultravox_model = ultravox.RealtimeModel(api_key="your_api_key")

    # Create a new realtime session
    stt_session = ultravox_model.session()

    # Process audio from a participant's microphone track
    for _, participant in ctx.room.participants.items():
        audio_stream = rtc.AudioStream(participant.tracks[1].pub.track)
        async for audio_event in audio_stream:
            stt_session.push_audio(audio_event.frame)
        break  # Process only the first participant for this example

    # Gracefully close the session when done
    await stt_session.aclose()

async def handle_stt_events(session: ultravox.RealtimeSession):
    # Placeholder for handling transcription events
    pass
This example walks you through initializing the Ultravox real-time model, capturing audio from a participant’s microphone, and sending it to the STT session for transcription. If you have any questions or need a hand with anything else, just let me know—I’m here to help! Sources: livekit-plugins/livekit-plugins-baseten/livekit/plugins/baseten/stt.py
b
Is there a way to send images to livekit + letta?
t
Absolutely, you can send images to LiveKit using the Letta integration! To include images in the agent's chat context, you can use
ImageContent
, either by providing base64-encoded data or by linking to external URLs. Here’s a little example to help you get started:
Copy code
import base64
from livekit.agents import ChatContext, ImageContent

# Example: Adding an image to the chat context
image_bytes = open("example.png", "rb").read()
chat_ctx = ChatContext()
chat_ctx.add_message(
    role="user",
    content=[
        ImageContent(
            image=f"data:image/png;base64,{base64.b64encode(image_bytes).decode('utf-8')}"
        )
    ],
)
If you prefer, you can also simply use an external URL like this:
Copy code
chat_ctx.add_message(
    role="user",
    content=[
        "Here is an image:",
        ImageContent(image="<https://example.com/image.jpg>")
    ],
)
Feel free to reach out if you'd like me to walk you through anything else or if you want to explore more about Letta's vision features. I’m here to help! Sources: Realtime data | LiveKit Docs | Large language model (LLM) integrations | LiveKit Docs | examples/data-streams/data_streams.py | Web and mobile frontends | LiveKit Docs | livekit-agents/livekit/agents/llm/_provider_format/openai.py