bitter-bird-28505
09/22/2025, 8:59 AMtall-airline-16756
09/22/2025, 8:59 AMtall-airline-16756
09/22/2025, 9:01 AMbitter-bird-28505
09/22/2025, 9:02 AMtall-airline-16756
09/22/2025, 9:02 AMRealtimeModel class. This model is great because it supports dynamic context injection through the update_chat_ctx method, letting you add retrieved information as deferred instructions.
2. Next, create a RAG retriever that taps into a vector database (like Annoy) to find relevant info based on user queries. The retrieved context is then added into the conversation as a system message wrapped in <instruction> tags.
3. Finally, integrate this retriever into your agent as a function tool, so it can dynamically fetch and inject context during real-time chats.
If you'd like, I’m happy to provide the full Python code for this setup or help you get started with components like the vector database. Just let me know how I can assist!
Sources: LiveKit Docs | delay-handling/agent.py | examples/README.md | rag/README.mdbitter-bird-28505
09/22/2025, 9:05 AMtall-airline-16756
09/22/2025, 9:06 AMimport asyncio
from livekit.agents import Agent, AgentSession, JobContext, WorkerOptions, cli, llm
from livekit.plugins.ultravox.realtime import RealtimeModel
from livekit.plugins import openai as openai_plugin
import annoy
import pickle
from pathlib import Path
class RAGRetriever:
def __init__(self, index_path: str, data_path: str):
self._index_path = Path(index_path)
self._data_path = Path(data_path)
self._annoy_index = annoy.AnnoyIndex(1536, 'angular')
self._annoy_index.load(str(self._index_path))
with open(self._data_path, 'rb') as f:
self._paragraphs_by_uuid = pickle.load(f)
async def search(self, query: str) -> str:
query_embedding = await openai_plugin.create_embeddings(
input=[query],
model="text-embedding-3-small",
dimensions=1536,
)
results = self._annoy_index.get_nns_by_vector(query_embedding[0].embedding, n=1)
if results:
paragraph_uuid = results[0]
return self._paragraphs_by_uuid.get(paragraph_uuid, "No relevant information found.")
return "No relevant information found."
class VoiceAssistantWithRAG(Agent):
def __init__(self):
super().__init__(
instructions="You are a helpful assistant.",
llm=RealtimeModel(model_id="fixie-ai/ultravox", voice="Mark"),
)
self._retriever = RAGRetriever("index.annoy", "data.pkl")
self.add_tool(self._rag_search)
@llm.function_tool
async def _rag_search(self, query: str):
context = await self._retriever.search(query)
chat_ctx = llm.ChatContext()
chat_ctx.add_message(role="system", content=f"<instruction>{context}</instruction>")
await self.session.update_chat_ctx(chat_ctx)
return f"Retrieved context for query: {query}"
async def entrypoint(ctx: JobContext):
session = AgentSession()
await session.start(agent=VoiceAssistantWithRAG(), room=ctx.room)
if __name__ == "__main__":
cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
For TTS and SST, here’s an example using OpenAI's real-time model:
from livekit.agents.voice import AgentSession
from livekit.plugins import openai
class WeatherAgent:
def __init__(self):
self.llm = openai.realtime.RealtimeModel(modalities=["text"])
self.tts = openai.TTS(voice="ash")
async def entrypoint(ctx):
session = AgentSession()
agent = WeatherAgent()
await session.start(agent=agent, room=ctx.room)
If you have any questions or need a hand walking through these examples, please don’t hesitate to ask—I’m here to help!
Sources: Realtime model integrations | LiveKit Docs | OpenAI TTS integration guide | LiveKit Docs | examples/other/text-to-speech/elevenlabs_tts.py | Groq TTS integration guide | LiveKit Docs | LMNT TTS integration guide | LiveKit Docsbitter-bird-28505
09/22/2025, 9:37 AMtall-airline-16756
09/22/2025, 9:38 AMfrom livekit import agents, rtc
from livekit.plugins import ultravox
async def run(ctx: agents.JobContext):
# Wait for the first participant
await ctx.wait_for_participant()
# Initialize the Ultravox plugin
ultravox_model = ultravox.RealtimeModel(api_key="your_api_key")
# Create a new realtime session
stt_session = ultravox_model.session()
# Process audio from a participant's microphone track
for _, participant in ctx.room.participants.items():
audio_stream = rtc.AudioStream(participant.tracks[1].pub.track)
async for audio_event in audio_stream:
stt_session.push_audio(audio_event.frame)
break # Process only the first participant for this example
# Gracefully close the session when done
await stt_session.aclose()
async def handle_stt_events(session: ultravox.RealtimeSession):
# Placeholder for handling transcription events
pass
This example walks you through initializing the Ultravox real-time model, capturing audio from a participant’s microphone, and sending it to the STT session for transcription. If you have any questions or need a hand with anything else, just let me know—I’m here to help!
Sources: livekit-plugins/livekit-plugins-baseten/livekit/plugins/baseten/stt.pybitter-bird-28505
09/22/2025, 1:42 PMtall-airline-16756
09/22/2025, 1:43 PMImageContent, either by providing base64-encoded data or by linking to external URLs. Here’s a little example to help you get started:
import base64
from livekit.agents import ChatContext, ImageContent
# Example: Adding an image to the chat context
image_bytes = open("example.png", "rb").read()
chat_ctx = ChatContext()
chat_ctx.add_message(
role="user",
content=[
ImageContent(
image=f"data:image/png;base64,{base64.b64encode(image_bytes).decode('utf-8')}"
)
],
)
If you prefer, you can also simply use an external URL like this:
chat_ctx.add_message(
role="user",
content=[
"Here is an image:",
ImageContent(image="<https://example.com/image.jpg>")
],
)
Feel free to reach out if you'd like me to walk you through anything else or if you want to explore more about Letta's vision features. I’m here to help!
Sources: Realtime data | LiveKit Docs | Large language model (LLM) integrations | LiveKit Docs | examples/data-streams/data_streams.py | Web and mobile frontends | LiveKit Docs | livekit-agents/livekit/agents/llm/_provider_format/openai.py