can I enable prompt caching for google realtime mo...
# ask-ai
c
can I enable prompt caching for google realtime models?
t
Thanks for reaching out, Aditya. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Thanks for reaching out! I wanted to share that prompt caching for Google Realtime models in LiveKit is handled internally, so it can’t be manually turned on, off, or adjusted through LiveKit’s API. This caching is managed by Google’s backend, and you can see how it’s working by checking the token usage metrics provided by the Google Realtime API. If you have any more questions or need help interpreting those metrics, I’m here for you! Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime... | plugins/google/README.md | tools/server/README.md
c
What are the best techniques to cache the tokens? I want to cache the system prompt
t
When working with Google Realtime models in LiveKit, I understand it can be a bit tricky since caching system prompts or tokens isn’t something you can directly configure. But don’t worry! There are a couple of ways you can optimize performance: • Minimizing Instruction Changes: Try setting the system prompt (instructions) just once at the start of your session. This helps avoid the extra work that comes with restarting sessions frequently. • Controlling Context Window Size: You can manage the chat history you send to the model by adjusting the
ChatContext
. The system will automatically discard older messages as you get close to the model’s context limit. Using these approaches can really help reduce unnecessary overhead. Keep in mind, the internal caching of the Google Realtime API isn’t something we can tweak from the outside, but these tips should make your experience smoother! Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/real... | plugins/google/src/beta/realtime/realtime_api.ts | tools/server/README.md
c
The system prompt seems to be sent automatically when it's agent's turn. Reference: For realtime models,
llm_prompt_tokens
includes both input tokens and cached tokens from the conversation context usage_collector.ts:32-34 . The high token count you're seeing is likely due to the cumulative nature of conversation context being sent with each turn. Root Cause Analysis The OpenAI realtime model implementation shows that
inputTokens
includes the full conversation context for each response realtime_model.ts:1316 . This means: • Turn 1: 7,552 tokens (initial context + user input) • Turn 2: 12,207 tokens (previous context + turn 1 response + new user input) The token details structure confirms this includes text, audio, and cached tokens api_proto.ts:225-234 .
t
I'm glad you're looking to manage the system prompt and reduce token usage for Google Realtime models in LiveKit! Here are some helpful tips to guide you through: • Set Initial System Prompt: You can start by using the
instructions
parameter during initialization to define your system prompt. • Update System Prompt Dynamically: If you need to change the system prompt during a session, the
update_instructions
method is there for you—just keep in mind that it will restart the session. • Compress Context Window: To handle longer conversations smoothly, try using
context_window_compression
to manage the context size effectively. • Truncate Chat Context: The system automatically truncates the chat context to 40 messages, which helps keep token usage in check. • Monitor Token Usage: It’s really useful to track your token consumption with the
handleUsageMetadata
method, so you can see input, output, and total token counts clearly. • Control Response Length: You can also set the
max_output_tokens
parameter to limit how long the model's responses are, making sure things stay concise. • Adjust Transcription Settings: Lastly, tweaking
input_audio_transcription
and
output_audio_transcription
lets you control how much tokens come from audio transcription. By using these strategies, you'll be able to optimize token usage and enjoy smoother performance. If you have any questions or want to dive deeper into any of these, I’m here to help! Sources: livekit-plugins/livekit-plugins-ultravox/livekit/plugins/ultravox/realtime/... | plugins/google/src/beta/realtime/realtime_api.ts
c
It's the system prompt which is the culprit which is being sent at every turn
t
To help you avoid sending the system prompt with every turn when using Google Realtime models, you can set it once using the
instructions
parameter during initialization. This way, the Google Realtime API keeps the system prompt on the server side, so it won’t be included in every client message. Here’s a simple example:
Copy code
llm = google.beta.realtime.RealtimeModel(
    instructions="You are a helpful assistant.",
    # other parameters
)
If at any point during your session you want to update the system prompt, just use
update_instructions()
. This will restart the session and apply the new instructions smoothly, without needing to resend them in every turn. Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/stt.py
c
What is the difference between defining the instruction in llm and defining it in the using the init function
t
Thanks for your question! Let me help clarify the difference between setting instructions during LLM initialization and using the
update_instructions()
method—it’s all about when and how the instructions take effect: • LLM Initialization: When you set instructions here, they become the base context for every session created from that model. Think of these as your starting point—they stay the same unless you decide to update them later. • Using `update_instructions()`: This is a handy way to change the system instructions on the fly during an active session. Just keep in mind, you’ll need to restart the session for the new instructions to kick in. Both approaches send instructions to the Google Realtime API through the
system_instruction
field. Initialization sets your foundation, while
update_instructions()
gives you flexibility to adjust things as you go. I hope this helps you feel more confident in choosing the right method for your needs! Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime... | plugins/google/src/beta/realtime/realtime_api.ts
c
Difference between
Copy code
class InterviewBot(agents.Agent):
    def __init__(self, context_vars=None) -> None:
        self.context_vars = context_vars
        super().__init__(
            instructions=PROMPT.format(**(context_vars or {})),
        )
AND
Copy code
session = AgentSession(
            llm=google.beta.realtime.RealtimeModel(
                model="gemini-2.5-flash-live-preview",
                language="en-US",
                temperature=0.8,
                voice="Puck",
                context_window_compression=types.ContextWindowCompressionConfig(  
                    sliding_window=types.SlidingWindow()
                ),
                max_output_tokens=250,
            ),
            allow_interruptions=True,
            user_away_timeout=60,
        )
t
Thanks for your question! Let me help clarify the difference between defining instructions in the
Agent
class initialization and in the
AgentSession
. It really comes down to their scope and timing: • Agent Class Initialization: Think of these instructions as the foundation—they set the agent’s core behavior that stays consistent across all sessions. These are synced with the
RealtimeModel
when a session starts and won’t change unless you decide to update them. • AgentSession: These instructions are more flexible and tailored to each session. You can use them to override or add to the agent’s base instructions for just one interaction, like when using the
generate_reply()
method. They only affect that specific response and don’t change the agent’s overall behavior. This setup is great because it lets your agents keep a steady personality while still being able to adapt their replies to fit different situations or user needs. If you have any more questions or want to dive deeper, I’m here to help! Sources: livekit-agents/livekit/agents/voice/agent_activity.py | examples/src/gemini_realtime_agent.ts
c
Will using the instructions in the AgentSession class isntead of Agent class could solve the problem?
t
It might be helpful to try using instructions in the
AgentSession
class instead of the
Agent
class, especially if the issue is related to the system prompt being sent repeatedly. By dynamically providing instructions for specific interactions with
generate_reply()
, you can avoid sending the full system prompt every time, which could make things smoother. That said, if the problem is coming from how the system prompt is handled at the session level (like if it’s being sent automatically by the backend), this method might not completely fix it. In those situations, focusing on optimizing the initialization instructions or managing context compression could be a better way to go.