calm-book-72252
09/22/2025, 9:38 PMtall-airline-16756
09/22/2025, 9:38 PMtall-airline-16756
09/22/2025, 9:39 PMcalm-book-72252
09/22/2025, 9:42 PMtall-airline-16756
09/22/2025, 9:42 PMChatContext. The system will automatically discard older messages as you get close to the model’s context limit.
Using these approaches can really help reduce unnecessary overhead. Keep in mind, the internal caching of the Google Realtime API isn’t something we can tweak from the outside, but these tips should make your experience smoother!
Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/real... | plugins/google/src/beta/realtime/realtime_api.ts | tools/server/README.mdcalm-book-72252
09/22/2025, 9:43 PMllm_prompt_tokens includes both input tokens and cached tokens from the conversation context usage_collector.ts:32-34 . The high token count you're seeing is likely due to the cumulative nature of conversation context being sent with each turn.
Root Cause Analysis
The OpenAI realtime model implementation shows that inputTokens includes the full conversation context for each response realtime_model.ts:1316 . This means:
• Turn 1: 7,552 tokens (initial context + user input)
• Turn 2: 12,207 tokens (previous context + turn 1 response + new user input)
The token details structure confirms this includes text, audio, and cached tokens api_proto.ts:225-234 .tall-airline-16756
09/22/2025, 9:44 PMinstructions parameter during initialization to define your system prompt.
• Update System Prompt Dynamically: If you need to change the system prompt during a session, the update_instructions method is there for you—just keep in mind that it will restart the session.
• Compress Context Window: To handle longer conversations smoothly, try using context_window_compression to manage the context size effectively.
• Truncate Chat Context: The system automatically truncates the chat context to 40 messages, which helps keep token usage in check.
• Monitor Token Usage: It’s really useful to track your token consumption with the handleUsageMetadata method, so you can see input, output, and total token counts clearly.
• Control Response Length: You can also set the max_output_tokens parameter to limit how long the model's responses are, making sure things stay concise.
• Adjust Transcription Settings: Lastly, tweaking input_audio_transcription and output_audio_transcription lets you control how much tokens come from audio transcription.
By using these strategies, you'll be able to optimize token usage and enjoy smoother performance. If you have any questions or want to dive deeper into any of these, I’m here to help!
Sources: livekit-plugins/livekit-plugins-ultravox/livekit/plugins/ultravox/realtime/... | plugins/google/src/beta/realtime/realtime_api.tscalm-book-72252
09/22/2025, 9:49 PMtall-airline-16756
09/22/2025, 9:50 PMinstructions parameter during initialization. This way, the Google Realtime API keeps the system prompt on the server side, so it won’t be included in every client message.
Here’s a simple example:
llm = google.beta.realtime.RealtimeModel(
instructions="You are a helpful assistant.",
# other parameters
)
If at any point during your session you want to update the system prompt, just use update_instructions(). This will restart the session and apply the new instructions smoothly, without needing to resend them in every turn.
Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/stt.pycalm-book-72252
09/22/2025, 9:52 PMtall-airline-16756
09/22/2025, 9:52 PMupdate_instructions() method—it’s all about when and how the instructions take effect:
• LLM Initialization: When you set instructions here, they become the base context for every session created from that model. Think of these as your starting point—they stay the same unless you decide to update them later.
• Using `update_instructions()`: This is a handy way to change the system instructions on the fly during an active session. Just keep in mind, you’ll need to restart the session for the new instructions to kick in.
Both approaches send instructions to the Google Realtime API through the system_instruction field. Initialization sets your foundation, while update_instructions() gives you flexibility to adjust things as you go. I hope this helps you feel more confident in choosing the right method for your needs!
Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime... | plugins/google/src/beta/realtime/realtime_api.tscalm-book-72252
09/22/2025, 9:53 PMclass InterviewBot(agents.Agent):
def __init__(self, context_vars=None) -> None:
self.context_vars = context_vars
super().__init__(
instructions=PROMPT.format(**(context_vars or {})),
)
AND
session = AgentSession(
llm=google.beta.realtime.RealtimeModel(
model="gemini-2.5-flash-live-preview",
language="en-US",
temperature=0.8,
voice="Puck",
context_window_compression=types.ContextWindowCompressionConfig(
sliding_window=types.SlidingWindow()
),
max_output_tokens=250,
),
allow_interruptions=True,
user_away_timeout=60,
)tall-airline-16756
09/22/2025, 9:54 PMAgent class initialization and in the AgentSession. It really comes down to their scope and timing:
• Agent Class Initialization: Think of these instructions as the foundation—they set the agent’s core behavior that stays consistent across all sessions. These are synced with the RealtimeModel when a session starts and won’t change unless you decide to update them.
• AgentSession: These instructions are more flexible and tailored to each session. You can use them to override or add to the agent’s base instructions for just one interaction, like when using the generate_reply() method. They only affect that specific response and don’t change the agent’s overall behavior.
This setup is great because it lets your agents keep a steady personality while still being able to adapt their replies to fit different situations or user needs. If you have any more questions or want to dive deeper, I’m here to help!
Sources: livekit-agents/livekit/agents/voice/agent_activity.py | examples/src/gemini_realtime_agent.tscalm-book-72252
09/22/2025, 9:56 PMtall-airline-16756
09/22/2025, 9:56 PMAgentSession class instead of the Agent class, especially if the issue is related to the system prompt being sent repeatedly. By dynamically providing instructions for specific interactions with generate_reply(), you can avoid sending the full system prompt every time, which could make things smoother.
That said, if the problem is coming from how the system prompt is handled at the session level (like if it’s being sent automatically by the backend), this method might not completely fix it. In those situations, focusing on optimizing the initialization instructions or managing context compression could be a better way to go.