I am using python sdk to dispatch voice agents whi...
# ask-ai
c
I am using python sdk to dispatch voice agents which uses gemini live as the LLM. The issue I'm facing is the token usage comes out pretty high. LOGS:
Copy code
UsageSummary(llm_prompt_tokens=326670, llm_prompt_cached_tokens=0, llm_completion_tokens=1093, tts_characters_count=0, tts_audio_duration=0.0, stt_audio_duration=0.0)", "level": "INFO", "name": "__main__", "timestamp": "2025-09-22T15:07:01.088677+00:00"}
Do the realtime models send the SYSTEM PROMPT at every turn which leads to this?
t
Thanks for reaching out, Aditya. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I wanted to share a little insight about how the system prompt works in LiveKit's Python SDK. It’s actually sent just once during the session initialization, so it doesn’t get sent at every turn. This helps keep token usage down during your interactions. If you’re noticing higher token usage than expected, it could be related to other things like the length of user inputs or the model’s responses. I’m here to help if you’d like me to take a closer look or offer some tips on optimizing your token usage. Just let me know! Sources: livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/real...
c
No it's just a normal audio conversation.
t
I'd love to understand your setup a bit better so we can find the best solution together. Could you share a few details? • Are you using any custom prompts or special configurations? • What’s the usual length of user inputs and the model’s responses? • Have you noticed any particular situations where token usage tends to increase? Your insights will really help us pinpoint where we might optimize things. Thanks so much for sharing!
c
Yes I am using a custom prompt which is around 2500 tokens. The total tokens for user input and model response is this
Copy code
{
  "message": "RealtimeModel metrics",
  "level": "INFO",
  "name": "livekit.agents",
  "ttft": 23.36,
  "input_tokens": 17246,
  "cached_input_tokens": 0,
  "output_tokens": 28,
  "total_tokens": 17274,
  "tokens_per_second": 0.9,
  "timestamp": "2025-09-22T15:06:56.553969+00:00"
}
No
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!
c
By the metrics I tracked that system prompt is sent on every call