"metrics": { "total_requests": 3, "tot...
# ask-ai
b
"metrics": { "total_requests": 3, "total_duration": 15.774682521820068, "total_ttft": 0.8553550243377686, "avg_ttft": 0.28511834144592285, "total_input_tokens": 12144, "total_output_tokens": 66, "total_tokens": 12210, "avg_tokens_per_second": 4.183919385300249, "total_audio_input_tokens": 3, "total_text_input_tokens": 12141, "total_image_input_tokens": 0, "total_cached_tokens": 0, "total_text_output_tokens": 66, "total_audio_output_tokens": 0, "total_image_output_tokens": 0, "cancelled_requests": 0 } above are the sum metrices of the session now i am using gemini live api how to calculate the cost the rates are different for text input text output and audio input and audio output how will calculate the cost here i dont see the audio output here
t
Thanks for reaching out, Aman. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
To calculate the cost using the Gemini Live API metrics, you’ll want to apply the pricing rates for text input, text output, audio input, and audio output to the token counts you have. I see you’ve got detailed token counts for text and audio input/output, which is great! The next step is to find the specific pricing rates to move forward. Here’s a simple way to calculate the cost: 1. First, check out the Google Gemini API pricing documentation to find the rates for text input, text output, audio input, and audio output. 2. Then, multiply the token counts for each category (like
total_text_input_tokens
,
total_text_output_tokens
,
total_audio_input_tokens
) by their corresponding rates. 3. Finally, add up the costs from all categories to get your total cost. If you’d like, I’m here to help you access the metrics programmatically or guide you through the calculation process. Just let me know! Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime...
b
i dont see the audio output tokens .. even though i had a audio session with the model via livekit voice agent
t
The absence of audio output tokens in your metrics likely means that the Gemini Live API didn’t include audio-specific token details in its response. The LiveKit metrics system is set up to capture audio output tokens, but it relies on the API providing that information. Here’s a gentle way to check this further: 1. Take a look at the
usage_metadata
from the API response and see if there’s an entry in
response_tokens_details
where the
modality
is marked as
AUDIO
. 2. If you don’t find such an entry, it probably means the API isn’t categorizing or reporting audio output tokens. If you find that the API response is missing this data, it might be worth reaching out for further investigation or clarification from the API provider. Please feel free to ask if you'd like a hand inspecting the
usage_metadata
or if you have any other questions—I’m here to help! Sources: livekit-agents/livekit/agents/metrics/utils.py | plugins/openai/src/realtime/api_proto.ts
b
currently i am using @session.on("metrics_collected") def _on_metrics_collected(ev: MetricsCollectedEvent): session_metrics.add_metric(ev.metrics) to collect the metrics give me code to get the usage_metadata
t
The
@session.on("metrics_collected")
event doesn’t give direct access to the raw
usage_metadata
. Instead, it provides structured metrics like
RealtimeModelMetrics
or
LLMMetrics
, which are thoughtfully derived from the
usage_metadata
. Here’s a friendly example of how you can work with the available metrics:
Copy code
python @session.on("metrics_collected") def _on_metrics_collected(ev: MetricsCollectedEvent): m = ev.metrics if isinstance(m, RealtimeModelMetrics): print(f"Input Tokens: {m.input_tokens}, Output Tokens: {m.output_tokens}") print(f"Audio Input Tokens: {m.input_token_details.audio_tokens}") print(f"Text Input Tokens: {m.input_token_details.text_tokens}") print(f"Audio Output Tokens: {m.output_token_details.audio_tokens}") print(f"Text Output Tokens: {m.output_token_details.text_tokens}")
If you’re looking for the raw
usage_metadata
, it’s not directly available through the public API. But no worries! You can enable OpenTelemetry tracing to capture detailed trace data, including token usage and other helpful attributes. I’m here to help if you want assistance setting up tracing or if you have any other questions — just let me know! Sources: examples/voice_agents/langfuse_trace.py | plugins/google/src/beta/realtime/realtime_api.ts | Logs, metrics, and telemetry | LiveKit Docs