When using the gpt realtime api, I am seeing a ver...
# ask-ai
m
When using the gpt realtime api, I am seeing a very high token count and also degradation as the conversation gets bigger longer than a few minutes. I suspect something with the context window and the realtime api. Can you give me some leads to help debug?
t
Thanks for reaching out, Austin. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I understand how the high token count and the performance slowdown you're noticing with the GPT Realtime API during long conversations can be a bit frustrating. This usually happens because of the way the context window and token management work. The system has a 40-message limit on the chat context, and once this limit is passed, it trims the context to keep things running smoothly. This trimming can sometimes affect how well the conversation flows and remembers earlier parts. To help you work through this, here are a few friendly tips you might find useful: • Keep an eye on token usage with the metrics available to you, so you know when you’re getting close to the limits. • Try managing the conversation context by summarizing or saving older parts of the chat, which can lighten the load. • Consider optimizing your inputs by reducing any unnecessary audio or text, which helps keep token counts down. • Look into adjusting the maximum tokens allowed for responses or explore settings like
contextWindowCompression
if that’s an option for you. These approaches can really help keep the context size manageable and improve performance. If you’d like, feel free to share more about your setup or how your conversations typically go—I’d be happy to offer more personalized guidance! Sources: livekit-plugins/livekit-plugins-aws/livekit/plugins/aws/experimental/realti... | plugins/openai/src/realtime/realtime_model.ts
m
how can I track token usage
t
I'm happy to help you track token usage in the GPT Realtime API! You can make use of the built-in metrics system, which emits
RealtimeModelMetrics
events. These events give you detailed insights into token usage, including input, output, and total tokens, along with breakdowns by modality like text or audio. Here’s a friendly guide on how to track token usage: 1. Listen for
metrics_collected
Events:
You can attach a listener to your
AgentSession
to capture these events and pull out token usage info. Here's an example to get you started:
Copy code
# Python Example
@session.on("metrics_collected")
def on_metrics_collected(ev):
    if ev.metrics.type == "realtime_model_metrics":
        print(f"Input tokens: {ev.metrics.input_tokens}")
        print(f"Output tokens: {ev.metrics.output_tokens}")
        print(f"Total tokens: {ev.metrics.total_tokens}")
Copy code
// JavaScript Example
session.on('metrics_collected', (ev) => {
    if (ev.metrics.type === 'realtime_model_metrics') {
        console.log(`Input tokens: ${ev.metrics.inputTokens}`);
        console.log(`Output tokens: ${ev.metrics.outputTokens}`);
        console.log(`Total tokens: ${ev.metrics.totalTokens}`);
    }
});
1. Use the
UsageCollector
to Aggregate Usage:
This handy class helps you gather metrics across a session, including token counts from LLM and Realtime API interactions. Here’s how you can use it:
Copy code
# Python Example
usage_collector = metrics.UsageCollector()
@session.on("metrics_collected")
def on_metrics_collected(ev):
    usage_collector.collect(ev.metrics)
summary = usage_collector.get_summary()
print(f"Total LLM prompt tokens: {summary.llm_prompt_tokens}")
Copy code
// JavaScript Example
const usageCollector = new metrics.UsageCollector();
session.on('metrics_collected', (ev) => {
    usageCollector.collect(ev.metrics);
});
const summary = usageCollector.getSummary();
console.log(`Usage: ${JSON.stringify(summary)}`);
I hope these steps make it easier for you to monitor and manage your token usage smoothly. If you have any questions or need a hand with the implementation, please don’t hesitate to ask—I’m here to help! Sources: livekit-agents/livekit/agents/metrics/utils.py | agents/src/metrics/base.ts | Logs, metrics, and telemetry | LiveKit Docs
m
ok I implemented this, and see the following after a 5 min voice conversation with a realtime agent.
Copy code
Metrics summary: UsageSummary(llm_prompt_tokens=110601, llm_prompt_cached_tokens=96768, llm_input_audio_tokens=10251, llm_completion_tokens=5020, llm_output_audio_tokens=3389, tts_characters_count=0, tts_audio_duration=0.0, stt_audio_duration=0.0)
can you help me understand these metrics
t
I'd love to help you understand the metrics in your
UsageSummary
better: • llm_prompt_tokens: This is the total number of tokens sent to the LLM as prompts, including both text and audio tokens. • llm_prompt_cached_tokens: These are tokens reused from cache, which helps reduce redundant processing and improve efficiency. • llm_input_audio_tokens: Tokens that come from audio input processed by the Realtime API. • llm_completion_tokens: Tokens generated by the LLM in its responses (completions). • llm_output_audio_tokens: Tokens generated as audio output by the LLM. • tts_characters_count: Total characters processed by the Text-to-Speech (TTS) system. • tts_audio_duration: Total duration (in seconds) of audio output generated by TTS. • stt_audio_duration: Total duration (in seconds) of audio input processed by the Speech-to-Text (STT) system. Looking at your specific metrics: • You have a high number of prompt tokens (
llm_prompt_tokens=110601
), with a great portion cached (
llm_prompt_cached_tokens=96768
), which is fantastic because it helps optimize performance. • Audio input contributed
llm_input_audio_tokens=10251
tokens, and the LLM generated
llm_completion_tokens=5020
tokens in its responses. • Audio output from the LLM accounted for
llm_output_audio_tokens=3389
tokens. • It looks like there was no Text-to-Speech (TTS) or Speech-to-Text (STT) processing involved (
tts_characters_count=0
,
tts_audio_duration=0.0
,
stt_audio_duration=0.0
). Please feel free to reach out if you'd like me to explain anything further or help you optimize these metrics. I'm here to support you! Sources: livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime... | Logs, metrics, and telemetry | LiveKit Docs
m
how can I manage the conversation history to reduce the amount of cached tokens?
t
To help you manage your conversation history and keep the number of cached tokens in check, here are some friendly strategies you might find useful: 1. Truncate Conversation History: Try limiting the number of messages you send to the LLM by trimming the context. For instance, you could keep just the last 15 messages before sending the request:
Copy code
async def truncate_context(assistant: VoicePipelineAgent, chat_ctx: llm.ChatContext):
    if len(chat_ctx.messages) > 15:
        chat_ctx.messages = chat_ctx.messages[-15:]  # Keep only last 15 messages

agent = VoicePipelineAgent(
    before_llm_cb=truncate_context,
)
2. Summarize History: Instead of sending the entire conversation history, summarizing previous interactions into a brief format can really help. This way, you reduce token load but still keep the important context intact. 3. Control Context Persistence: When handing off between agents, avoid passing the full history unless it’s necessary. Selectively sharing context can make a big difference in minimizing token usage. 4. Monitor History Usage: It’s a good idea to include only the relevant parts of the conversation in the context to prevent processing redundant tokens. These tips should make a noticeable difference in reducing token usage and boosting performance. If you’d like, I’m here to help you implement any of these strategies—just let me know! Sources: OpenAI Realtime API integration guide | LiveKit Docs | VoicePipelineAgent | LiveKit Docs | VoicePipelineAgent | LiveKit Docs | Azure OpenAI Realtime API and LiveKit | LiveKit Docs | Live Conversations with AI using ChatGPT and WebRTC
m
how can I access this chat_ctx in an agent function call
t
I'm happy to help you access
chat_ctx
in an agent function call! You can do this by using
self.chat_ctx
within any class that inherits from
Agent
. This gives you a handy, read-only view of the entire conversation history, including all messages, function calls, and outputs. Here’s a simple example showing how to access
chat_ctx
inside a function tool:
Copy code
from livekit.agents import Agent, llm

class MyAgent(Agent):
    @llm.function_tool()
    async def get_weather(self, location: str):
        # Access full conversation history
        for item in self.chat_ctx.items:
            print(item.type, getattr(item, "role", None), getattr(item, "text_content", None))
        return f"The weather in {location} is sunny."
If you’d like to make changes to the context, you can create a copy of
chat_ctx
and update it using
update_chat_ctx()
. Here’s how you might do that:
Copy code
from livekit.agents.llm import ChatContext

@llm.function_tool()
async def my_tool(self):
    # Create a modified copy of the chat context
    new_ctx = self.chat_ctx.copy()
    new_ctx.add_message(role="assistant", content="I logged this internal step.")
    # Apply the updated context
    await self.update_chat_ctx(new_ctx)
If anything feels unclear or if you want to dive deeper into this, please don’t hesitate to reach out. I’m here to support you every step of the way! Sources: Workflows | LiveKit Docs | livekit-agents/livekit/agents/voice/agent_activity.py
m
how would I clear the chat ctx ?
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!