```Tokens: ↑ input 3.57M • cache hit 0.00% • ↓ out...
# success-stories
e
Copy code
Tokens: ↑ input 3.57M • cache hit 0.00% • ↓ output 21.68K • $ 0.00
I have no clue what happened in my self hosted setup but I am now able to have very long running conversations and finally got to witness the context compression in action. Using a full sized qwen3 coder 480B model. I'm able to make really good documentation for new codebases. I don't think the model has actually processed 3.57 million tokens in this chat, but regardless, a week ago things would break after just a few tool calls. Would love to get an accurate metric of the context length in my requests, but regardless, I'm very VERY happy right now.
👌 7
🙌 2
l
Awesome, if you want to share the prompts that you’ve used to be successful I’m sure the people in the community would appreciate it!
e
I've made no modification to the default tool-calling/starter prompts haha.