eager-appointment-68614
12/11/2025, 8:42 PMTokens: ↑ input 3.57M • cache hit 0.00% • ↓ output 21.68K • $ 0.00
I have no clue what happened in my self hosted setup but I am now able to have very long running conversations and finally got to witness the context compression in action. Using a full sized qwen3 coder 480B model. I'm able to make really good documentation for new codebases. I don't think the model has actually processed 3.57 million tokens in this chat, but regardless, a week ago things would break after just a few tool calls. Would love to get an accurate metric of the context length in my requests, but regardless, I'm very VERY happy right now.limited-student-10747
12/12/2025, 7:42 PMeager-appointment-68614
12/17/2025, 3:20 AM