This message was deleted.
# troubleshooting
s
This message was deleted.
l
From my understanding the requests should be timed out at 5 min mark. Internally, the processing classes keep a track of the timeout and in case it is exceeded cancel the execution and throw the desired error Are you facing any troubles while using the query timeout or seeing any conflicting behaviour to above?
d
It’s not an exact science, but I was comparing the query timeout clocked by Superset and charts on Datadog. After the 5 minutes mark, it looks like there are still high volume activity on Historical side.
l
Does the volume reduce after a while, i.e. while the load is high at 5 min mark, does the load subside after say 2-3 mins?
d
it didn’t subside after 2-3 minutes. More like 15-20 minutes
l
Which version of Druid are you on? And are there any specific type of queries that you are running which are showing this behavior? It might be a bug similar to https://github.com/apache/druid/pull/12271.
d
We are at 0.23.0
I haven’t delve deep into what type of query yet
l
Ohh okay. Are there any dubious errors or warnings in the logs which might point to the queries not cancelling?
d
So this is interesting. We just had a major outage due to a number of bad “LIKE” queries. The most interesting relevant part of this is that all of Historical threads were stuck spinning forever performing groupBy. After 2 hours the cluster didn’t recover and the same threads that were stuck still stuck. We had to perform complete Historical restart to recover from this issue. I will paste relevant jstack when I am in front of my computer later.
👀 1
Starting to think that extremely long G1 GC pause is the culprit. (Which will happen if the user runs crazy big
LIKE
queries)