This message was deleted.
# troubleshooting
s
This message was deleted.
a
so I'm assuming some setting or default may have changed between 0.23.0 and 24.0.0
That's possible. A flavor of 24.0.0 is running on many of our prod clusters and we haven't seen such a drop in throughput.
You could take stack traces of a slow task to see if it's doing something expensive.
n
Can you walk me through doing that (or point me in the right direction at least)? We're running druid containerised within kubernetes
a
one way is to issue
-3
signal via
kill -3
command to the task. so if task pid is 2, you can do
kill -3 2
in the container. This will result in thread dump getting logged in the task logs. Then you can look at the stack trace in the task logs.
👍 1
you would want to capture them from the task JVMs (
internal peon
)
if you post the
.std
or
.html
files we can take a look
n
Apologies for digging up an old thread but I've just got into a position to look at this again. Within the peon logs, we are seeing (debug level) fetch timeouts to kafka that appear to coincide with the drop in ingest (feels kinda obvious). I'm keen to work out what's causing these as the kafka cluster itself is running fine with no errors in the logs. I'm assuming the peon task might be blocking on something which is causing the fetch request to timeout somehow. I'm unable (easily) to export the flame graph from site. I have a 5 minute sample of an indexing task misbehaving so I'm keen to learn how to analyze and understand the process. For kafka ingest, what do I need to look for in the flame graph (specific threads and states etc)?
a
there will be a
task-runner-*
thread that's particularly important.