Hi Team, Suddenly after the a week of druid upgrad...
# general
s
Hi Team, Suddenly after the a week of druid upgrade from 25 to 27 and Spliting the coordinator in to "coordinator and overlord separate process" we are suddenly seeing spike of the below errors: almost 80% of the tasks are being failed.
Copy code
Task [index_kafka_ds_radio_node_rollup_364f1ac5c2f7f56_opjlflih] failed to return start time, ki..
Task [index_kafka_ds_50201fa781026ee_higdnkik] failed to respond to [set end offsets] in a timely...
I have tried increasing the Number of retries (ChatRetries to 20) in ingestion spec and http timeout to 30S. we are still facing same issues. we are using kubernetes deployment and in the overlord logs, its shows connection refused for the middlemanagers (i dont see any restarts of MM)