This message was deleted.
# troubleshooting
s
This message was deleted.
i
we see index_kafka tasks that succeded in log, i.e.
Copy code
2022-11-21T13:08:34,252 INFO [task-runner-0-priority-0] org.apache.druid.indexing.worker.executor.ExecutorLifecycle - Task completed with status: {
  "id" : "index_kafka_sp_campaigns_realtime_aggregation_eef1b68744abc48_jfncjiei",
  "status" : "SUCCESS",
  "duration" : 1798848,
  "errorMsg" : null,
  "location" : {
    "host" : null,
    "port" : -1,
    "tlsPort" : -1
  }
}
however task status if FAILED with errorMsg: “errorMsg”: “No task in the corresponding pending completion taskGroup[4] succeeded before completion timeout ela...” our completionTimeout was 30mins(increased it to 60 mins) but there is gap in data which we would like to fill
@Michael Taranov @Ori Amichay
r
For a successful task even though the handoff failed, the segment would've been published to deep storage and would be eligible for loading through the run rules coordinator duty. could you check if there is a gap in your data loaded in druid? It would be done via UI in datasources tab or using API (https://druid.apache.org/docs/latest/operations/api-reference.html#segment-loading-by-datasource)
m
Hi @Rohan Garg Are you talking about following ? “Availability detail” under data sources
r
Hi Michael! yes, there's availability % as well in that tab
m
In our case we didn’t have any availability issues (or at least no indication of this) . Just number of ingestion tasks failed because of
completionTimeout,
which @Igor Berman mentioned, still not sure why it may take so much time.
r
oh I see, I thought there was availability concern as well due to this message :
but there is gap in data which we would like to fill
I think I misunderstood that Regarding more investigation on the long time for handoff, you can probably check the runtime of
RunRules
coordinator duty around that time (it is emitted as the metric
coordinator/time
) and see if it was high as well. If there are lot of segments being handed-off around the same, we've seen issues around their completion. Multiple fixes are being done in the upcoming release to reduce handoff time.
✅ 1
i
thanks @Rohan Garg for the info, we will check coordinator/time metric
we see high times for tasks I’ve found in coordinator logs some locking with kill tasks that we schedule to remove unused segments
Copy code
2022-11-21T11:54:15,763 INFO [TaskQueue-Manager] org.apache.druid.indexing.overlord.TaskLockbox - Cannot create a new taskLockPosse for request[TimeChunkLockRequest{lockType=EXCLUSIVE, groupId='kill_sp_campaigns_realtime_aggregation_after_dedup_pjjhnijp_2022-11-14T00:00:00.000Z_2022-11-21T00:00:00.000Z_2022-11-21T03:00:01.515Z', dataSource='sp_campaigns_realtime_aggregation_after_dedup', interval=2022-11-14T00:00:00.000Z/2022-11-21T00:00:00.000Z, preferredVersion='null', priority=0, revoked=false}] because existing locks[[TaskLockPosse{taskLock=TimeChunkLock{type=EXCLUSIVE, groupId='index_kafka_sp_campaigns_realtime_aggregation_after_dedup', dataSource='sp_campaigns_realtime_aggregation_after_dedup', interval=2022-11-20T23:00:00.000Z/2022-11-21T00:00:00.000Z, version='2022-11-20T23:00:00.599Z', priority=75, revoked=false}, taskIds=[index_kafka_sp_campaigns_realtime_aggregation_after_dedup_2899dc64f6e52cd_cpfinpbf, index_kafka_sp_campaigns_realtime_aggregation_after_dedup_5cccc69339458cf_mhoddmpl, index_kafka_sp_campaigns_realtime_aggregation_after_dedup_a173f06e2e4bc07_hpbcbmki, index_kafka_sp_campaigns_realtime_aggregation_after_dedup_3bcd3eee475f061_lndfdjfb, index_kafka_sp_campaigns_realtime_aggregation_after_dedup_4e232d8b7824f50_nejlbefp, index_kafka_sp_campaigns_realtime_aggregation_after_dedup_bc56833a08847ae_imcnnool, index_kafka_sp_campaigns_realtime_aggregation_after_dedup_6f07318c19267f8_oieefadb, index_kafka_sp_campaigns_realtime_aggregation_after_dedup_7bf9902c552596f_hcnhlbbk, index_kafka_sp_campaigns_realtime_aggregation_after_dedup_57cc76904e3f9a9_kdedccld]}]] have same or higher priorities
We have cron that submits kill tasks to cleanup data periodically, so we most probably caused this “deadlock” by setting end of interval to be within retention window. I’m going to change this cron so it will submit kill tasks so that interval wont intersect with retained window
r
Thanks for the logs - they show that the kill task wasn't able to get the lock on the interval
2022-11-14T00:00:00.000Z/2022-11-21T00:00:00.000Z
since kafka ingestion was writing to that interval. It doesn't indicate any deadlock, just that the system has decided not to provide write lock to two tasks simultaneously. Also, druid always prefers ingestion tasks over kill tasks for locks. If you wish to avoid these conflicts proactively, then as you mentioned you can ensure that the kill interval is not overlapping the ingestion one