Hey Druid Team! I have some doubts on `indexing-se...
# dev
h
Hey Druid Team! I have some doubts on
indexing-service
that can help me in fixing this issue. I am not able to find how Druid makes sure that replica tasks consuming from same partitions are made sure to not get scheduled on same workers. I don't find any patch preventing it in TaskMaster, TaskRunner or in WorkerSelectStrategy. I'm assuming as replica tasks are next to each other in TaskQueue, this would prevent to get scheduled on same workers (I might be wrong). I'm asking this because, the issue I pointed gets triggered when let's say on a worker, Task -> A, Task Group -> G moves to PUBLISHING and a new actively reading task -> B consuming from same TaskGroup -> G get scheduled on same worker, task A's StreamAppenderator thread
-appenderator-abandon
and
TASK]-publish
is not able to terminate. Is there any affinity on these appenderator threads to the TaskGroup or partitions which are preventing them from terminate ? This increases the probability of the issue occuring if replicas are increased and workers are decreased, active task gets failed when we reach this state. Any help on the open questions above would be appreciated. Thanks!