<@U0318CX1NH4> Starting :thread: to discuss <https...
# dev
k
I feel having per work configs for pr https://github.com/apache/druid/pull/16889 is like really bad for the end user
The end user is not worried about per worker semantics but global limits
k
Yeah, I would have preferred having a global ratio myself. But in effect, even per worker ratio gives you the same (or close enough) overall ratio.
Also,
parallelIndexTaskSlotRatio
already has the same per-worker logic.
So, it probably makes sense to continue in the same vein.
k
But in effect, even per worker ratio gives you the same (or close enough) overall ratio.
Not really since we can have logic like "all controllers of msq go to tier A" and worker go to tier B
TierB is on spot instances
k
Hmm, yeah, with tiered workers, it becomes a problem.
k
Some folks have container packing optimizations
I think we should maintain a state in the taskRunner
and then mutate it on schedule and on remove
We might have to persist it somewhere
k
Sure, having it on the TaskQueue/TaskRunner seems fine to me.
k
on populate it on the fly
Which should be fine since we already poll all running tasks on overlord startup
iirc
k
The other task slot ratios used by Druid i.e. compactionTaskSlotRatio and killTaskSlotRatio are enforced by the coordinator.
Yeah, it need not be persisted.
k
Exactly
k
Also, do you feel these should be dynamic configs instead of runtime properties? I feel having them be dynamic config could be useful.
k
Yeah dynamic configs are better since it would not require an overlord restart
k
compactionTaskSlotRatio and killTaskSlotRatio are already dynamic.
k
If we have enough admin auth security which we should already have then I think its worth thinking in that direction
k
In that case, let's also deprecate
parallelIndexTaskSlotRatio
since it is a per-worker config, and we should deprecate it in favor of the new config, whenever we add it.
➕ 1
Since you mentioned tiers, do you think defining the task slot ratios makes more sense on a tier level?
In that case, we could piggy back on the worker tier dynamic config. We would just need to add a new field there and handle it in
WorkerSelectUtils.selectWorker()
somehow.
k
I feel we can start with a global/default tier limit and work on making tier specific limits based on adoption
👍 1
m
Hi @Karan Kumar @kfaraz, just to summarise, are you suggesting going back to the original option with TaskQueue limit but extended with customisable task types stored in map? In case of global limit absoluteLimit map storing integer limits would also make sense. I do like an idea about dynamic configuration, but i'm not sure how to make it fly, as exampled compactionTaskSlotRatio and killTaskSlotRatio are related to the coordinator duties and seem to be checked inside of run() there
k
are you suggesting going back to the original option with TaskQueue limit but extended with customisable task types stored in map
Yes You would most likely have to add stuff to https://github.com/apache/druid/blob/d982727a298526948961e470df0eade6344f1ea8/inde[…]a/org/apache/druid/indexing/overlord/http/OverlordResource.java
Basically add a new json to
DefaultWorkerBehaviorConfig
private field TaskLimits taskLimits;