Thinking about Threshold prioritization strategy (...
# dev
m
Thinking about Threshold prioritization strategy (https://druid.apache.org/docs/latest/configuration/#threshold-prioritization-strategy) and was wondering if the following changes would make sense or not (haven’t tested any of these and don’t know if they will be useful in practise or not) • What if we can stack the violatesThreshold and penalize a query more if it violates multiple threshold. i.e.
Copy code
int toAdjust = 0
    if (violatesPeriodThreshold) {
      toAdjust += adjustment;
    }
    if (violatesDurationThreshold) {
      toAdjust += adjustment;
    }
    if (violatesSegmentThreshold) {
      toAdjust += adjustment;
    }
    if (violatesSegmentRangeThreshold) {
      toAdjust += adjustment;    
    }
    if (toAdjust != 0) {
      final int adjustedPriority = theQuery.context().getPriority() - toAdjust;
      return Optional.of(adjustedPriority);
    }
• What if we can set the adjustment value for each Threshold seperately? i.e.
Copy code
int toAdjust = 0
    if (violatesPeriodThreshold) {
      toAdjust += periodThresholdAdjustment;
    }
    if (violatesDurationThreshold) {
      toAdjust += durationThresholdAdjustment;
    }
    if (violatesSegmentThreshold) {
      toAdjust += segmentThresholdAdjustment;
    }
    if (violatesSegmentRangeThreshold) {
      toAdjust += segmentRangeThresholdAdjustment;    
    }
    if (toAdjust != 0) {
      final int adjustedPriority = theQuery.context().getPriority() - toAdjust;
      return Optional.of(adjustedPriority);
    }
The motivation for the first change is that if a query that violate N thresholds, it should be penalize more (not equal) to another query that violate N-1 thresholds. The motivation for the second change is that some violate are worst than other. i.e. periodThreshold is not that bad compare to segmentRangeThreshold. The prioritization value would then carry over to the Historical and can help with resources prioritization on Historical processing threadpool (related to this discussion https://apachedruidworkspace.slack.com/archives/C030CMF6B70/p1745436989786489). CC:@Gian Merlino @Clint Wylie
g
seems like an interesting approach
are you thinking of extending the threshold strategy or creating a new one?
m
extending the threshold strategy
g
it's sort of like a (simple) cost based strategy
m
making the current threshold strategy more flexible
it’s sort of like a (simple) cost based strategy
Yes. Extending from the thresholds we already have
the issue we have with the current threshold strategy is that it’s giving a lot of false positive where we are penalizing queries that are actual fast/quick and it’s hard to make these a different priotization value that the one that is really slow/expensive
g
yeah to me this makes sense