Slackbot
10/18/2022, 4:12 PMThomas
10/18/2022, 4:14 PM{
"millisToWaitBeforeDeleting": 900000,
"mergeBytesLimit": 524288000,
"mergeSegmentsLimit": 100,
"maxSegmentsToMove": 1000,
"percentOfSegmentsToConsiderPerMove": 100,
"useBatchedSegmentSampler": true,
"replicantLifetime": 15,
"replicationThrottleLimit": 1000,
"balancerComputeThreads": 200,
"emitBalancingStats": false,
"killDataSourceWhitelist": [],
"killPendingSegmentsSkipList": [],
"maxSegmentsInNodeLoadingQueue": 1000,
"decommissioningNodes": [],
"decommissioningMaxPercentOfMaxSegmentsToMove": 70,
"pauseCoordination": false,
"replicateAfterLoadTimeout": false,
"maxNonPrimaryReplicantsToLoad": 2147483647
}kfaraz
10/18/2022, 4:23 PMsegment/moved/count . It denotes the number of segments chosen for balancing in a coordinator cycle. The other metric to look at would be segment/assigned/count. This is the number of segments that were missing from historicals and were assigned for loading. The assigned count is high when there are new published segments in the system that must be made available on historicals or when segments are being moved across tiers (if you have multiple tiers).Thomas
10/18/2022, 4:25 PMkfaraz
10/18/2022, 4:26 PMcachingCost?Thomas
10/18/2022, 4:27 PMcachingCostThomas
10/18/2022, 4:28 PMcachingCost it was 6Thomas
10/18/2022, 4:29 PMThomas
10/18/2022, 4:29 PMkfaraz
10/18/2022, 4:30 PMmaxSegmentsInNodeLoadingQueueThomas
10/18/2022, 4:31 PMThomas
10/18/2022, 4:31 PMThomas
10/18/2022, 4:32 PMkfaraz
10/18/2022, 4:32 PMcoordinator/time before and after the change?
With the filter duty = org.apache.druid.server.coordinator.duty.RunRuleskfaraz
10/18/2022, 4:33 PMThomas
10/18/2022, 4:35 PMcachingCost and got more “stable”kfaraz
10/18/2022, 4:35 PMcachingCost is supposed to be faster.Thomas
10/18/2022, 4:36 PMkfaraz
10/18/2022, 4:36 PMsegment/assigned/count look? Are we assigning enough segments in each run for loading?Thomas
10/18/2022, 4:38 PMkfaraz
10/18/2022, 4:40 PMcachingCost. Because the number of segment assignments is really the only thing that can be affected by the strategy. Since the runs are also faster and the number of assignments is okay, the issue is something else.kfaraz
10/18/2022, 4:44 PMsegment/assigned/count much smaller than the load queue size? If not, you could consider increasing it further if you have more segments being published.Thomas
10/18/2022, 4:48 PMkfaraz
10/18/2022, 4:51 PMThomas
10/19/2022, 8:38 AMsegment/assigned/count is smaller than the segment.loadQueuekfaraz
10/19/2022, 8:41 AMsegment/assigned/count is always smaller than maxSegmentsInNodeLoadingQueue, then we don't need to increase it further.kfaraz
10/19/2022, 8:44 AMThomas
10/19/2022, 8:44 AMThomas
10/19/2022, 8:44 AMcostkfaraz
10/19/2022, 8:45 AMThomas
10/19/2022, 8:47 AMcoordinator.20221012.log:2022-10-12T14:16:51,586 ERROR [Coordinator-Exec--0] org.apache.druid.server.coordinator.duty.BalanceSegments - [_default_tier]: Balancer move segments queue has a segment stuck: {class=org.apache.druid.server.coordinator.duty.BalanceSegments, segment=xxxx_2021-11-19T00:00:00.000Z_2021-11-20T00:00:00.000Z_2021-11-30T17:20:23.435Z_6, server=DruidServerMetadata{name='druid-historical-xxx:8083', hostAndPort='druid-historical-xxx:8083', hostAndTlsPort='null', maxSize=6000000000000, tier='_default_tier', type=historical, priority=0}}Thomas
10/19/2022, 8:48 AMcost strategy and I’m not seeing those unavailable segments anymorekfaraz
10/19/2022, 8:48 AMkfaraz
10/19/2022, 8:49 AMAnd btw, yesterday I’ve switched back toInteresting.strategy and I’m not seeing those unavailable segments anymorecost
cachingCost does have some cost computation-based issues but it shouldn't really cause unavailability of any segment.kfaraz
10/19/2022, 8:52 AMCurrent replication: [...] . Do you see logs like this for the segment in question?Thomas
10/19/2022, 9:15 AMkfaraz
10/19/2022, 9:15 AMThomas
10/19/2022, 9:23 AMThomas
10/19/2022, 9:23 AMThomas
10/19/2022, 9:24 AM