This message was deleted.
# troubleshooting
s
This message was deleted.
t
And here is the Coordinator dynamic config :
Copy code
{
  "millisToWaitBeforeDeleting": 900000,
  "mergeBytesLimit": 524288000,
  "mergeSegmentsLimit": 100,
  "maxSegmentsToMove": 1000,
  "percentOfSegmentsToConsiderPerMove": 100,
  "useBatchedSegmentSampler": true,
  "replicantLifetime": 15,
  "replicationThrottleLimit": 1000,
  "balancerComputeThreads": 200,
  "emitBalancingStats": false,
  "killDataSourceWhitelist": [],
  "killPendingSegmentsSkipList": [],
  "maxSegmentsInNodeLoadingQueue": 1000,
  "decommissioningNodes": [],
  "decommissioningMaxPercentOfMaxSegmentsToMove": 70,
  "pauseCoordination": false,
  "replicateAfterLoadTimeout": false,
  "maxNonPrimaryReplicantsToLoad": 2147483647
}
k
It is okay if there are always items in the load queue. This is most likely because of balancing. The coordinator constantly tries to find the optimal home for segments for improved query performance. So you will always see some segments moving around. You can verify this by looking at the metric
segment/moved/count
. It denotes the number of segments chosen for balancing in a coordinator cycle. The other metric to look at would be
segment/assigned/count
. This is the number of segments that were missing from historicals and were assigned for loading. The assigned count is high when there are new published segments in the system that must be made available on historicals or when segments are being moved across tiers (if you have multiple tiers).
t
Hello, thanks for answer. But I’m often seeing unavailable segments which is annoying for some internal teams.
k
Did you start seeing this issue after upgrading to 24 or after using
cachingCost
?
t
After using
cachingCost
Currently, i have an avg of 45 segments unavailable per day. Before
cachingCost
it was 6
And the segments I’m talking about are not fresh ones
Got some from Feb 2022 like an hour ago
k
Hmm, that is problematic. Have you tried increasing the load queue size?
maxSegmentsInNodeLoadingQueue
t
Yep, it was 500, increased to 1000
I noticed that historicals and coordinator was not using a lof of resources when balancing was happening
I wanted to speed up the process
k
Do you have an idea of the metric
coordinator/time
before and after the change? With the filter
duty = org.apache.druid.server.coordinator.duty.RunRules
I want to see if the coordinator runs are taking more time after you started using cachingCost strategy.
t
It decreased after changing to
cachingCost
and got more “stable”
k
That's good.
cachingCost
is supposed to be faster.
t
Yup
k
How does
segment/assigned/count
look? Are we assigning enough segments in each run for loading?
t
Same trend before and after
k
Hmm, then I am inclined to believe that the issue is probably not with
cachingCost
. Because the number of segment assignments is really the only thing that can be affected by the strategy. Since the runs are also faster and the number of assignments is okay, the issue is something else.
Is the
segment/assigned/count
much smaller than the load queue size? If not, you could consider increasing it further if you have more segments being published.
t
Thanks a lot for taking the time to troubleshoot this 🙂 I’m gonna check this tomorrow and post results in the thread
k
Sounds good 🙂
t
Hello 🙂 The
segment/assigned/count
is smaller than the
segment.loadQueue
k
Hi 🙂 . Hmm, if the
segment/assigned/count
is always smaller than
maxSegmentsInNodeLoadingQueue
, then we don't need to increase it further.
Do you see anything in the logs for the segments that are unavailable?
t
Yup, I saw some stuck segments
Which I did not see using
cost
k
Could you share the exact logs?
t
Copy code
coordinator.20221012.log:2022-10-12T14:16:51,586 ERROR [Coordinator-Exec--0] org.apache.druid.server.coordinator.duty.BalanceSegments - [_default_tier]: Balancer move segments queue has a segment stuck: {class=org.apache.druid.server.coordinator.duty.BalanceSegments, segment=xxxx_2021-11-19T00:00:00.000Z_2021-11-20T00:00:00.000Z_2021-11-30T17:20:23.435Z_6, server=DruidServerMetadata{name='druid-historical-xxx:8083', hostAndPort='druid-historical-xxx:8083', hostAndTlsPort='null', maxSize=6000000000000, tier='_default_tier', type=historical, priority=0}}
And btw, yesterday I’ve switched back to
cost
strategy and I’m not seeing those unavailable segments anymore
k
This seems to be a stuck balancing move, this should not cause segment unavailability.
And btw, yesterday I’ve switched back to
cost
strategy and I’m not seeing those unavailable segments anymore
Interesting.
cachingCost
does have some cost computation-based issues but it shouldn't really cause unavailability of any segment.
The load rules would also be printing logs like
Current replication: [...]
. Do you see logs like this for the segment in question?
t
Yes, I do
k
Oh, cool. What do those say?
t
INFO [Coordinator-Exec--0] org.apache.druid.server.coordinator.rules.LoadRule - Dropping segment
And I’m seeing a lot of them
It looks like a same segment was present on multiple historicals (no replication set for this datasource)