This message was deleted.
# troubleshooting
s
This message was deleted.
g
generally these operations are limited by metadata store performance: scaling up your metadata store generally helps
like, using a server type with more CPUs/memory
we're looking into improvements in this area that will limit the load on metadata store as the number of allocations increases. hope to have something to ship here in the next release or two!
👍 1
btw, i'm wondering, why not use hour segment granularity? i've seen that as the most common segment granularity for streaming data
k
@Nick M, thanks for reporting the issue. We have seen the issue with the allocation of pending segments on some other clusters as well. We are actively working to resolve this issue, hopefully in the upcoming Druid release.
overlord logs are continuously allocating pending segments
I don't think a 15min granularity would help with the segment allocation slowness. For your cluster, I agree with Gian, you should try going with an HOUR granularity. You also mention that with a granularity higher than 15 min, you end up with too many segments per time chunk (after compaction). You could try increasing the number of rows per segment specified in the auto-compaction config, and use
range
partitioning on 3 to 4 dimensions for better data distribution. Could you please also elaborate on how the "coordinator slows down"? If it's segment loading and datasource availability that takes time, you can try increasing the load queue size through the coordinator dynamic config. Hope this helps 🙂
👍 1
n
generally these operations are limited by metadata store performance: scaling up your metadata store generally helps
We're running the bitnami postgresql-ha deployment for metadata. It's a cluster of 3 (1 master and 2 replicas where read requests are load balanced and write requests are forced to the master). CPU and memory does not seem to be close to exhausted according to the metrics. And it's backed by some decent disks and the latency doesn't seem to be an issue either. What are others using for metadata? Does mysql offer better performance?
btw, i'm wondering, why not use hour segment granularity? i've seen that as the most common segment granularity for streaming data
We found that when using hour segment granularity, we ran into issue with the coordinator cycles taking too long due to the number of segments in the timechunk (post compaction we have upwards of 3000 segments at an hourly granularity). I think we're hitting this https://github.com/apache/druid/issues/11700
Could you please also elaborate on how the "coordinator slows down"? If it's segment loading and datasource availability that takes time, you can try increasing the load queue size through the coordinator dynamic config.
The coordinator cycle takes a while due to the VersionedIntervalTimeline issue above (or was on 0.23 when we decided to move to 15 minute granularity). The cycles took so long that eventually it stopped handing off segments during ingest with the default completion timeout
k
After compaction, what is the typical size of each of your segments in terms of rows and bytes?
Reducing the granularity from hour to 15 min will only increase the number of total segments. You can also look at the
coordinator/time
metric to identify the problematic coordinator duty. The timeline build happens on a separate thread, so it shouldn't affect the coordinator cycles, except maybe through the
MarkAsUnusedOvershadowedSegments
duty, we can be sure once we look at the above metric. If it does turn out to be the
MarkAsUnused
duty that is taking most of the time, you can temporarily disable it by setting
millisToWaitBeforeDeleting
in the coordinator dynamic config to a very high value. (a fix for this duty has already been merged https://github.com/apache/druid/pull/13287 and will be released in Druid 25).
n
After compaction, what is the typical size of each of your segments in terms of rows and bytes?
Around 5 million rows with each row around 200 bytes so each segment is ~1GB in size
The timeline build happens on a separate thread, so it shouldn't affect the coordinator cycles, except maybe through the
MarkAsUnusedOvershadowedSegments
duty, we can be sure once we look at the above metric. If it does turn out to be the
MarkAsUnused
duty that is taking most of the time, you can temporarily disable it by setting
millisToWaitBeforeDeleting
in the coordinator dynamic config to a very high value. (a fix for this duty has already been merged https://github.com/apache/druid/pull/13287 and will be released in Druid 25).
This duty is indeed the one that takes the most time in the coordinator duty cycle. Looking forward to Druid 25 as that will fix one of the biggest issues we've seen with druid. Does that patch change the acceptable number of segments per timechunk (2000 as per https://github.com/apache/druid/issues/11700#issuecomment-918374527)?
k
The patch significantly reduces the number of segments that the duty has to check for marking as unused. So the runtime of the duty also goes down.
Around 5 million rows with each row around 200 bytes so each segment is ~1GB in size
So you are getting around 1TB data per hour? Since your segments already seem big enough, you cannot increase the segment size further to reduce the total number of segments. Instead, try enabling roll-up if you are not already using it.
👍 1
n
So you are getting around 1TB data per hour?
Since your segments already seem big enough, you cannot increase the segment size further to reduce the total number of segments. Instead, try enabling roll-up if you are not already using it.
Yeah, around 1TB / hour at the moment but that will increase when we can stabilise things at this speed. Roll-up is on the horizon but in this early stage of deployment is not yet in place and won't be for a while I don't think
s
@Nick M - you may want to look at your pending segments table and see if it needs to be cleaned up.
n
It’s configured to auto cleanup but at a 15 minute segment granularity will grow to be a little under a million records in a day or so.
s
oh yeah, that’s a lot.
I think the auto cleanup looks at the time at which the earliest task was submitted for a datasource. If you happen to have any zombie tasks or kill tasks for a datasource that are blocked, auto cleanup may not be cleaning up. Alternatively, you can increase the frequency of cleanup.
You may also want to look at creating indices in your metadata store to improve query perf.
g
an FYI about allocating segments— there's another patch from @kfaraz that helps here: https://github.com/apache/druid/pull/13369 we've been running it on a cluster with nearly 1000 tasks for the last couple of days, & 98%ile action runtimes are down from 3 minutes to 13 seconds. Maximum (over this time period) is 15s; we had seen over 5 minutes prior to rolling out this patch.
👍 1
he's really on fire recently with the scalability patches! 🧑‍🚀 🚀
😄 1
👍 1
there's been three others related to kafka supervisor performance that went up recently: • https://github.com/apache/druid/pull/13334 ⬅️ more aggressive deduping of work in the main loop • https://github.com/apache/druid/pull/13328 ⬅️ fewer metadata store calls per run in the main loop • https://github.com/apache/druid/pull/13354 ⬅️ async comms from supervisor -> task; contact all tasks at once and don't block worker threads
we're targeting 25.0 for all of these and for the Coordinator patch he mentioned a couple weeks ago (https://github.com/apache/druid/pull/13287)
n
These all look very promising. When is 25.0 due for release?
k
The release is scheduled for the mid of December. The release branch has already been cut, backports are in progress. We should have a release candidate by next week.
🙏 1