This message was deleted.
# general
s
This message was deleted.
j
Hi Younes, There is partition "boosting" which introduces a high cardinality pseudo-column to the range key set to even out segment distribution. This PR: https://github.com/apache/druid/pull/15474#discussion_r1448388430 indicates it is being added to Group By queries, which suggests it should already be in a (more recent?) release of the product. I don't have information on the base "clustering" behavior, other than I am aware that MSQ should be adding something to the range key to increase cardinatlity in an effort to create more even resulting segment sizes.
y
I'm using the latest version 28.0.1. I'll check to see if this feature is included...
l
It is expected that a partition is processed by a single worker (otherwise we’d be “partitioning” it further). What is not expected is that the single partition is too large, that it is a pain point for the user. Since it’s an issue with the segment generation, it is highly likely that what John mentioned is the issue, since we haven’t seen inaccuracy in other regions of generating partitions. To see if it’ll help to upgrade to 29.0.0 (which is going to be released soon) - check if the query is a Group By query, and also has a Clustered by
What’s the size of the large partition? That’ll also result in large sized segments, which is also not optimal.
Adding more items in your clustering spec would help (if your usecase allows)
y
Thank you, I'll try that
Ah, that one is for 29.0
I have 8 partitions, 348M rows, 218M rows in 1 partition
l
Umm, that’s really bad. That’ll mean that the segment will have 218M rows
I am guessing what really happened is what John and I have described. Dataset with skewed values for the clustered by dimension. It also means that rest 100mn rows went into other 7 partitions, which is also unoptimal. I’d recommend either upgrading to 29.0.0 and reingesting, when it comes out OR adding more dimensions to the clustered by, so that the partition sizes become lower.
y
I added a high cardinality column to the clustering. And that worked nicely!!!!
Thankyou 🙂
🙌 1
k
@Younes Naguib You should try the previous query which generated large segments with druid 29.0.0 and see if it changes something. It would be a good validation for the druid community. Also druid 29.0.0 will be out this week.
👍 1