This message was deleted.
# general
s
This message was deleted.
s
Have you enabled metrics? You can see how many segments each of them scanned in the query metrics. Specifically in broker's
query/segments/count
, you can also see specific segment processing time in Historical's
query/segmentAndCache/time
Also, how do you expect to query this? Hash is typically better to minimize segments needed when using an equality condition at query time on the partitioning column. It is more susceptible to skew in the data and can therefore create more imbalance in the segment sizes. With range partitioning, particularly in MSQ, segment size balancing is better, but you do have the possibility of needing to scan more segments in the equality condition case because a value can span more than one partition. I'm very curious to hear more about your results.
a
Metrics is a good way to do that comparison. I don't think that scanning more segments is necessarily a problem. Even if the request lands on a data server, data server has on-disk indices to skip most of the rows.
a
Scanning more segments has become an issue, that's why I am considering range dim partitioning. Right now cpu reaches 100% in historicals, querying last 7days of data replication =2. There are 900 segments in a day near 500-700 MB.
Currently with hashed based partitioning there is no pruning happening, which is why cpu is reaching upto 100%.
s
Any common filtering columns among the queries? In the hash case did you filter on the partitioning column? I know the obvious answer is yes. But just verifying 😁 it seems very odd to me that it wouldn’t prune. Is the filter an equality?
a
Yes we had equality condition and filtering on partitioning column.
@Sergio Ferragut
query/segments/count
is not enabled by default. Do we need to do code changes for that?