Another question, do replicas of segments matter o...
# general
y
Another question, do replicas of segments matter on performance? Does it increase parallelism?
j
it should, if the query needs more parallelism. However if you have far more segments to read than you have Historicals in the cluster (which is usually the case), then I don't think it affects performance noticeably.
y
I am trying to maximize performance the windowing query for large data. I do not see spikes on cpu on historical. I followed basic cluster tuning.
I saw the query segment time, how to reduce it?
j
Afaik Window Functions run primarily on the Broker, they are post-aggregate. but if you are seeing long segment scan times, can you tell us how many segments are being scanned for the query, and what the Avg and P95/P98 scan times are? And if you are using streaming ingestion, also if you can split out node type of Peon vs Historical.
y
# of segments are 30. let me look for P95/P98 scan times btw, thank you for your input.
Can you tell me what is the name of the metric for P95/P98 scan time?
query segment time is about 8k ms
j
I think avg and P95/P98 are both derived form the segment scan time metric ... the client tool must use those. So your segment scan times are around 8 sec? That's pretty slow. How big are the segments (rows, GB)?
y
700k rows and size around 400MB each.
It really depends on query filter condition. I also used ranged partition.
j
how many processors on your Historicals? With only 30 segments, range partitioned, there is a chance that you are query pruning and opening up only a few segments, which won't have much parallelism. In which case I would think the make the segments smaller to reduce scan times.
y
24 core cpu for historical and 3 of them.
So are you saying I am only utilizing 30 core of it?
I see, since it is ranged, I probably utilize less than 30 segment. Segment that falls into filter condition will be utilized. Due to data distribution, I think ranged partition is not a good choice. I think I should use hash partitioning.
j
What is your segmentGranularity? and how many segments per time interval?
y
I changed to hash to distribute data equally. Performance is improved in GroupBy query.