Hi, I noticed some strange behavior when setting `...
# troubleshooting
t
Hi, I noticed some strange behavior when setting
realtime.segment.flush.threshold.rows
for my realtime tables. It seems that the actual number of rows per segment becomes some value smaller than the value I set. For example, I'll set this to 1000000, but in the segment metadata,
segment.flush.threshold.size
would be 500000 and the segment does only ingest 500000 rows. This seems to only happen for some tables, and sometimes it is shrunk by a factor 2 or 4. Just wondering if there is any other setting I'm missing that is causing this?
m
Can you check if it gets divided by number of partitions?
t
Oh that might be it. In one example, I have about 20 partitions and it gets divided by 4. Is this expected?
m
Yes if you have each server consuming 4 partitions
If the doc is not explicit enough we should fix that cc: @Mark Needham
t
I see, does the number include replicas as well?
I couldn't find anything in the docs explaining this
m
It boils down to how many partitions a server has to consume (function of how many partitions there are, how many servers, and what’s the replication)
t
got it, thanks!
Does this also apply to setting
realtime.segment.flush.threshold.segment.size
? it looks like when I use that, alongside
realtime.segment.flush.autotune.initialRows
, it actually makes each segment that size.
m
No, what I mentioned above only applies to
realtime.segment.flush.threshold.rows
👍 1
m
have added an explanation of what Mayank said above to https://docs.pinot.apache.org/configuration-reference/table#indexing-config
🙏 2