Hello everyone :wave: I’m currently running the Pi...
# troubleshooting
l
Hello everyone 👋 I’m currently running the Pinot Recommendation engine, and have a question regarding the output I’m seeing. One of the kafka topics we have has 100,000 messages/second, and Pinot is recommending 400 partitions for this topic.. I can see in the code it simply takes the number of messages and divides it by 250. Is 400 partitions really recommended here? Or should I disregard it.. Not sure what the impact is either way
This is the output..
Untitled.txt
m
how many partitions does the Kafka topic use at the moment?
as far as I understand Pinot will create one consumer (i.e. one thread) per Kafka partition and will import from those partitions concurrently. So I think it's mostly for increasing ingestion throughput. I'm not sure how many messages / second you should have per partition though. @Mayank would be able to give a better answer for that
l
The recommender seems to indicate 250 messages per partition..
m
400 sounds a lot though, - that'd be as you say 250 messages / second / partiton
seems low
l
I just don’t know what an acceptable messages/s is
If 2500/s is fine I could go with 40 partitions
Also another question, is it OK to change the partitioning of a table?
not sure if it stores partitioning per segment or not
If I add more partitions, I’ll need to update rmy column partition map
m
@Lars-Kristian Svenøy you can start with 32 partitions as well. How many Pinot servers and what’s their cpu mem config
l
I’ve currently got 16 partitions, 16 servers, with 2 CPU and 16 GB memory each
Unfortunately at the moment I’ve got uneven load on my partitions
so it looks like I also need to change the key I’m partitioning on
but then this key will no longer be usable for partition based pruning
m
16 might also work. Side note, you may want to reduce the number of cores and increase their cpu/mem, will reduce query fanout
m
@Mayank can you change the partitions once it's all configured with Pinot
or would that mess things up?
m
Functional everything works
l
Thanks guys, this is very helpful.
Now I’m curious though, because the provisioning helper tells me that if I keep my partitions as 16, my segments will be absolutely huge
m
Changing the partitions count works still works. Changing the partition function or column will break partitioning for older segments and Pinot will think it is not partitioned
l
1 hour segments = 2.26G
m
depending how you configure the segment flush threshold, Pinot will work out how many messages need to go in each segment to keep it a consistent size
l
with 16 partitions
m
https://docs.pinot.apache.org/operators/operating-pinot/tuning/realtime#fine-tuning-the-segment-commit-protocol If you scroll to the bottom of here:
Copy code
"realtime.segment.flush.threshold.rows": "0"
"realtime.segment.flush.threshold.time": "24h"
"realtime.segment.flush.threshold.segment.size": "450M"
if you set threshold.rows to be 0
it will work it out for you
l
yep that’s what I’m doing
m
and they're really big?
m
ah I see. With that ever time rate you are generating 2G per hour?
l
Per partition yes
m
Is the schema too fat?
l
no it’s fairly lightweight
but it’s 100,000 events per second
I do have a key I could reliably partition on, but it isn’t often queried for.
But I could then theoretically up my partitions to let’s say 200 for the topic
But I’m not sure if that’s OK or not
With 200 partitions, I could get my segments down to 370 MB
I just don’t know if that is madness or not
This table currently has 10,579 segments, and a total of 399506097497 documents
m
You can use managed offline flow to reduce number of segments.
l
I am using the realtime to offline segments task
m
400B records? Nice!
l
yeah 🙂
m
Yeah, then you can go upto 1-2G of those segments size in offline
l
And it’s going to get heavier
m
Quite impressive!
l
Sounds good thank you, I’ll give it a shot
One more question
Is it fine to have different partitioning for realtime vs offline?
s
@Lars-Kristian Svenøy if you can copy-paste the provisioning helper's output that will help a lot.
l
I pasted it in the top of this conversation @Subbu Subramaniam
Thank you all 👍
s
Ah, missed it, thanks
I think you have the answers you needed, but just to add, the recommendation engine assumes that you need all the segments in memory in the realtime host.
l
Yep figured that earlier. Would be nice to have a setting to let it know a deepStore is being used
s
Please do file an issue. thanks
l