Is there any issue with having a large number of t...
# troubleshooting
e
Is there any issue with having a large number of tables on a single cluster in terms of zk latency, metadata size, etc.? We currently have ~75 tables.
s
we have 4000+ tables.
e
Thanks:) Could it be constant uploading of offline segments that increases zk latency - we've only seen this issue in the past month.
s
You can choose to split the pinot and helix controllers. It has helped us. You should increase the network speed between zk and controllers as well.
e
you mean use split commit? or is this something else?
let us know if there's a doc or example, I'd be happy to try it out, sounds great:)
@Subbu Subramaniam - is this the doc we can follow to separate the helix and pinot controllers? https://cwiki.apache.org/confluence/display/PINOT/Controller+Separation+between+Helix+and+Pinot
✔️ 1
s
@Jack do we have a doc on how to separate the pinot and helix controller hardware ? We do have the steps in the design doc, but much of that is verified (when we did it, we had to make sure that we can also come back out of the split mode. You don't need to worry about it, since we have tested it for several years now).
🙌 1
e
thanks so much @Subbu Subramaniam!
j
Right, that’s the correct doc. 👆 Feel free to try it out! :)
e
ok:) Thanks @Jack!
k
I dont think splitting is needed at 75 tables
e
thanks
k
Elon, are you using metadata uri push or data push?
e
Still using data push and for most tables we're still not using peer download (just started to roll it out)
We are working on migrating to metadata push, would that reduce zk network io?
k
do you really know its ZK network? do you have graphs?
e
Yes - for the past 2 months the zk network io has been steadily increasing
also have graphs for zk latency which seems to have started increasing about 2 weeks ago
In spikes
we saw messages like "io exception packet len too large" in the controller - after increasing the
jute.maxbuffer
those messages went away and the zk network io decreased a little bit
Just increased on the brokers
Is it dangerous to increase
jute.maxbuffer
on the servers?
k
Then you might be having too many small segments issue
🙌 1
e
Nice, that sounds like it's the issue! We have many realtime tables that are very small segments - we use daily push frequency, should we also have max time set to 24 hours?
If we move all of those tiny tables to their own tenant would that help?
they are realtime only tables
a
how many segments is too many?
e
and is it the amount of segments or the size?
s
I was thinking overnight and arrived at the same conclusion that you shoul dnot need to split at 75. We split when were at 2000 or so. Sounds like you have identified the problem to be that of small segments. Check why you have such small segments. Is it tht your inbound stream is too slow? Can you reduce the number of partitions? Or, can you change the time to 48 hours so that larger segmengs can be made? I hope you are using the segment size option in the realtime config
e
Thanks!
We noticed that we have segment lineage nodes that are 80-90kb and we see timeouts for
endReplaceSegments
- which is 10 mins. Would many concurrent uploads to the same table with a large lineage zk node cause zk latency to increase?
s
oh, you are using merge/rollup to combine segments? We use it in our 4000-table cluster, and have not seen any problems so far. @Seunghyun any comments?
e
nice:) We're increasing zk session timeouts to see if that helps
what we see is that some segments are out of sync with deep store, maybe a missed zk message? and it seems this is the likely cause of the failure for
endReplaceSegments
- does that sound likely?
s
e did have problems with timeouts earlier, and we set it to a larger value. I will let @Seunghyun help u here
e
thanks @Subbu Subramaniam!
s
@Elon If your ZK is too slow and make it work the
endReplaceSegments
, there’s a config to set the timeout for
endReplacement
on the minion side
😮 1
e
nice! I can take a look but would you know off the top of your head which code I can look at?
s
Please refer
MinionConf
. There’s the config called
pinot.minion.endReplaceSegments.timeoutMs
🙌 1
You need to set it as part of the minion starter config
e
Thanks! Really appreciate the help, this is the best community:) I promise to make some contributions back:)
s
@Elon which issue can i assign to you ? 🙂
😁 1
e
any that you want! lmk I would be super happy to help in any way I can:)
s
By curiosity, are you trying out merge/rollup?
e
we are but still using pinot-0.8.0
s
FYI, we don’t support realtime tables yet
e
I thought offline
s
it will work with offline.
e
ah,you mean it converts to offline, got it
right now most inserts are from trino - I use the code from the library - and am working on metadata push instead of direct segment push. I have a pr for the trino connector which I will be updating.
last question: we constantly are uploading offline segments, to refresh, kind of like an "upsert" but we replace segments. Is this not a good way to use pinot?
sorry, one more: would switching to partitioned replica group segment assignment and replica group routing be recommended to reduce load on zk?
a
it should be a different discussion, but I think Pinot should fail before is silently get's into inconsistent state
should we make those exceptions more critical? Hard stop and restart critical?
e
inconsistent meaning: segments in gcs not downloaded to servers, or replicas have different versions of a segment (different crc), or replication factor < desired for a certain amount of time.