We noticed that some brokers are missing routing t...
# troubleshooting
e
We noticed that some brokers are missing routing table entries for segments. This results in different query results depending on which broker is used. After restarting the broker multiple times it picks up the segments (which are not newly added - a lot of them are 2+ days old). Here is the error message:
Copy code
2022/01/07 11:39:58.204 WARN [BaseInstanceSelector] [HelixTaskExecutor-message_handle_thread] Failed to find servers hosting segment: MYSECRETTABLE-1641542240329_2022-01-06_2022-01-06_6 for table: MYSECRETTABLE
Is there any way to reduce/eliminate these errors? We're looking at zk client configuration in helix...
➕ 1
The errors are non-deterministic, we're seeing if it correlates with zookeeper latency.
Also, this table exists with no advanced routing or partitioning - would it help if we use
replicaGroup
instance selector and partitioned replica group segment assignment and move the table to a custom tenant?
m
There’s an api in swagger to force rebuild the routing resource for a table that should update the routing table.
Seems like the broker missed the notification of new segment added. This would need debugging around the time the segment was created and looking at helix/zk/broker related logs on what communication was going on.
e
I see increased latency in zookeeper around the time of the missed messages
Greater than 30 seconds, which seems to be the zk clients timeout in helix 0.9.8 - we're still using pinot-0.8.0 and zk 3.5.5 - are there any fixes in later releases (we will be upgrading to 0.9.2 soon)? @Mayank
m
No fixes specific to this problem that I am aware of @Elon
👍 1
e
Ok, I noticed that zk network io increased, we upload segments constantly (and have not yet switched to metadata push - we use regular upload segment api). Could that contribute to zk message size
Either way, I will share my findings here, in case there can be some improvements... which I would be happy to contribute
@Mayank does that sound likely (upload segment frequency -> zk network io)?
m
Is the message size > 1MB?
e
there were some that were, and we had to increase the jute.maxbuffer size in zk client for the controller
do you recommend that we set that to the same value on the zk server?
m
I am trying to recollect if there’s a Helix hard limit cannot be increased. If so, and if there were notifications that were indeed part of message batch that got dropped, that could explain the whole thing.
e
oh, nice!
let me know if there's anything I can try here to verify that
I do know for sure that increasing jute.maxbuffer on the controllers reduced network io by a lot
zk network io ^^
We have many tables that use the default segment assignment/routing - would moving those tables to the partitioned replica group or replica group assignment strategy help reduce zk message size?
d
@Elon fyi there is a hard limit here in helix 0.9.8 (what pinot is currently using): https://github.com/apache/helix/blob/e275c54220423fcbb54f6a3596763fe6dd74e129/helix-core/src/main/java/org/apache/helix/ZNRecord.java#L64. helix has been patched to make the limit configurable via jute.maxbuffer in 0.9.10. i also think jute.maxbuffer needs to be set on both client and server for it to work correctly per https://zookeeper.apache.org/doc/r3.4.8/zookeeperAdmin.html
If this option is changed, the system property must be set on all servers and clients otherwise problems will arise.
e
Thanks so much @David Yang! We are still using pinot-0.8.0/helix-0.9.8 - we noticed that increasing
jute.maxbuffer
on the controller eliminated the "java.io.IOException: Packet len XXX" errors and zk cpu stopped spiking. Would you recommend we set it on zookeeper or it would have no effect since
ZNRecord
has a max size?
If the size has a 1mb limit how can we get those packet len errors? Is it because of child records?
d
per the zk guide, i think setting it on the server should be done regardless. doing this without the updated helix version will eliminate some class of issues but not all. we were specifically trying to avoid issues with large idealstate, so i can’t comment on the specific broker issue you saw. with jute.maxbuffer increased but without helix updated, we were able to recreate issues writing idealstate >1MB via pushing new segments.
m
Thanks @David Yang for sharing your insights.
e
thanks @David Yang @Mayank! really appreciate the help
what we are seeing now: segment lineage for a table is blowing up from 92kb yesterday to 211kb today - would that cause timeouts for
endReplaceSegments
and increase zk latency?