We noticed that some brokers seem to have missing ...
# troubleshooting
e
We noticed that some brokers seem to have missing routing entries, ex. missing segments and/or servers. When the broker starts up, we see messages like:
Copy code
WARN [BaseInstanceSelector] [HelixTaskExecutor-message_handle_thread] Failed to find servers hosting segment: MYTABLE-1641542240329_2022-01-06_2022-01-06_6 for table: MYTABLE
When we issue the rebuild routing table api to the broker it just rebuilds an identical routing table with the same missing servers/segments. Is there any way to resolve that? We are also looking... could it be related to
ZkCacheBaseDataAccessor
not refreshing?
Would reducing the # of tables or total segment count (i.e. spread tables over many broker tenants) that a broker serves improve this?
m
Are those segments in IS/EV?
e
Yes
It just happened again:) Seems like after a broker oom's and is restarted it does not get a list of all servers and sometimes all segments, i.e. same broker request to all the other brokers produces a consistent result that is different from the restarted broker's result
But after restarting, sometimes 3x, the routing table is built and matches the segments in other brokers
Would removing the routing table and then rebuilding it help? i.e. using the api calls on the broker
j
i wonder if https://github.com/apache/pinot/issues/7578 is the same issue
i know that one say it’s around adding a new broker, but it seems you’re just running into this on startup
e
thanks @Johan Adami!
the issue we see it that the routing table exists in the broker but is incomplete, and rebuilding the routing table does not change the result. Only a restart (1-3x) does.
And we see those "Failed to find servers hosting segment" error (exact text in the first message)
j
is this the same table you were worried about the IS/EV being > 1mb? I wonder if somehow it’s not always getting back the full data from ZK. The broker isn’t supposed to go health until IS == EV for it, so it feels strange it would just be missing entries even after 1-2 restarts
e
No, this is a different table that has larger segments ~100 - 200mb per segment
j
ok last random hypothesis. do you have
-XX:+ExitOnOutOfMemoryError
set when you run the broker? And how certain are you that restarting it is actually restarting? we’ve seen some cases where we send SIGTERM and the broker just hangs especially when it’s in the middle of building the routing table.
e
Yes we have that option set and the broker does restart. I curl them all directly with a "select count(*)" or other query to verify they all return the same results
When they don't I compare the routing tables and find missing segments and/or servers
Even when we do deployments, sometimes brokers start up with incomplete routing tables (as compared with other brokers). Would reducing the # of tables and segments by creating broker tenants help with this?
@Johan Adami @Mayank whenever you have a chance. Really appreciate the help and I will continue to debug, I'll share what I find also.
m
@Jackie for any inputs^^
🙏 1
j
@Elon Can you please check if the servers are in good state? If the segment is showing
ONLINE
in EV, check the online servers and see if they are in
liveInstances
and if they if queries disabled in
Helix InstanceConfig
e
Yes, we checked that - ideal state == EV and all segments online, exactly 3 replicas. The servers are all online and in live instances and none are disabled (checked that as well).
I manually curl every broker and only 1 returns with the missing data.
Then on restart (or 2-3) it's fine.
j
Does the bad broker ever disconnected from ZK?
Also, which version of Pinot are you running?
e
We see zk timeouts in the logs but only on zk side - but then we restart
we're still on 0.8.0
Was this fixed or improved in 0.9.2/0.9.3?
I don't see messages from broker about zk timeouts, but they show up in zk logs.
j
Checked the code and seems no fix since
0.8.0
👍 1
e
Thanks!
j
Do you see some explanation after the warning log?
e
No explanation, just timeout
Could this be related to
ZkCacheBaseDataAccessor
code?
j
I mean following this line:
Failed to find servers hosting segment: MYTABLE-1641542240329_2022-01-06_2022-01-06_6 for table: MYTABLE
Here is the log I find in the code:
Copy code
LOGGER.warn(
          "Failed to find servers hosting segment: {} for table: {} (all ONLINE/CONSUMING instances: {} and OFFLINE "
              + "instances: {} are disabled, counting segment as unavailable)", segment, _tableNameWithType,
          onlineInstancesForSegment, offlineInstancesForSegment);
e
Yep, that's the message!
could it be some issue with zk? like an inconsistency? Is that even possible?
j
You mean no explanation following the log in the parentheses?
e
Checking , might have to dig the log messages up, will get back to you shortly
it's the online servers and segments:
Copy code
WARN [BaseInstanceSelector] [HelixTaskExecutor-message_handle_thread] Failed to find servers hosting segment: MYTABLE-1642556066145_2021-12-03_2021-12-03_4 for table: MYTABLE (all ONLINE/CONSUMING instances: [Server_pinot-us-central1-server-zonal-41.pinot-us-central1-server-headless.pinot.svc.cluster.local_8098, Server_pinot-us-central1-server-zonal-42.pinot-us-central1-server-headless.pinot.svc.cluster.local_8098, Server_pinot-us-central1-server-zonal-43.pinot-us-central1-server-headless.pinot.svc.cluster.local_8098] and OFFLINE instances
But those segments were available on all other brokers and the segments themselves were on live server instances - EV == IS.
There appears to be correlation between gc's (count + time) right before and during the broker restart, this could be why there are missed messages - but they do not seem to be correlated with zk session timeouts
j
It is also possible that broker is somehow running into GC's and not able to process the callbacks
💡 1
e
Strange - it's on startup (the gc's and incomplete routing tables) - will see if I can find anything and will share...
c
if the broker is unable to process callbacks won't we see this on the Helix metrics ?