Hi, observed that increasing the zk client timeout...
# troubleshooting
e
Hi, observed that increasing the zk client timeout in the pinot zookeeper does not prevent a zk client timeout from helix, which is hardcoded. We see these errors when the brokers are under heavy gc pressure, gc pauses, etc.
Copy code
org.apache.helix.manager.zk
public class ZkClient extends org.apache.helix.manager.zk.zookeeper.ZkClient implements HelixZkClient {
...
public static final int DEFAULT_SESSION_TIMEOUT = 30 * 1000;
@Junkai Xue @Kishore G
Would it make sense to make this configurable?
I do agree w @Kishore G’s concern that the longer the timeout the longer an outage is not detected
m
Not to side track the conversation @Elon, but are you seeing 30s GCs?
e
Thanksfully we didn't but we did see a spike in gc activity/duration
Still investigating, could be a readiness probe failing
p
we have the same issue, we should have a configurable zk timeout
not some constant value
k
There is a jvm flag system property to configure it
p
I couldn't spot it in the code, if you have the variable we can give it a try
k
Copy code
public static final String FLAPPING_TIME_WINDOW = "helixmanager.flappingTimeWindow";

  // max disconnect count during the flapping time window to trigger HelixManager flapping handling
  public static final String MAX_DISCONNECT_THRESHOLD = "helixmanager.maxDisconnectThreshold";

  public static final String ZK_SESSION_TIMEOUT = "zk.session.timeout";

  public static final String ZK_CONNECTION_TIMEOUT = "zk.connection.timeout";