Hi Team. We are planning to use a deployment of P...
# general
p
Hi Team. We are planning to use a deployment of Pinot in Production with the following configurations:
Copy code
Brokers: 3 Replicas
Controllers: 3 Replicas
Servers: 5 Replicas
Zookeepers: 3 Replicas
However, we are facing an issue with our controller deployment. When we run queries through presto, we are not able to get the result consistently. Many times the query returns the following error:
Unexpected response status: 500 for request  to url <http://pinot-controller.dataplatform.svc.cluster.local:9000/tables/><table-name>/instances, with headers {Accept=[application/json]}, full response {"code":500,"error":"Failed to get full list of /pinot/CONFIGS/PARTICIPANT"}
We did an analysis of the error and found out that this is happening because of only specific controllers being able to fetch
/pinot/CONFIGS/PARTICIPANT
for a specific table. For example if we fire a CURL request to controller replica 1 for Table A and get a response 200, we get an error for the other two controller. On further analysis, we found that we only get a 200 Response, IF the load balancer directs the query for a table to its leader controller. We were hoping for help on this issue. Temporary solutions would be to reduce to a single controller which is not recommended for production deployments. Communities help here would be greatly appreciated!
m
All controllers should be able to see the exact same cluster state, so I am wondering if this is a setup issue. Can you check the Pinot UI from different controllers and compare what you see in the ZK browser
p
Thanks for the reply. Yes, for some reason, only one controller is able to load the ZK browser. Only the leader is able to respond to queries while the other two are causing issues. We tried restarting the cluster and it worked for around 2-3 minutes before the connection broke again. Looking into what went wrong in the deployment process. Do let me know if you have any ideas where should we look first.
On seeing the logs, controller 0 and 2 have this common error log:
Copy code
java.io.IOException: Packet len4675551 is out of range!
	at org.apache.zookeeper.ClientCnxnSocket.readLength(ClientCnxnSocket.java:122) ~[pinot-all-0.10.0-SNAPSHOT-jar-with-dependencies.jar:0.10.0-SNAPSHOT-078c711d35769be2dc4e4b7e235e06744cf0bba7]
	at org.apache.zookeeper.ClientCnxnSocketNIO.doIO(ClientCnxnSocketNIO.java:86) ~[pinot-all-0.10.0-SNAPSHOT-jar-with-dependencies.jar:0.10.0-SNAPSHOT-078c711d35769be2dc4e4b7e235e06744cf0bba7]
	at org.apache.zookeeper.ClientCnxnSocketNIO.doTransport(ClientCnxnSocketNIO.java:363) ~[pinot-all-0.10.0-SNAPSHOT-jar-with-dependencies.jar:0.10.0-SNAPSHOT-078c711d35769be2dc4e4b7e235e06744cf0bba7]
	at org.apache.zookeeper.ClientCnxn$SendThread.run(ClientCnxn.java:1223) [pinot-all-0.10.0-SNAPSHOT-jar-with-dependencies.jar:0.10.0-SNAPSHOT-078c711d35769be2dc4e4b7e235e06744cf0bba7]
Controller 1 seems fine from the logs.
m
How many tables / segments do you have?
p
We have 54 tables (27 RT and 27 Offline). We will be increasing this soon.
Will look into this and come back to you thanks!
m
I think it might be segments within table (Znode size).
👍 1
p
Hi! We were able to resolve this issue. For the benefit for others that might face this issue, apply the following JVMFLAG/JAVA_OPTS to both the zookeeper and controller deployment. You can set the value depending on your error message. For our case, we increase to 50 MB.
Copy code
-Djute.maxbuffer=<value>