issue pinot.JPG
# troubleshooting
s
issue pinot.JPG
m
Check external view on status of these segments. Then check server logs for these segments
s
ok let me check
now the pods are in crashloopbackoff state ..mainly zookeeper pods not starting .. cc: @Mayank @Xiang Fu
now after 30 mins .. zookeeper pods are started and in ready state but broker pods are in crasjloopbackoff state
cc: @Shailesh Jha @Mohamed Kashifuddin @Mohamed Hussain
broker pods are running but not ready state
s
image.png
✔️ 1
s
above is recent broker logs
s
Brokers restarts with this error.
✔️ 1
m
@Shailesh Jha I've seen that error before when trying to connect to an inaccessible zookeeper node. Can you check if ZK can be reached?
s
Hi @Mark Needham Thanks for responding. Currently ZK pods also restarting
m
Seems like you might be under provisioned for the workload
s
you mean under provisioned for hard disk or ram or cpu?
m
cpu and mem
s
what situations can cause this error code 305 @Mayank @Xiang Fu @Jackie @Subbu Subramaniam @Neha Pawar if hard disk space is exhausted and hard disk crashes .. can the same error as mentioned in this post will come up . .I need to do RCA for the above error but cant see much error log on log files pushed to Loki
m
@Xiang Fu @Jackie @Mayank Is there any config in values.yaml through which we can enable HPA in pinot ? how to enable hpa in pinot kubernetes environment? https://github.com/apache/pinot/blob/master/kubernetes/helm/pinot/values.yaml Cc: @Sadim Nadeem @Shailesh Jha
x
You can add hpa template similar to external service
Please submit PR for it and I can help review your code changes
s
sure @Xiang Fu .. @Mohamed Kashifuddin please raise the PR also what is the root cause behind Error code 305 @Xiang Fu
please check this thread for more details on this error code 305 .. most probably hard disk exhausted but dont see concrete evidence backing it
x
what’s the current status, does zk keep restarting and have you checked zk disk usage?
s
no .. we have to do a fresh deployment using helm since zookeeper kept on restarting once we ran the below command while debugging:-
kubectl delete po prod-dev-pinot-zookeeper-0 prod-dev-pinot-zookeeper-1 prod-dev-pinot-zookeeper-2 prod-dev-pinot-broker-0 prod-dev-pinot-broker-1 prod-dev-pinot-broker-2 prod-dev-pinot-controller-0 prod-dev-pinot-controller-1 prod-dev-pinot-controller-2  prod-dev-pinot-server-0 prod-dev-pinot-server-1 prod-dev-pinot-server-2  -n  prod-dev-pinot
earlier we were facing the error code 305 that segments were missing in all of the tables
also we saw that pinot-server restarted after 32 days the same day the above error code 305 came ie 7th Jan 3 am IST
and we saw prometheus alert on 7th Jan around 3 am for disk-pressure
here you can see ..one of the pinot server pod restarted after 32 days automatically
one of the errors that we got after zookeeper started restarting after performing kubectl delete few hours after getting 305 error on pinot sql query editor that segments missing
please find below the prometheus alerts that we got at 3am IST on 7th jan before pinot server pod restarted automatically and we started getting error code 305 on pinot sql query editor and segments missing error .. Note: regarding disk .. we did saw that /sda1/etc were exhausted upto 80% for pinot-server-1 and pinot-server-2 pods on doing
kubectl -n dev-pinot exec dev-pinot-server-1  -- df -h
also after redeploying usign helm since zookeeper kept on crashing .. we were not able to restore the past historical data from persistent volume claims on google kubernetes engine
prometheus alert at 3 am IST 7th jan when pinot-server pod automatically restarted
@Xiang Fu
zk disk usage was below 20% and no such issue with zk disk exhaustion .. but not sure about pinot-server disk usage @Xiang Fu
x
ic then the issue could be server Disk out spsces
✔️ 1