what are some ways can we ensure High Availability...
# troubleshooting
s
what are some ways can we ensure High Availability / Disaster Recovery / Failover of my pinot cluster running on google kubernetes engine .. • Any advice will be very helpful to ensure almost 100% up time.. • We have alerts configured for CPU/ram/disk on slack .. • We have logging configured on loki • we have 3 broker , controller , server, zookeeper pods each with enough CPU, ram and disk as per load(incoming feeds frequency on kafka topics) .. • we have only Realtime tables Note: we encountered a crash after 32 days most probably because of hard disk exhausted and because of restarting zookeeper pods using kubectl delete while debugging.. cluster became unstable and attached persistent volume claims data got lost .. and we had data loss.. cc: @Mayank @Xiang Fu @Subbu Subramaniam @Jackie
a
@Sadim Nadeem did you get answer on above queries? We are also facing same issue. Today we also lost our two tables, and then complete data lost due pod restart.
s
Nope..
cc : @Mohamed Kashifuddin @Shailesh Jha
m
What crashed in your case, ZK?
x
make sure you have a persisted PVC e.g. AWS EBS
if you use local disk then zk pod restart will lost data for sure