This message was deleted.
# general
s
This message was deleted.
x
It’s not required to run the prepare scripts on EKS worker node but the host where you run
helm
or
kubectl
.
a
@Xiaogang Wen Thanks. I did run the prepare scripts on the host this time but I'm still getting the same error this time as well.
y
turn off zookeeper tls and try again. NIO doesn’t support TLS. when you turn on zookeeper tls, the init script using NIO failed to connect to zk.
a
@Yu Wei Sung Zookeeper TLS is not enabled. I'm getting this error with TLS not enabled. What is NIO?
y
check the pod status first.
What’s the zk pod/svc status?
a
Zookeeper pod is up and ready. But I see the above posted error in the ZK logs. Bookie, proxy and broker pods are stuck on init.
y
what’s the zk service?
a
You mean the type of service?
It's ClusterIP
y
port?
what endpoint slice?
a
8000/TCP,2888/TCP,3888/TCP,2181/TCP
y
8000 port? on zk?
how many containers in the pod?
a
Just the 1 container
y
only 1 endpoint… is this single node?
a
Nope 4 nodes
But Zookeeper has only 1 replica, I don't know why. I didn't change anything in the values.yaml
y
if you check back, the log shows “WARN”. this might not be the error
a
All other Pulsar containers are waiting to start though
Pulsar-bookkeeper-verify-clusterid container inside bookie pods has these logs
y
check the zk log. the job is waiting for zk to become ready.
a
But Zookeeper pod is ready. As can be seen here.
y
bookie init is failing. you need to check bookie init
the order is zk -> bk -> broker. zk is ok. the. bk is stuck.
a
I posted the logs of bk init a bit above
I can't find anything in this Slack or google when I search the log output
y
that’s bookkeeper pod. there should be a job
pulsar-bookie-init?
a
That's the logs of init container inside the bookie pod
Ah yes I see that job too and I see that it's not completed
y
before bookie statefulset kick in, there is a k8s job. that job will setup zk znode, so bookie sts will connect to zk.
that’s missing
a
I can't access the logs of that job now though. Because the job pod is no longer up.
y
then your bookie won’t be ready and broker won’t be up. try uninstall/clean up and try again.
a
Ok let me try that and this time I'll try to capture the logs of that job
Uninstalled, deleted PVC and then installed Pulsar again. I was able to capture logs of that job.
y
job finished. the znode created. bookie should be up. check the status of init-container from bookie sts
a
Bookie 0 and 2 are running and ready. Bookie 1 is stuck at init.
Bookie 3 is stuck at init too
I do see this error in the init container in the bookie pod this time
Broker pod also started running and then crashing. These are broker logs
y
function worker on broker?
what’s the broker pod resource limits? memory/cpu?
the broker looks good but failed at function worker….
a
No limits, only requests are set. cpu: 200m and memory: 512Mi
What is function worker?
y
function workers let you run pulsar functions. I “guess” the broker pod has function worker setup but the memory is too small to run broker and function workers in the same pod.
a
@Yu Wei Sung But there are no resource limits set. Do you think the node doesn't have enough memory?