This message was deleted.
# general
s
This message was deleted.
y
1. turn off gclog? 2. redirect gclog location under work_dir?
m
Is it possible to control the gclog from the chart values?
"-Xloggc/pulsar/gc.log"
try with this param and replace with working dir as Yu Wei Sung mentioned and see if you are able to move forward
or you can comment out the gc logging under the same values yaml
m
Thank you! this fixed my issue! But now i’m facing a new issue, the job was completed successfully and when the bookies are trying to start i’m seeing these error logs:
pulsar-bookie-1 pulsar-bookie 2023-04-19T07:41:05,878+0000 [main] ERROR org.apache.bookkeeper.bookie.LegacyCookieValidation - There are directories without a cookie, and this is neither a new environment, nor is storage expansion enabled. Empty directories are [/pulsar/data/bookkeeper/journal/current, /pulsar/data/bookkeeper/ledgers/current]
pulsar-bookie-1 pulsar-bookie 2023-04-19T07:41:05,883+0000 [main] ERROR org.apache.bookkeeper.server.Main - Failed to build bookie server
pulsar-bookie-1 pulsar-bookie org.apache.bookkeeper.bookie.BookieException$InvalidCookieException:
pulsar-bookie-1 pulsar-bookie 	at org.apache.bookkeeper.bookie.LegacyCookieValidation.checkCookies(LegacyCookieValidation.java:116) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
pulsar-bookie-1 pulsar-bookie 	at org.apache.bookkeeper.server.Main.buildBookieServer(Main.java:422) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
pulsar-bookie-1 pulsar-bookie 	at org.apache.bookkeeper.server.Main.doMain(Main.java:272) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
pulsar-bookie-1 pulsar-bookie 	at org.apache.bookkeeper.server.Main.main(Main.java:255) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
BTW: I deleted the old pv and pvc’s and new volumes were created and I still encounter this issue
r
run the decommission bookie command, that will remove the entries in the zookeeper
m
from which pod? because I can’t run it from any of the bookie pods
OK, from one of the zookeeper pod as I understand..
now I see this error after I ran the decommission:
2023-04-19T08:01:07,798+0000 [main] ERROR org.apache.bookkeeper.tools.cli.commands.bookies.DecommissionCommand - Received exception in DecommissionBookieCmd
org.apache.bookkeeper.replication.ReplicationException$UnavailableException: No auditor elected, though Autorecovery is enabled. So giving up.
at org.apache.bookkeeper.client.BookKeeperAdmin.triggerAudit(BookKeeperAdmin.java:1543) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.client.BookKeeperAdmin.decommissionBookie(BookKeeperAdmin.java:1579) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.tools.cli.commands.bookies.DecommissionCommand.decommission(DecommissionCommand.java:96) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.tools.cli.commands.bookies.DecommissionCommand.apply(DecommissionCommand.java:81) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.bookie.BookieShell$DecommissionBookieCmd.runCmd(BookieShell.java:1947) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.bookie.BookieShell$MyCommand.runCmd(BookieShell.java:246) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.bookie.BookieShell.run(BookieShell.java:2345) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.bookie.BookieShell.main(BookieShell.java:2436) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
2023-04-19T08:01:07,894+0000 [main-SendThread(localhost:2181)] WARN  org.apache.zookeeper.ClientCnxn - An exception was thrown while closing send thread for session 0x1001228f9f1000d.
org.apache.zookeeper.ClientCnxn$EndOfStreamException: Unable to read additional data from server sessionid 0x1001228f9f1000d, likely server has closed socket
at org.apache.zookeeper.ClientCnxnSocketNIO.doIO(ClientCnxnSocketNIO.java:77) ~[org.apache.zookeeper-zookeeper-3.8.0.jar:3.8.0]
at org.apache.zookeeper.ClientCnxnSocketNIO.doTransport(ClientCnxnSocketNIO.java:350) ~[org.apache.zookeeper-zookeeper-3.8.0.jar:3.8.0]
at org.apache.zookeeper.ClientCnxn$SendThread.run(ClientCnxn.java:1282) ~[org.apache.zookeeper-zookeeper-3.8.0.jar:3.8.0]
only after I deleted the whole cluster including volumes and redeployed it, everything started working
✅ 2
m
Sorry you had to delete the whole cluster. I remember running into this here: https://github.com/apache/pulsar-helm-chart/pull/266. I agree that removing it solves the problem. You might have been able to avoid this by manually changing the ownership of the bookie's files.
h
You’d better not modify the pvs of the bookie pod, it will lead to the bookie pod can’t start up
m
@Hang Chen are you saying not to remove the PVs? I agree with that. However, based on my testing, you can safely chown the files to be writable by the newest docker image.
h
you can safely chown the files to be writable by the newest docker image.
What do you mean by chown the file to be writable?
m
If you upgrade to docker image of 2.10 or later, you need to make all existing bookie and zookeeper files writable by the root group
h
Oh I see, Pulsar docker image of 2.10+ uses non-root to start up and it has no permission to access the old pv with root owner
m
Right
m
for 2.11 upgrade, we need to set security context for pulsar components
Copy code
securityContext:
+         runAsNonRoot: false
+         runAsUser: 0
m
Why? That shouldn’t be necessary. The docker image is meant to run as a non root user to increase security.
m
Hi, After upgrading successfully one cluster, I proceeded to the other one and now i’m facing this error:
2023-04-23T14:27:21,835+0000 [main-EventThread] ERROR org.apache.bookkeeper.proto.PerChannelBookieClient - Cannot connect to pulsar-bookie-2.pulsar-bookie.pulsar.svc.cluster.local:3181 as endpoint resolution failed (probably bookie is down) err org.apache.bookkeeper.proto.BookieAddressResolver$BookieIdNotResolvedException: Cannot resolve bookieId pulsar-bookie-2.pulsar-bookie.pulsar.svc.cluster.local:3181, bookie does not exist or it is not running
2023-04-23T14:27:21,835+0000 [main-EventThread] ERROR org.apache.bookkeeper.proto.PerChannelBookieClient - Could not connect to bookie: null/pulsar-bookie-2.pulsar-bookie.pulsar.svc.cluster.local:3181, current state CONNECTING :
org.apache.bookkeeper.proto.BookieAddressResolver$BookieIdNotResolvedException: Cannot resolve bookieId pulsar-bookie-2.pulsar-bookie.pulsar.svc.cluster.local:3181, bookie does not exist or it is not running
at org.apache.bookkeeper.client.DefaultBookieAddressResolver.resolve(DefaultBookieAddressResolver.java:63) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.proto.PerChannelBookieClient.connect(PerChannelBookieClient.java:535) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.proto.PerChannelBookieClient.connectIfNeededAndDoOp(PerChannelBookieClient.java:667) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.proto.DefaultPerChannelBookieClientPool.obtain(DefaultPerChannelBookieClientPool.java:121) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.proto.DefaultPerChannelBookieClientPool.obtain(DefaultPerChannelBookieClientPool.java:116) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.proto.BookieClientImpl.readEntry(BookieClientImpl.java:509) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.proto.BookieClientImpl.readEntry(BookieClientImpl.java:495) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.proto.BookieClientImpl.readEntry(BookieClientImpl.java:489) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.client.ReadLastConfirmedOp.initiate(ReadLastConfirmedOp.java:81) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.client.LedgerHandle.asyncReadPiggybackLastConfirmed(LedgerHandle.java:1440) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.client.LedgerHandle.asyncReadLastConfirmed(LedgerHandle.java:1403) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.client.LedgerOpenOp.openWithMetadata(LedgerOpenOp.java:225) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.client.LedgerOpenOp.lambda$initiate$0(LedgerOpenOp.java:119) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at java.util.concurrent.CompletableFuture.uniWhenComplete(CompletableFuture.java:863) ~[?:?]
at java.util.concurrent.CompletableFuture$UniWhenComplete.tryFire(CompletableFuture.java:841) ~[?:?]
at java.util.concurrent.CompletableFuture.postComplete(CompletableFuture.java:510) ~[?:?]
at java.util.concurrent.CompletableFuture.complete(CompletableFuture.java:2147) ~[?:?]
at org.apache.bookkeeper.meta.AbstractZkLedgerManager$4.processResult(AbstractZkLedgerManager.java:506) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.zookeeper.ZooKeeperClient$19$1.processResult(ZooKeeperClient.java:997) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.zookeeper.ClientCnxn$EventThread.processEvent(ClientCnxn.java:634) ~[org.apache.zookeeper-zookeeper-3.8.0.jar:3.8.0]
at org.apache.zookeeper.ClientCnxn$EventThread.run(ClientCnxn.java:553) ~[org.apache.zookeeper-zookeeper-3.8.0.jar:3.8.0]
Caused by: org.apache.bookkeeper.client.BKException$BKBookieHandleNotAvailableException: Bookie handle is not available
at org.apache.bookkeeper.discover.ZKRegistrationClient.getBookieServiceInfo(ZKRegistrationClient.java:248) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
at org.apache.bookkeeper.client.DefaultBookieAddressResolver.resolve(DefaultBookieAddressResolver.java:43) ~[org.apache.bookkeeper-bookkeeper-server-4.15.3.jar:4.15.3]
... 20 more
Any idea anyone?
An Update: I reverted back to 2.9.3 and restarted the upgrade in this sequence: 1. Upgraded successfully zookeeper to 2.11.0 2. Disabled autorecovery from within bookie pods 3. Started with bookie upgrade to 2.11.0 now the upgrade is stuck on these messages for almost 3 hours on the first bookie pod:
[11330.233s][info][safepoint] Safepoint "FindDeadlocks", Time since last: 109301 ns, Reaching safepoint: 76091 ns, At safepoint: 7540 ns, Total: 83631 ns
[11345.233s][info][safepoint] Safepoint "FindDeadlocks", Time since last: 14999710027 ns, Reaching safepoint: 87091 ns, At safepoint: 35820 ns, Total: 122911 ns
[11345.233s][info][safepoint] Safepoint "FindDeadlocks", Time since last: 176902 ns, Reaching safepoint: 78191 ns, At safepoint: 14240 ns, Total: 92431 ns
Any idea how to solve this issue?
m
I’m not sure what’s wrong there, but I will note that pulsar clusters should be upgraded sequentially. You should go 2.9.x to 2.10.y then 2.10.y to 2.11.z.
h
Cannot resolve bookieId pulsar-bookie-2.pulsar-bookie.pulsar.svc.cluster.local:3181, bookie does not exist or it is not running
You’d better check if bookie-2 is running or not
m
It’s stuck on 0/1 Running which means it’s not ready Anyway i’m trying now to first of all upgrade to 2.10.4 and proceed to 2.11.0
y
check if you have “readOnlyRootFilesystem” or “allowPrivilegeEscalation=false” in your security context. somehow, init-container doesn’t escalate when trying to write config to /pusar/conf (a part of overlay fs), readonly.