Slackbot
10/28/2022, 1:32 PMArmanit Garg
10/28/2022, 1:32 PMdruidDiscoveryAnnouncement-<node-role>
druidDiscoveryAnnouncement-id-hash
druidDiscoveryAnnouncement-cluster-identifier
Similarly, when it unannounces itself, it will remove the above labels. On the discovery side, we initialize a K8s watcher which looks for pods with all the above labels.
The problem lies with the fact that we can have peons running in a single pod with middle managers. When a peon dies, it will unannounce itself and that removes the common labels (druidDiscoveryAnnouncement-id-hash, druidDiscoveryAnnouncement-cluster-identifier) which were also being used by the middle manager process to announce itself. This leads to the K8s watcher filtering out the middle manager process running on the same pod since the required label does not exist anymore. So the middle manager ends up zombie'ing out because the overlord doesnt see it anymore.Armanit Garg
10/28/2022, 1:33 PMdruidDiscoveryAnnouncement-<node-role>
druidDiscoveryAnnouncement-id-hash-<node-role>
druidDiscoveryAnnouncement-cluster-identifier-<node-role>
so that there are no common labels being used by any druid process and that keeps announces and unannounces clean.
Unfortunately, this will break backwards compatibility for existing deployments using this extension. This is because the K8s watcher will start looking for nodes with the druidDiscoveryAnnouncement-cluster-identifier-<node-role> label and already running nodes will not follow that convention.Armanit Garg
10/28/2022, 1:34 PMIan Roberts
10/28/2022, 3:51 PMArmanit Garg
10/28/2022, 3:59 PMIan Roberts
10/28/2022, 4:08 PM