This message was deleted.
# troubleshooting
s
This message was deleted.
a
When a druid process announces itself, it will add the following labels to a pod:
Copy code
druidDiscoveryAnnouncement-<node-role>
druidDiscoveryAnnouncement-id-hash
druidDiscoveryAnnouncement-cluster-identifier
Similarly, when it unannounces itself, it will remove the above labels. On the discovery side, we initialize a K8s watcher which looks for pods with all the above labels. The problem lies with the fact that we can have peons running in a single pod with middle managers. When a peon dies, it will unannounce itself and that removes the common labels (
druidDiscoveryAnnouncement-id-hash
,
druidDiscoveryAnnouncement-cluster-identifier
) which were also being used by the middle manager process to announce itself. This leads to the K8s watcher filtering out the middle manager process running on the same pod since the required label does not exist anymore. So the middle manager ends up zombie'ing out because the overlord doesnt see it anymore.
The simplest solution would be to update the labels to add the node-role to the key:
Copy code
druidDiscoveryAnnouncement-<node-role>
druidDiscoveryAnnouncement-id-hash-<node-role>
druidDiscoveryAnnouncement-cluster-identifier-<node-role>
so that there are no common labels being used by any druid process and that keeps announces and unannounces clean. Unfortunately, this will break backwards compatibility for existing deployments using this extension. This is because the K8s watcher will start looking for nodes with the
druidDiscoveryAnnouncement-cluster-identifier-<node-role>
label and already running nodes will not follow that convention.
Implementing a backwards compatible solution is possible but will be quite messy because I think we will have to initialize 2 K8s watchers (code for reference) because afaik K8s labels dont really support an OR on keys. So im open to any other ideas folks might have about how we can fix this. Also happy to contribute a fix for this if we can come up with a good solution.
i
throwing this out there and may be completely off base. Could this be adapted to indexer process which is experimental leaving the existing labels alone? Then there is an option to switch to non zookeeper indexers leaving backwards compatability? Not sure but was just thinking about this and it was interesting
a
For our use case we want to stick to using middle managers and peons. If I am understanding your suggestion correctly, using indexers wouldnt really solve the problem right? It would just not run into this edge case.
i
true, was just brainstorming