This message was deleted.
# troubleshooting
s
This message was deleted.
c
What specifically do you mean by "multi region"?
d
Setup 2 Druid clusters but they are both backed by the same S3 and PG. And then I can use DNS/LB trickery to swing traffic between the two.
I am investigating absolutely zero downtime for Druid cluster at the moment…
c
Oh fascinating
d
If Cluster A’s Historical and MM use tierA and Cluster B’s Historical and MM use tierB… I think it cloud work. LMAO, could work.
r
cloud work 😛
well, if you have they connected to the same ZK, they will be just one cluster still
d
I think ZK would be non-issue, even if you have a dedicated Zookeeper per Kubernetes cluster.
r
if you have separate ZK, each will think that they have a coordinator leader and may corrupt the database, but as long you only set up a historical, it's fine for a fast- recovery until you kill the other ZK/coordenators
d
because internal traffic (brokered by ZK) never escapes to outside.
and Historical segment files don’t contain cluster information, which is a great design! 👍
c
With the kubernetes extensions instead of ZK you can control which cluster owns the leader via kubernetes APIS.
If the ZK issue becomes problematic
r
I think the main issue is to not ever have 2 coordenators as laeder running
having a second 'dr' cluster with just historicals could indeed work, I think, if they use the same ZK and have a different tier in order to not cause each query routed on the 'active' cluster goes across a potential expensive network round-trip and for the MM in the 'dr' cluster, connected to the same ZK, and add the task replica > 1 to have it HA but I don't remember if there's tiers for MM but I think that's useless bc if the database or ZK fails on the active region, the dr MM would not be able to commit and serve queries when the task finsihes
d
MM has category, which can be abused in similar manner 😛 Even better, it can be updated dynamically from the Overlord Dynamic Config 😄
MM is less of a concern for me because the MM cluster can be down on the DR cluster. Historical is much more important because data sync takes super long time.
g
It's not that crazy 🙂
I've seen people do something similar for blue/green deploys: launch a whole new tier and then terminate the old one
❤️ 1
Just keep in mind that the clusters will be linked in some ways and so you still have a chance of correlated failure (for example if the shared metadata store or shared deep storage has problems)