This message was deleted.
# general
s
This message was deleted.
g
typically in the release notes for each version there is a section about any compatibility / upgrade related issues, that also discusses potential rollback concerns
the main one is the
frontCoded
change, where if you are using
frontCoded
dictionaries, the format in 26 is not readable by 25
note that by
frontCoded
dictionaries are not enabled by default, so you probably aren't using them, unless you specifically turned them on in https://druid.apache.org/docs/latest/ingestion/ingestion-spec/#indexspec
j
yeah I read them, not using that currently
g
cool
so in that case rollback should just be a matter of reversing the upgrade
i.e. do the rollback in the reverse order of how you update the servers when you upgrade them
j
we are using helm charts so, would it be a matter of changing back to 25.0.0 in the image?
g
coordinator first, then broker, then data servers when doing a rollback. (upgrade is data servers first, then brokers, then coordinators)
hmm I'm not sure how rollback works with the helm charts specifically— I'm not sure if it rolls things in the "proper" order
j
coordinator first, then broker, then data servers when doing a rollback. (upgrade is data servers first, then brokers, then coordinators)
for data servers you mean middleManagers and Historical right - what about router?
g
if it doesn't, then any issues you have should be transient at least (possibly some API errors while there's mixed versions deployed "out of order", i.e., if you have a historical running an older version than a broker, that may give rise to API errors until the historical is updated)
yeah for data servers i mean MM and historical
router would go along with broker (the order between those isn't super important, they can be simultaneous)
j
thanks
I am planning then to go from 26.0.0 to 27.0.0 if dev behaves well (it is for the time being) - there are more things to check here https://github.com/apache/druid/releases/tag/druid-27.0.0#27.0.0-upgrade-notes-and-incompatible-changes tho but I think I am not using them
s
@Julian Reyes I haven't done this myself, but I have heard of folks using the Druid Operator for upgrade/rollback operations because it controls the order of upgrade.
s
I just did some upgrades in our lab environment from 25 to 28.0.1 using the Druid operator. With the operator, it was a case of changing the version on the image in the yaml file for the cluster then letting the operator handle the restarts. Other than an issue with the router needing to have more threads configured, the upgrade went smoothly. I did it in 2 steps, 25->26 and 26->28.0.1. This was based on a comment in the 28 upgrade notes about JSON columns (which we are using)
🙌 1
j
so you went directly from 26 to 28.0.1, interesting, I need to read that json notes
we do not use operator tho, we installed druid using helm and argocd, perhaps we could incorporate the operator
I have doubts with the rollback in the case where SQL schema needed to be updated
s
The operator definitely makes things much easier to manage, especially as the cluster grows in size
Yeah, for our instance, we are just consuming data from Kafka as Avro formatted messages, no SQL schemas or anything outside of Kafka
j
we consume directly from kinesis. What I referred in SQL is the update that needs to happen for a new column to be added I think
s
Ahh...in our case we were able to allow it to happen automatically so no need to manually update the tables
j
I was able to upgraded Dev cluster to latest (28.0.1). Like you Sean, I went 25.0.0 to 26.0.0 and then straight to 28.0.1. One issue I came across is that I have a postStart lifecycle that copies postgresql to
/opt/druid/lib/
however upgrades failed because postgresql was also upgraded to
42.6.0
. Updated that part and upgraded was successfully
s
Great! I have my postgresql as a separate deployment so it is independent of the Druid upgrade. The one big issue I hit was the broker process taking forever to start up as the historicals were slow to respond as they started up. Probably just a setting in newer Druid that I did not have configured
j
The pros cluster is pretty big, we have 12 historical and 12 middle manager so I thing same thing will happen since historical in prod take a long time to start up
Also, do we know if
druid_indexer_runner_k8s_podTemplate_base
should go to into ConfigMap or into coordinator yml deployment as env variable