Hey all. I’m looking for some usecases or inferenc...
# random
s
Hey all. I’m looking for some usecases or inferences on how all of you use Autoscale with flink. Pls share your thoughts? For both for session and app mode.
flink 1
a
Hey @sairam yeturi! Can you please elaborate more about what you’re looking for? A common example in our case is that most of our pipelines have daily fluctuations, so we scale out and in based on the time of day. We tried using a metric to indicate the required scale, but we couldn’t get the unit of scale right - especially since every rescaling requires restarting the job.
flink 1
s
That’s cool, I was looking at how everyone uses Autoscale, in app mode or session mode clusters… do you use native kubernetes HPA for scaling? What key metrics you use if it’s metric based scaling. Also, during scaling as you mentioned restart issues persist? @Amir Halatzi
a
@sairam yeturi yes, we used native HPA because 1.17 hasn’t been released yet 🙂 We tried to use Kafka lag as our metric, and it proved hard to find a rule of thumb for setting the scaling window - partly because of restarts: not only does updating the deployment will cause the job the restart, but sometimes task managers will take some time to be provisioned and then restart whenever another task manager joins the cluster
flink 1
s
Thanks amir. Appreciate your feedback. Hey all- Any more perspectives from you this flink community on reactive mode for scaling vs native HPA k8s for scaling ✌️ what’s your pick?
Quick check if anyone else has any inputs thanks again.
t
Hi @Amir Halatzi, what will change in 1.17 that will affect autoscale?
s
@Tal Sheldon how do you use scaling today?
t
Currently elastic scaling (reactive mode) - this was the only way we could have real autoscale. BUT since we are not yet in production with Flink - it's not battle tested on our end. It could change. infra - K8S
👍 1
s
Do you see any challenges in the current setup with elastic scaling
a
Hey @Tal Sheldon! I can't find it in the docs ATM, by afaik the autoscale feature in the k8s operator requires version 1.17
y
hi @sairam yeturi our flink is deployed in yarn, and we implement an feature of rescale which will not restart the job, which mean that we will change the streamGraph and jobGraph in runtime, and then stop old execution and start the new execution in runtime
👍 1
s
Thank you this is helpful to know