This message was deleted.
# general
s
This message was deleted.
a
What parts would you want to scale up and down?
m
Hey @D K I thought this was a nice description of the ways that you might want to scale Druid.
1. MiddleManager Nodes: In Druid, the MiddleManager nodes are responsible for data ingestion and can be scaled up or down based on the workload. If you're experiencing high loads or increased ingestion rates, you can add more MiddleManager nodes to handle this load. 2. Historical Nodes: These nodes are used for storing and querying data. You can scale these nodes horizontally to increase query capacity and data storage. 3. Coordinator and Overlord Nodes: While these nodes are less about data processing and more about management, they play a crucial role in scaling. The Coordinator node assigns segments to Historical nodes, and the Overlord node manages the distribution of ingestion tasks to MiddleManagers. 4. Auto-scaling: For dynamic scaling, especially in cloud environments like AWS, GCP, or Azure, you can use their respective auto-scaling features. However, this requires external setup and configuration. Druid itself doesn't provide a built-in auto-scaler but can work with these cloud-based auto-scaling tools. 5. Kubernetes: If you are running Druid in a Kubernetes environment, you can leverage Kubernetes' auto-scaling capabilities to scale your Druid pods based on metrics like CPU and memory usage. 6. Third-Party Tools and Scripts: Some users implement custom scripts or use third-party tools to monitor the performance metrics of Druid clusters and dynamically adjust the number of nodes based on predefined thresholds or rules. 7. Manual Scaling: While not dynamic, manual scaling based on anticipated workloads (like scaling up before expected high-traffic events) is also commonly practiced.
j
So to sum up, "dynamic" in the sense that any of the services can be scaled up or down while the cluster remains fully online, yes Druid was architected to support that in all areas of operation. "dynamic" in the sense of auto-scaling, where Druid takes care of the scaling based on changes in workload over time, afaik that is only handled by Druid in Kafka and Kinesis streaming ingestion: • https://druid.apache.org/docs/latest/development/extensions-core/kafka-supervisor-reference#task-autoscaler-properties • https://druid.apache.org/docs/latest/development/extensions-core/kinesis-ingestion#task-autoscaler-properties There is also mention of autoscaling MiddleManagers in conjunction with the underlying infrastructure: • https://druid.apache.org/docs/latest/design/overlord#autoscaling