This message was deleted.
# general
s
This message was deleted.
j
Hi Sanyog, I don't know if there is a rule of thumb for overall cluster sizing, a lot depends on the format of the data (which determines final resting size in Druid), ongoing ingestion activity, and ongoing query activity. I would think to start with the Data Nodes, determine what you need to hold two copies of Druid segments on local disk. But the size is dependent on the format of the incoming data. So first question -- what is the format of your 5TB data? If JSON then you may see a 10:1 reduction in size, meaning you only need 500GB x 2 replicas = 1TB of overall disk space. If so, then (using AWS instance typess as an example) m5d data nodes may work fine. If the data is already in a compressed format (e.g. parquet) then you may need the full 5TB x 2 = 10TB disk, in which case i3 data nodes may be better. From there, you size query nodes based on query activity, and Master nodes generally 2 or 4 cpus each. Here are a couple of examples of clusters that I have seen used.
s
Hi John, thanks for this insight. This really helped
👍 1