All- The historical process on our production Drui...
# general
a
All- The historical process on our production Druid cluster has suddenly crashed with an OOM error and is failing to start up. Exception: Native memory allocation (mmap) failed to map.... Tried to increase the memory allocation as below JVM: 8GB --> 12GB MaxDirectMemorySize: 34GB --> 75GB But it still fails to start up. We noticed that the process fails after acquiring ~61.5 GB of memory (from "top" command). Host memory: 125GiB What is the max "MaxDirectMemorySize" that can be allocated for this host? Is it only 50% of the host memory ~62GB? Thanks, AR.
k
any idea how many segments you have on each historical node? If it’s above 65k you might need to look at https://druid.apache.org/docs/latest/operations/basic-cluster-tuning#linux-limits
a
Hi, One of the tables has 53K segments. Next one is ~12K. But yes in total, it is 80k - 85K segments. But would this cause historical to go OOM? Wouldn't it throw a different error if ulimit settings were breached? Thanks AR.
Also, we have 4 historicals. So all 85K wouldn't be on one server. I checked the ulimits and it is set to 65K. Will check on "vm.max_map_count".
vm.max_map_count = 65530
k
it’s not the ulimit - it’s the vm.max_map_count - try increasing it to something like
500000
a
Thanks Kyle. Will try to get that done. But would this cause an OOM in the historical? Thanks, AR.
k
yes, it is an OOM - just not the kind you might not normally think of
a
Thanks Kyle. Will try to get the done. In the meantime, how do we get the cluster up? If we mark some of the segments unused using the API, it should reduce the number of segments to be loaded and we should be able to bring the cluster up. Correct?
k
yes, that makes sense. You could also clear the segment cache (delete the local copy of the segment files on the historical)
if it has too many segments copied locally, it will fail to startup
a
I was going to ask that. So to bring it up, we should: 1. Mark segments as unused using the API. 2. Delete the local segment cache so that it doesn't try to load everything from it. Anything else that we should do?
HI Kyle, Would marking segments as unused suffice should the segments have to be deleted? Thanks, AR.
Is it that one segment file --> one memory map? I was checking the table with the largest segment count. It is a daily table but for one particular year, it has around 40K segments. I am thinking of reindexing and compacting the data for that year which should reduce the number of segments. Would this help in Druid not breaching the current "vm.max_map_count" limit? Thanks, AR.