Any thoughts on why middlemanager CPU usage would ...
# general
s
Any thoughts on why middlemanager CPU usage would be under 10% even though there are 50+ indexing tasks running on that 64cpu node and all of them are working fine? I'm not complaining but I'm curious to know if there is any room for optimization 🙂
j
Hi Shantha, Is the entire server running at 10%, or MM process only is using 10% MM tends to not do very much (hence the name) so not surprising if it's cpu usage is minimal.
s
Hi John, the entire server is running at 10%
j
In the Imply product the MM and Historical are co-located ... maybe that was done due to the lower cpu usage ... but I would think still you should be able to make good use of the cpu even if just for MM and ingestion tasks. What does your supervisor spec look like? Maybe the row buffer, persist and build parameters could be optimized better? Are you ingesting JSON, or a compressed format like Parquet? JSON uses more cpu for parsing, the compressed data format are supposed much more cpu-efficient ...
s
I have ~30 dimensions and a few thetasketch metrics. According to ingest/events/processed, the supervisor processes ~3 million events. I'm using the defaults for most of the supervisor spec tuning config settings. I'm reading using the avro_stream input format, not sure if that helps.
j
Can you share your supervisor spec? You can obfuscate the datasource and field names if you want ... but would help to see all of the parameters involved. Thx.
b
Are the tasks realtime indexing or batch indexing? Typically querying is more CPU intensive than ingestion. So the MM may be at 10% utilization while ingesting with no query activity, but significantly higher if you're ingesting and querying realtime tasks simultaneously.