In the Imply product the MM and Historical are co-located ... maybe that was done due to the lower cpu usage ... but I would think still you should be able to make good use of the cpu even if just for MM and ingestion tasks.
What does your supervisor spec look like? Maybe the row buffer, persist and build parameters could be optimized better?
Are you ingesting JSON, or a compressed format like Parquet? JSON uses more cpu for parsing, the compressed data format are supposed much more cpu-efficient ...