Jakob Riebe
08/02/2024, 2:10 PMIndexMergerV9.multiphaseMerge - src) when publishing segments from stream ingestion (e.g. kafka)?
This strategy can be configured in batch ingestion and compaction by setting maxColumnsToMerge!=-1 but not for stream ingestion.
I took a look at the relevant code section (StreamAppenderator.mergeAndPush) and it appears that this is already prepared:
mergedFile = indexMerger.mergeQueryableIndex(
indexes,
schema.getGranularitySpec().isRollup(),
schema.getAggregators(),
schema.getDimensionsSpec(),
mergedTarget,
tuningConfig.getIndexSpec(),
tuningConfig.getIndexSpecForIntermediatePersists(),
new BaseProgressIndicator(),
tuningConfig.getSegmentWriteOutMediumFactory(),
tuningConfig.getMaxColumnsToMerge() // <-- always -1 for stream ingestion tasks (default implementation in `AppenderatorConfig` is never overridden)
);
So basically this would "only" require to make maxColumnsToMerge configurable in the respective xxxTaskTuningConfig for kafka/kinesis/rabbitmq/etc. and to update the UI (API/WebConsole).
Are there any reasons against using multiphase merge in stream ingestion at all or is this simply not (yet) implemented?
Thanks in advance!Gian Merlino
08/07/2024, 8:47 PMJakob Riebe
08/08/2024, 11:54 AMGian Merlino
08/08/2024, 4:46 PM