Hi Druid Team, Are there any plans to allow using ...
# dev
j
Hi Druid Team, Are there any plans to allow using the multiphase segment merging strategy (
IndexMergerV9.multiphaseMerge
- src) when publishing segments from stream ingestion (e.g. kafka)? This strategy can be configured in batch ingestion and compaction by setting
maxColumnsToMerge!=-1
but not for stream ingestion. I took a look at the relevant code section (StreamAppenderator.mergeAndPush) and it appears that this is already prepared:
Copy code
mergedFile = indexMerger.mergeQueryableIndex(
            indexes,
            schema.getGranularitySpec().isRollup(),
            schema.getAggregators(),
            schema.getDimensionsSpec(),
            mergedTarget,
            tuningConfig.getIndexSpec(),
            tuningConfig.getIndexSpecForIntermediatePersists(),
            new BaseProgressIndicator(),
            tuningConfig.getSegmentWriteOutMediumFactory(),
            tuningConfig.getMaxColumnsToMerge()  // <-- always -1 for stream ingestion tasks (default implementation in `AppenderatorConfig` is never overridden)
        );
So basically this would "only" require to make
maxColumnsToMerge
configurable in the respective
xxxTaskTuningConfig
for kafka/kinesis/rabbitmq/etc. and to update the UI (API/WebConsole). Are there any reasons against using multiphase merge in stream ingestion at all or is this simply not (yet) implemented? Thanks in advance!
g
plans— not that i'm aware of although it would be a good idea; it seems like a miss that the config doesn't exist for realtime. a PR to add it would be appreciated!
j
Thanks for the prompt reponse 👍 I'm not sure if I can provide a meaningful PR as I'm not proficient in Java but I'll try anyway 😅. Assuming you're familiar with the relevant Druid internals - are there any more parts that need to be updated despite the actual configs and the API/Console?
g
that's all i can think of; i think it should plug in well