Hi All, Is there a way to mark the segments a tabl...
# general
a
Hi All, Is there a way to mark the segments a table "used" based on interval? All the segments in the table had to be marked "unused" as it had too many small segments and was causing the historical to go OOM. Related thread: https://apachedruidworkspace.slack.com/archives/C0303FDCZEZ/p1713280553965159 To temporarily resolve the issue in the above thread, we had to mark all segments of a table "unused". Then the historical started up fine and we were able to compact other tables with many segments to reduce the total number of segments. Now we want to reload this table incrementally in intervals and compact it before loading the next interval. Looking at the console, the only option that I see is to mark all segments "used". But this might cause the historical to go OOM again - we have scheduled a change to increase the "max_map_count" setting but that will only be done in a day or two. Is it possible to load the table in chunks of intervals? If I load the data for a day, the datasource gets created in Druid. Then can I use the option "Mark as used segments by interval" and choose a historical interval which is marked as "unused"? Would this work? Thanks, AR.
k
Have a look at the API - I think this one is what you’re looking for: https://druid.apache.org/docs/latest/api-reference/data-management-api#mark-a-group-of-segments-used
a
Thank you Kyle. I was thinking of using the API but the approach that I suggested worked. We loaded data for one day from the source which recreated the datasource in Druid. Then we were able to incrementally "mark segments in interval as used" from the web console. With this, we are currently loading data for each month and compacting it before proceeding to the next month. Thanks, AR.
j
Hi AR, One "trick" you can use is to mark the segments as used but with 0 replicas loaded on historicals. This is "cold tier" mode. The segments will not be loaded on any historical but are still queryable by MSQ ... and I believe can also be seen by batch jobs, so can still be compacted. Let us know if you try that approach. Thanks. John
a
Hi John, I have read about this but I believe this is not available on the version of Druid that we are currently using (v24.0.2). Thanks, AR.