This message was deleted.
# troubleshooting
s
This message was deleted.
k
Yes, compaction does download all the segments for an interval to get the metadata before it can proceed. Thus, it faces the risk of running out of disk space. I think there is a GitHub issue for an improvement to avoid downloading the segments. Auto-compaction also launches separate tasks for each interval grain, which is certainly not needed while doing a re-index. I wonder if auto-compaction could just submit a Druid re-index task instead of using a specialized CompactionTask.
d
I think that would be way better, and the code will be more general as well. Perhaps, in the future CompactionTask can be deprecated and removed, just tell all customers to migrate to Druid ingestion.
m
Compaction task has the advantage of user not needing to specify all the specs (dimensions, metrics, etc) and is the reason it needs to download all the segments for an interval
k
So, after the segments are downloaded and the specs are determined, it effectively becomes a regular re-indexing task, right? If that's the case, the right direction for now would be to figure out a way to have enough metadata that helps us eliminate or reduce the segment downloads needed to figure out the specs.
g
yes, definitely; for this reason it's related to the catalog effort
catalog, if used, would provide us with enough info to do a compaction without inspecting all segments first
šŸ’Æ 1
d
ah, i see. If I re-ingest from Druid datasource then I need to describe the dimensions and metrics explicitly? I was about to ask if there’s a property that I can pass to Druid ingestion spec to just use all the dimensions and metrics as-is.
m
So, after the segments are downloaded and the specs are determined, it effectively becomes a regular re-indexing task, right?
-> yes
If that's the case, the right direction for now would be to figure out a way to have enough metadata that helps us eliminate or reduce the segment downloads needed to figure out the specs.
-> yes, one way is to store the ingestionSpec used for the creation of the segment in the metadata store. This is similar to storing lastCompactionState that we store for all segment created by a compaction job. Another way is the catalog effort
ah, i see. If I re-ingest from Druid datasource then I need to describe the dimensions and metrics explicitly?
-> yes, if you forget to explicitly specify any metric or dimension then you can lose data!
I was about to ask if there's a property that I can pass to Druid ingestion spec to just use all the dimensions and metrics as-is.
-> that’s what the (Auto)Compaction is there for.
🤯 1
d
Sorry for this dumb question, how come the Console UI is able to know all of the columns when the ā€œwizardā€ is helping the user craft the ingestion JSON?
@Vadim knows something compaction job don’t šŸ˜„
v
The console uses the
sampler
API. It samples the data for any inputSource šŸ˜›
g
so, it could miss some columns, potentially, if your segments have different schema from segment to segment
d
ooo gotcha, glad that I didn’t YOLO this Druid re-ingestion to prod.
Hm… actually, our group always versioned the ingestion spec in Git. This may still work.