Hi All, I have question regarding kill tasks locks...
# general
i
Hi All, I have question regarding kill tasks locks vs kafka ingestors locks: 1. it seems kill tasks are WAITING due to some locks taken by kafka ingestors, is it ok? something I can do about it? 2. kill task is supposed to remove only unused segments so why it takes any locks? 3. suppose we started to use concurrent append feature (i.e. "useConcurrentLocks": true ) both under compaction & kafka ingestors, do we have option to pass same context for kill tasks?
a
suppose we started to use concurrent append feature (i.e. "useConcurrentLocks": true ) both under compaction & kafka ingestors, do we have option to pass same context for kill tasks?
Kill tasks cannot run concurrently with ingestion using these locks unfortunately
kill task is supposed to remove only unused segments so why it takes any locks?
I believe this is due to the legacy flag
markAsUnused
where used segments can also be marked as unused and killed
Since it has been deprecated, maybe we don't need to hold locks anymore but I'm not sure if there are other implications
i
Hi @Amatya Avadhanula thanks for answers regarding legacy flag of
markAsUnused
you mean that kill task itself can mark segments as unused and kill(aka delete) segments in same "run"?
a
Yes
Copy code
if (markAsUnused) {
      numSegmentsMarkedAsUnused = toolbox.getTaskActionClient().submit(
          new MarkSegmentsAsUnusedAction(getDataSource(), getInterval())
      );
      <http://LOG.info|LOG.info>("Marked [%d] segments of datasource[%s] in interval[%s] as unused.",
               numSegmentsMarkedAsUnused, getDataSource(), getInterval());
    } else {
      numSegmentsMarkedAsUnused = 0;
    }
i
a
BTW I just saw that there are tasks such as
MoveTask
RestoreTask
and
ArchiveTask
These fetch unused segments. So perhaps the KillTask lock prevents these from running together
i
Basically I'm not sure if I want to mark everything unused. Since there are both used and unused segments for recent intervals. Compaction task now runs in parallel to kafka ingestors, using concurrent locks. However multiple compactions create additional load on deep storage and I was trying to run kill tasks to remove old (already compacted) segments...
a
If you are using concurrent append and replace, please do not manually clean deep storage locations
Different segments may reuse the same file in deep storage
i
yes, i don't want to do it
👍 1
but i do want to cleanup old (compacted) segments
maybe i need to pass parameter that compaction should delete compacted segments?
a
Compaction does not delete compacted segments. I think one way would be to try adding a parameter to kill tasks so that they do not take locks. Or perhaps allow them to use APPEND locks (I'm not sure if this is a good idea)
i
a
I think that is for creating tombstones for index_parallel / compact tasks
j
I would think Kill Tasks require time interval locks in case you try to re-mark those segments as used again (e.g. retention policy change) while the kill task is running ... would that be the case?
a
Marking segments as used with the coordinator apis may not check for task locks as far as I know. I'm not 100% sure
i
@John Kowtko no, not doing policy changes or re-marking as used. My intention is only to remove compacted segments. i'll post later payload
@Amatya Avadhanula do you know how to control parameters of kill tasks issued by coordinator?
a
Which parameters specifically?
i
@Amatya Avadhanula I mean what you suggested above
Copy code
I think one way would be to try adding a parameter to kill tasks so that they do not take locks.
a
Ah, no. The change that I suggested would require code changes
i
oh, i see
a
You would have to override the
isReady
method for
KillUnusedSegmentsTask
to use
APPEND
locks instead of
EXCLUSIVE
I'm not 100% sure that it wouldn't have other repercussions. It might work for your environment though
i
@Amatya Avadhanula regarding your comment of
Copy code
If you are using concurrent append and replace, please do not manually clean deep storage locations
Different segments may reuse the same file in deep storage
does segment table contains data that can help me to analyzer which file is used and which file is already not used by any segment? I see payload that contains
Copy code
"path": "<hdfs://taz-il/production/druid_deep_storage/sp_campaigns_realtime_aggregation_by_request_time/20240413T000000.000Z_20240413T010000.000Z/2024-04-13T00_00_01.528Z/7_f908dc30-38eb-4489-b0f4-21d6a4f14fa9_index.zip>"
In addition, do you think I should fill request for improvement in this area?
a
You would have to scan the entire segments table for this
I recall that you said that you don't mark data as used later. In that case the kill task can handle load spec duplication
In addition, do you think I should fill request for improvement in this area?
Sure, please create an issue. Are you also interested in contributing to it?
i
can you please elaborate about
Copy code
I recall that you said that you don't mark data as used later. In that case the kill task can handle load spec duplication
I'm interested in contributing 🙂 I'm currently working on something else - https://github.com/apache/druid/pull/16266 (maybe it's less important)
🙌 1
I indeed not marking anything as used manually. Only ingestors & compactions in my case. So what kill task should do?
a
I mean that you don't intend to mark unused segments as used after they have been marked by the coordinator
i
yes, true that
but what does it mean that "In that case the kill task can handle load spec duplication"
a
Let's say you have segments S1 and a deep storage location D such that S1 -> D
A kill task generally deletes S1 from the metadata store and deletes D from deep storage
But suppose you have a relation like S1 -> D <- S2
If you kill only S1, D must not be deleted
This will be handled by the kill task. Only when no used segment S_i uses D, does the kill task delete D from deep storage
i
yes
so you mean better not to do it manually
a
Yes
i
and improve kill task
a
Yep
i
ok, understood
thanks !
👍 1
a
i
opened this issue, will wait for some comments and then will start work on it https://github.com/apache/druid/issues/16361
a
Thanks for the issue, and the PR! Will take a look soon
Left some comments. I initially said that an APPEND lock would work fine but I have some concerns about multiple kill tasks running on the same interval because of this change
i
thanks, will address them shortly
I've pushed changes per your suggestion. If it looks ok, I'll add tests. I had 1 question though - where is the place that coordinator creates kill task? and how I can control coordinator issued kill task context
a
I'll take a look soon