This message was deleted.
# general
s
This message was deleted.
a
@sriramdas sivasai, could you share some more information on the setup so we can help you. A few questions: 1. What version of Druid is the cluster on? 2. Just to confirm, you’re referring to manually submitted kill tasks, and not the coordinator-issued kill tasks from the KillUnusedSegments duty, correct? The latter has dynamic configuration properties like
killDataSourceWhitelist
,
killTaskSlotRatio
and
maxKillTaskSlots
- https://druid.apache.org/docs/latest/configuration/#coordinator:~:text=none-,killTaskSlotRatio,-Ratio%20of%20total` 3. Do you have any active ingestion going on for the datasource that you’re also trying to kill that results in lock conflicts? Could you share the kill task payload along with any other relevant logs from the stalled task? 4. Does kill for other datasources succeed?
s
1. we are using 27.0.0. 2. yes, manually submitted kill tasks 3. there is an on going ingestion happening, but there will not be a conflicts. No log has started yet 4. other datasources are getting succeeded.
Copy code
{
  "type": "kill",
  "id": "api-issued_kill_ds1001_clkndhhj_1000-01-01T00:00:00.000Z_2024-03-02T00:00:00.000Z_2024-03-03T20:18:40.016Z",
  "dataSource": "ds1001",
  "interval": "1000-01-01T00:00:00.000Z/2024-03-02T00:00:00.000Z",
  "context": {
    "forceTimeChunkLock": true,
    "useLineageBasedSegmentAllocation": true
  },
  "groupId": "api-issued_kill_ds1001_clkndhhj_1000-01-01T00:00:00.000Z_2024-03-02T00:00:00.000Z_2024-03-03T20:18:40.016Z",
  "resource": {
    "availabilityGroup": "api-issued_kill_ds1001_clkndhhj_1000-01-01T00:00:00.000Z_2024-03-02T00:00:00.000Z_2024-03-03T20:18:40.016Z",
    "requiredCapacity": 1
  }
}
a
can u check if u have a running index_kafka task
for a long period of time, im talking hours or days
a
Do you see the same issue when you resubmit the kill task after you cancel the stalled one? How many unused segments do you have in
ds1001
? To find out, you can run a query like this in the metadata store:
select count(*) from druid_segments where used = false and datasource = 'ds1001'
Also, note that several kill task improvements were made in 28.0.0 and above. If there is a ton of unused segments in the kill interval, then using a combination of
limit
and
batchSize
introduced in newer versions (28.0.0 and over) should help.
m
FYI, we are running into the same issue so https://github.com/apache/druid/issues/16030 - for this issue
a
I responded in the ticket