This message was deleted.
# troubleshooting
s
This message was deleted.
d
What does the log inside MM say? It supposed to say “successfully deleted hdfs://xyz” file
a
Hi @Didip Kerabat, thank you for your response. The above lines are extract from the MM log itself. I do not see a line stating it is deleting some files from HDFS
d
Does your YARN show any activity?
a
I dont have access to my YARN right now, will check and update the thread
g
It looks like it hasn't even tried to delete anything
Is it possible your datasource does not have any unused segments in the interval you provided?
a
Hi @Gian Merlino, correct, it is not deleting at all. I verified, my datasource has a retention of
P3D
, and I see underlying segments in HDFS deep storage from 2021.
g
the interval in the job
2020-01-01T00:00:00.000Z/2020-12-31T00:00:00.000Z
does not include 2021 -- if you did not do a different job including 2021 as well, perhaps this is why
a
Sorry, my bad, I should have specified that clearly. I had two separate jobs for 2020 and 2021; and in underlying HDFS I see both 2020 and 2021
g
ah, got it
in that case i am not sure what happened! in the most recent log for the 2021 task, do you see any attempts to delete files?
d
You should really debug this inside your YARN. YARN will say if the applicationId that supposed to delete is working or not.
a
Thanks for your pointers, @Didip Kerabat and @Gian Merlino! I did some investigation from the metadata MySQL table as well, and here are my two observations - 1. I had to specify a smaller time range for clearing the underlying HDFS segments. Since the number of segments in the HDFS directory was too large, even when running
hdfs dfs -ls
, I had to specify extra memory to get the output using
export _JAVA_OPTIONS="-Xmx8g
. Looks like when Druid’s Kill process was also trying to access the HDFS location, it was running into memory issues. When I specified smaller windows, the kill task was able to delete the HDFS segments. 2. The kill process did not delete all the segments in my specified interval, because in the MySQL
druid_tasklocks
table, I could still see the listing of these segments. Some segments dated back to 2020, and 2021. (My datasources’ retention is set to P3D).
d
is this kill job manual? The automatic kill job supposed to kill 1 day only
a
Correct, @Didip Kerabat, this kill job was manually triggered.
d
i see i see. You will run into same problem with compaction as well. In general, you have to be careful with really large time range
We automate kill with our airflow job by submitting 1 day at a time to avoid this issue.
a
Got it, do we have any recommendations on the time range we should use?
d
We usually do things 1 day at a time, even though the actual time range is large.
a
Cool! Thanks! 👍
d
it’s best to automate that using something like airflow because it requires some logic: chunking, retry, backoff, etc. We do this for backfill, kill, and compaction.
a
makes sense.. Long time back we also had cron scripts that would issue kill tasks, but somehow they got lost in transition from one infra to the other. Thanks again for all your pointers! They really helped!
On another note, @Didip Kerabat, @Gian Merlino - Where do we specify feature requests? Please let me know if this would make sense - • From the Druid console, when we click on
Delete unused segments (issue kill task)
for a datasource, it starts
api_issued_kill*
task, and the interval for this task is selected from year 1000 to 3000. • But for the data source for which I started this task, has continuous real time ingestion, and batch ingestion. So there is always a lock on some time range between the years 1000 and 3000. As a consequence of which, the kill task stays in
PENDING
state. • So, the feature request here being, can we specify the interval when we want to issue the kill task from the console? Thinking about the recent conversation in this thread, we also have to be mindful of the time ranges.
d
cross post it to #C030K0Z7S1Z?
a
Sure