Hello everyone, we have been facing an issue with ...
# general
s
Hello everyone, we have been facing an issue with the overlord process has not responding when we enable the compaction for our data sources. Currently, we have close to 30 + data sources (98% of them ingesting the data through the Kafka and remaining are static tables). We are dealing close to 110TB of data in Druid and ingestion of 25Billion Events everyday. Due to heterogeneous nature of data and with our configurations it’s creating approx 120k segments every day and we have total of 1.5 M segments, we are trying to enable the compaction and its causing the lag for the Kafka ingestion stream and compaction is very slow. We have 350 MM which are running 1 task each and we set the compaction ratio to 0.5. Can you please suggest any configurations to make overlord process robust and does not cause the ingestion lag increase. I have tried setting Kafka ingestion to async true. Currently using Druid 27.0.0
a
1. How do you enable compaction while streaming ingestion is active for the datasource? 2. When do you observe lag? 3. How long does the task take to complete? Is it much greater than the configured taskDuration? 4. What is the segment granularity and the late message rejection period for these supervisors? 5. Do you have late arriving data (currentTime - 10 * segmentGranularity)? 6. What is the
task/action/run/time
like?
s
1. We set a compaction config to perform the compaction with delay of 1day 2. After a while (we observed lag got increased to 8B after a day we enable the compaction for all data sources and interesting thing is, the data source which increased the lag doesn’t have compaction enabled, because this is huge stream compare to others) 3. Task is being completed in 3 hours and duration is set to 2 hrs + completion timeout set to 3 hrs 4. Segment granularity to hour and late rejection to 3 days 5. Late arrival data is possible to for last 7 days 6. Will have to check Other observation, the same stream was constantly maintaining the lag to less than 20M and sometimes to 0. But after enable the compaction to other data sources its causing this. Found that overload going in to unhealthy state in kubernetea dashboard and Druid ui is going in proxy mode
a
1. What is the
task/action/run/time
like?
This is an important metric, especially if you have late arrival data with several tasks. Do you see the lag go up periodically and come down on its own gradually and repeat?
s
As far as I observe, it always grow up and bit stable some times but didn’t observe its going down
Observe that tasks are being failed with below error
a
Ah it may be an issue with slow handoff
Could you also please check the max value of
ingest/handoff/time
?
s
Yea slow handoff. I see it waited for so long and exited. Unfortunately I don’t have dashboard to check them right now, will see if I can get the metric.
Do you think increasing the completion timeout help to wait more and handoff the segment ?
a
This patch may help you if ingest/handoff/time is slow: https://github.com/apache/druid/pull/15952
s
And how do we ensure to make the handoff faster and atleast not to fail ? Increasing the completion timeout ?
Is the PR yet to merge it to release ? When are we expecting ?
a
Polled and found %,d segments in the database
could you search for this log in your leader coordinator?
It would help to know what frequency this log occurs with
Is the PR yet to merge it to release ? When are we expecting ?
Druid 30 possibly
s
Do you think increasing the completion timeout help to wait more and handoff the segment properly ?
a
Yes, it would help
Polled and found %,d segments in the database
could you search for this log in your leader coordinator?
Please check this
And also the max values of
ingest/handoff/time
,
coordinator/global/time
,
task/action/run/time
s
Checking
Yes. Found it and it’s returning 1.5 M segments in the database
a
Great, at what frequency do you see this log repeat?
s
12 mins based on the log
a
I see, that's quite high. It should ideally happen every minute. But it is still not enough to cause a timeout of 2 hours
s
But in the coordinator we configured the Druid.coordinator.period=600S and start delay of 300S
a
Druid.coordinator.period=600S
Why was this change made?
s
Was suggested in the community before where metadata queries were taking long time if I remember correctly
a
Is smartSegmentLoading enabled?
You can check this in your coordinator's dynamic config
s
Yes
a
How many segments in the historical load queues? Is this count decreasing?
You can see these beside the historical entries in the servers tab