This message was deleted.
# general
s
This message was deleted.
j
Oh, there is an indexing period that determines the actual cycle ... the kill.period is just the "minimum" amount of time that can be between kill tasks ... if your indexer period is, say, 30 minutes, then chances are you won't hit the second 30 minute cycle right on the nose, so you will have to wait until the next 30 minute cycle. So the way to improve this is to set the indexing period to 5 or 10 minute intervals. e.g. here is one example from a cluster I work with:
Copy code
druid.coordinator.period.indexingPeriod=PT300S
druid.coordinator.kill.on=true
druid.coordinator.kill.period=PT30M
druid.coordinator.kill.durationToRetain=P60D
druid.coordinator.kill.maxSegments=1000
āœ… 1
j
Awesome @John Kowtko, thanks.
Hmm, according to this I should not be experiencing this issue: https://github.com/apache/druid/pull/14831
j
I'm not sure this PR (at least by the description) changes the behavior that I understand. As I understand it the indexer wakes up periodically to check things and starts jobs when they need to be started. So by default it may do this once every 30 minutes. Separately, the kill period isn't specifying a cycle interval, it is specifying "do not run more than two kill cycles within N minutes of each other". So if the indexer wakes up and does it's checking every 30 minutes, and if the last kill cycle was 59 minutes and 55 seconds ago, and kill.period is 60 minutes, then the indexer would not run another kill task this time around. But in another 30 minutes it will wake up, the last kill job was 89 minutes and 55 seconds ago, which is > 60 minutes, so it will run another kill job.
j
Well, previous to this change Druid triggered hourly and after it's now not triggering hourly
j
I am thinking that statement was more about preventing the kill task from accidentally running on the same interval twice in a row. Compaction has the same issue (and that may be getting fixed as well)
j
I was going off the release notes: Release Note • The value of
druid.coordinator.kill.period
can now be as low as
druid.coordinator.period.indexingPeriod
. It was earlier required to be strictly greater than the
indexingPeriod
. • The leader Coordinator now keeps track of the last submitted
kill
task for a given datasource to avoid submitting duplicate kill tasks.
j
Maybe it was just timing, i.e. after the version update the cycle "clocks" are triggering at slightly different times. having kill.period=1HR, and default indexing period is 30M, sounds like a recipe for kill to just miss the indexing cycle and have to wait for the next one.
j
So I guess I should set
druid.coordinator.kill.period
=
druid.coordinator.period.indexingPeriod
=
PT30M
in order to trigger it hourly now? šŸ™‚
j
I would set the kill.period to 30 min, and the indexing period to 5 min ... let coordinator check frequently ... then it will fire the kill jobs every 30-35 minutes.
j
Actual Kill Period =
druid.coordinator.kill.period
+
druid.coordinator.period.indexingPeriod
šŸ™‚
j
I thought that is the highest it could be ...
j
Oh, it was:
druid.coordinator.kill.period
>
druid.coordinator.period.indexingPeriod
But now it is:
druid.coordinator.kill.period
>=
druid.coordinator.period.indexingPeriod
But that change made the actual kill period longer
Actually, according to the PR we no longer have to set
druid.coordinator.kill.period
a
@JRob what John said šŸ™‚
druid.coordinator.kill.period
suggests a weak lower bound on how often the coordinator will execute kill tasks, which depends on the last kill time kill was spawned. Maybe there is scope to improve the docs for this configuration? Feel free to raise a PR if you agree: https://druid.apache.org/docs/28.0.1/configuration/#:~:text=druid.-,coordinator,-.kill.period
j
As a developer this makes sense. As a user, I would just want to be able to set the kill period and expect it will work. It's not clear from the current config options how to achieve that. As I thought about how to write this into the documentation it started getting really weird, like "`druid.coordinator.kill.period` sets a minimum on the actual kill period. The actual kill period depends on the indexing period so the actual kill period will be a multiple of
druid.coordinator.period.indexingPeriod
greater than
druid.coordinator.kill.period
so, for example, to achieve an hourly kill cycle with the default indexing period of PT30M, one could set
druid.coordinator.kill.period
to a value between PT31M and PT59M, but this will vary depending on how long indexing tasks take on your Druid cluster." I do think that the new "auto cleaner" feature added in 28.0 is a nice addition and likely what I will be using now, especially since it's the new default setting.
j
I agree a simpler "cycle time" per utility would be easier to get your head around. I imagine this indexer + utility model was created with many utilities in mind, allowing you to reduce cycle time for all utilities by adjusting only one parameter.
āœ… 1
a
Yeah, agreed as well. I’m not sure why
druid.coordinator.kill.period
exists. That config predates some of the guard rails that were added more recently, so it must have been for a conservative measure. But now, I think we could just piggy back on
druid.coordinator.period.indexingPeriod
. Fwiw, you could set
druid.coordinator.kill.period
and
druid.coordinator.period.indexingPeriod
to the same values and things will be slightly more predictable. I will see if we can deprecate
druid.coordinator.kill.period
to make things more simple here.
āœ… 1
j
Agreed