Maytas Monsereenusorn
04/23/2025, 7:36 PMGian Merlino
04/23/2025, 9:54 PMGian Merlino
04/23/2025, 9:54 PMpriority. a low priority query can use the entire pool, but if a higher priority query comes along, then it will take over the poolGian Merlino
04/23/2025, 9:55 PMlane is a concept that IMO doesn't fully make sense on Historicals; its main purpose is to control which queries are accepted for running, but by the time a query gets to a Historical, it's already been accepted (by the Broker that sent it)Gian Merlino
04/23/2025, 9:55 PMMaytas Monsereenusorn
04/24/2025, 12:24 AMare you wanting low-priority queries to not use the entire processing thread pool? (maybe restrict to 8 threads on a 32 thread machine)?Yes. Exactly.
and that’s adequately protected byThat’s not entirely true though as Druid historical does not preempt already running processing threads. For example, if you set query timeout to 5 minutes. Most of your queries are sub-second. In this example, our system only have query A and B. Query A always run for 100ms. Query B is an expensive query that will run for >5minutes (and timeout) was put in low lane. Query B ran run before query A and did have a lane capacity on the low lane. Broker ran Query B. Query B then ran on Historical taking up all available processing threads. Each processing thread will be taken up for 5 minutes. Query A was execute right after query b but now has to wait for 5 minutes for processing threads on historicals. Users of query A then file a ticket because a query that is suppose to run in 100ms is not taking 5minutes/timeout.. a low priority query can use the entire pool, but if a higher priority query comes along, then it will take over the poolpriority
If lane is a concept to protect/reserve resource. Why don’t we do it all the way? i.e. if we want to guarantee that a high priority query will always have resource reserved for it, then I think it make sense to apply it to historical as well.is a concept that IMO doesn’t fully make sense on Historicals; its main purpose is to control which queries are accepted for running, but by the time a query gets to a Historical, it’s already been accepted (by the Broker that sent it)lane
Clint Wylie
04/24/2025, 12:27 AMClint Wylie
04/24/2025, 12:28 AMMaytas Monsereenusorn
04/24/2025, 12:32 AMMaytas Monsereenusorn
04/24/2025, 12:34 AMClint Wylie
04/24/2025, 1:38 AMClint Wylie
04/24/2025, 1:39 AMClint Wylie
04/24/2025, 1:42 AMClint Wylie
04/24/2025, 1:43 AMClint Wylie
04/24/2025, 1:44 AMmax(1, someQueryCostModifier)Clint Wylie
04/24/2025, 1:45 AMClint Wylie
04/24/2025, 1:47 AMClint Wylie
04/24/2025, 1:48 AMClint Wylie
04/24/2025, 1:49 AMClint Wylie
04/24/2025, 1:51 AMClint Wylie
04/24/2025, 1:52 AMClint Wylie
04/24/2025, 1:52 AMGian Merlino
04/24/2025, 2:17 AMGian Merlino
04/24/2025, 2:19 AMGian Merlino
04/24/2025, 2:20 AMGian Merlino
04/24/2025, 2:21 AMGian Merlino
04/24/2025, 2:22 AMIt is difficult-to-impossible for either Druid, or the application querying it, to know ahead of time what priority a query should have. The best approach is typically to use very simple heuristics such as “queries that process more segments should be lower priority” or “queries for the ‘download’ feature should be lower priority”. But the amount of time to process a segment varies greatly based on things like value cardinality, filter selectivity, complexity of expressions, etc. This limits the success that people have with prioritization. To improve this, we want to:
• implement a quota system that tracks CPU time used by ‘quota-carrying entities’. Such entities may be users (i.e. identity as determined by an authenticator) or may be groups of users.
• prioritize incoming queries based on recent CPU usage by the relevant ‘quota-carrying entity’.
• reprioritize currently-running queries, which may necessitate canceling and re-running them, if the new priority causes them to switch to a lower lane that is currently full.
Ben Krug
05/01/2025, 7:27 PMMaytas Monsereenusorn
05/01/2025, 11:24 PMI have only really seen one situation where this doesn’t work well— if processing of a single segment takes a very long time for some reason. Sometimes this can happen, although it’s rare in my experience for it to be more than a few seconds. So, in this situation, it could take a few seconds for the higher priority query to “take over”. But it shouldn’t take minutes
If it does take minutes something seems “wrong” to meWe do have some queries that takes on avg 1-2 minutes to process a single segment (attached metric on P50, P90, P99
query/segment/time )
Maybe the query is the problem but we provide Druid as a platform and allow our end user to ingest whatever they want (they write the ingestionSpec) and query whatever they want (they write the query)…and sometimes those query are not the fastest/most efficient 😞Maytas Monsereenusorn
05/01/2025, 11:36 PMSELECT
DATE_TRUNC('day', __time) "timestamp",
is_x,
1000.0 * SUM(CASE WHEN foo IN ('a','b') THEN 0 ELSE CASE WHEN is_x = False THEN bar / 0.9 ELSE baz END END) / SUM(cnt) FILTER (WHERE is_y != 'true') "value"
FROM ds
WHERE __time BETWEEN '2024-01-01T00:00:00.000Z' AND '2025-04-21T23:59:59.999Z'
GROUP BY 1,2 ORDER BY 1,2
LIMIT 2000000
The time range contains about 20,000 segments, each segment about 3.5-5.5 million rows 150MB-250MB.
Note that removing/rewriting without the CASE fixes the slowness in the query but the lack of guardrail / resource management on the historical resulted in this query taking over the historical for ~2minutes (blocking everything else).Maytas Monsereenusorn
05/01/2025, 11:39 PMimplement a quota systemThis sounds reasonable to me. I guess similar to like Hadoop YARN Fair Scheduler.
Maytas Monsereenusorn
05/01/2025, 11:41 PMa tricky thing to balance here is that deprioritized heavy queries can also become so starved and increase the likelyhood of timing out if there are too many interactive queries because they spend all of their time just getting skipped in line in the queue waiting for segments to be processedI would argue that starving the slow queries (which usually points to badly written query, query that could be rewritten, optimized, and query that would eventually timeout anyway) is better than making all the other queries slow. One unhappy user is better than 100s of unhappy users haha
Maytas Monsereenusorn
05/01/2025, 11:44 PMIt is difficult-to-impossible for either Druid, or the application querying itWould we be able to use segment metadata stats for this (size, cardinality, type, etc)? Maybe type of aggregators/filters? If a query can be vectorized or not? Or would those be too expensive to compute upfront?
Maytas Monsereenusorn
05/06/2025, 7:56 PMGian Merlino
05/07/2025, 12:35 AMGian Merlino
05/07/2025, 12:37 AMWould we be able to use segment metadata stats for this (size, cardinality, type, etc)? Maybe type of aggregators/filters? If a query can be vectorized or not?I think the issue is there's too many variables, especially when a query starts to involve multiple columns. Maybe AI could figure it out? Only half-joking?
Gian Merlino
05/07/2025, 12:37 AMWe do have some queries that takes on avg 1-2 minutes to process a single segment🤯 OMG. no wonder you have problems with the current priority system.
Gian Merlino
05/07/2025, 12:39 AMMaytas Monsereenusorn
05/07/2025, 8:44 PMDo you actually need these queries to finish? I wonder if a simple solution could be having a timeout at the level of individual-segment processing. i.e., query-level timeout of 5 minutes, but segment-level timeout of 15 seconds. If it takes more than 15 sec to process a single segment then something is probably really inefficient with the query.I am also thinking/agreeing that we probably don’t need these expensive queries to finish. (Similar to what I mentioned at https://apachedruidworkspace.slack.com/archives/C030CMF6B70/p1746142880974949?thread_ts=1745436989.786489&cid=C030CMF6B70). These queries will often timeout on the query-level timeout anyway. These queries are also often indication of something that a user is doing wrong (badly written query). Having these queries fail and those users optimizing/fixing the queries is much better than letting it run (and impacting all the other users/queries on the Druid cluster). Do we already have a timeout at the level of individual-segment processing? I haven’t thought of this solution but now that you mentioned it, I think it makes a lot of sense. I also think it should be pretty simple to configure and use (compare to query lane thresholds tuning, etc).
Gian Merlino
05/07/2025, 11:12 PMGian Merlino
05/07/2025, 11:13 PMMaytas Monsereenusorn
05/16/2025, 7:15 AMGian Merlino
06/17/2025, 5:39 AMJesse Tuglu
06/17/2025, 5:51 AMJesse Tuglu
07/21/2025, 8:56 PMJesse Tuglu
07/21/2025, 8:59 PMGian Merlino
08/02/2025, 9:10 AMGian Merlino
08/02/2025, 9:12 AMGian Merlino
08/02/2025, 9:13 AMGian Merlino
08/02/2025, 9:15 AMJesse Tuglu
08/04/2025, 6:55 AM