This message was deleted.
# general
s
This message was deleted.
y
Load-based Rate Limiting Solution 1. sync key metrics about the current load of data nodes to Redis 2. brokers feedback the load of data nodes based on the current load status
a
y
@Abhishek Agarwal Thanks for your input. Yes, we have enabled the high/low laning strategy. We only assign 20% of the threads to process heavy queries but still encounter the data node overwhelmed issue.
j
There are a couple more options I can think of: • If these are different datasources you could place them on different Data Tiers, which will physically separate them to run on different Data Nodes. Even if it is the same datasource, if you have the resources available to double-ingest that datasource you could place one copy on each tier. • if the heavy queries can run for several seconds or more, then consider using Query from Deep Storage (https://druid.apache.org/docs/latest/api-reference/sql-api#query-from-deep-storage) which will bypass the Historicals completely Thanks. John
y
@John Kowtko Thanks for sharing your thoughts. Yes, we also used the tiering data nodes solution when we were using large shared cluster architecture but now we split the shared cluster into small clusters. In a small cluster, we usually assume all datasources have same priority. The MSQ task engine looks promising! But just another question comes into my mind is that - Is it possible to schedule heavy queries to MSQ engine only when the data nodes are heavy loaded? It seems we still need broker to feedback the load of data nodes.
j
MSQ uses a different endpoint, so your application would have to know when to use this vs the regular Druid SQL API endpoint. As for figuring out when the query system is heavily loaded. I know there are Clarity metrics available that you could look at for Jetty threads and queuing, and also metricas for segment scan waits, but I don't know offhand if there is an API call you could make to the Broker/Router to ask how many queries are currently running ....
👍 1