Hi Krishna,
Yes, TopN queries should benefit significantly from the LIMIT clause, because the limit is pushed down to the compute node for initial filtering of results. The second paragraph on this page mentions some of this:
https://druid.apache.org/docs/latest/querying/topnquery
k
Krishna
02/28/2024, 9:43 PM
Thanks @John Kowtko
Krishna
02/28/2024, 10:55 PM
i think we are using threshold already in the query . so it is getting pushed to historicals .
Krishna
02/28/2024, 11:14 PM
Any memory parameters to improve performance we r seeing high latency when it has sketches and lot of dimensions and metrics in the query
Krishna
03/26/2024, 5:04 PM
@John Kowtko Ingestion was hourly and then we do daily rollup . what we noticed is TopN queries on daily rollup for a month range is taking around 17 secs . but when i didn't roll up to daily for the same month , same TopN query is returning in < 5 secs . segment size for daily rollup is around 1GB and segment size for hourly rollup is around 700 MB . any thing i need to look at to improve the query against daily rollup datasource ?
j
John Kowtko
03/29/2024, 12:03 AM
I don't know much about memory alocation for sketches ... hopefully someone else can answer to that.
If HOUR segments product better query time than DAY segments, then maybe you need the extra parallelism allowed by more segments being accessed in the query. I have seen that in other clusters when the segment scans are cpu intensive.