kaivalya apte
02/17/2022, 11:21 AMDISTINCTCOUNT, I have a use case where I want a distinct count based on a column, which has very high cardinality. 900million rows out of which 890million might be distinct. DISTINCTCOUNT function fails because of OOM. I see that DISTINCTCOUNT is implemented using a HashSet which will load all the distinct values in memory. I was thinking if a bloom filter with low false positivity rate may help here. Thoughts?Peter Pringle
02/17/2022, 11:23 AMRichard Startin
02/17/2022, 11:23 AMdistinctCountThetaSketch or distinctCountHLL functions. They each have parameters to trade resource consumption for accuracykaivalya apte
02/17/2022, 11:24 AMkaivalya apte
02/18/2022, 10:55 AMid and date column to count the distinct, which is not supported by HLL or sketches. Is there any other way I could do that? I am considering following options:
⢠combine id and date into a single column just used for distinct queries.
⢠using a partial upsert config with an additional column count which can be incremented whenever there is a duplicate event and then just count the rows will give me distinct rows and for total rows I could SUM() on count column.
Any other ideas?Richard Startin
02/18/2022, 11:00 AMRichard Startin
02/18/2022, 11:01 AMkaivalya apte
02/18/2022, 11:08 AMkaivalya apte
02/18/2022, 11:08 AMRichard Startin
02/18/2022, 11:22 AMkaivalya apte
02/18/2022, 11:23 AMRichard Startin
02/18/2022, 11:25 AMRichard Startin
02/18/2022, 11:26 AMkaivalya apte
02/18/2022, 11:30 AM