Slackbot
01/11/2023, 6:36 AMBrandon Tan
01/11/2023, 6:41 AMMichael Tsibelman
01/11/2023, 6:44 AMMichael Tsibelman
01/11/2023, 6:45 AMBrandon Tan
01/11/2023, 6:47 AMfake the time columnMichael Tsibelman
01/11/2023, 6:50 AMBrandon Tan
01/11/2023, 6:51 AMMichael Tsibelman
01/11/2023, 6:52 AMBrandon Tan
01/11/2023, 6:53 AMBrandon Tan
01/11/2023, 6:55 AMAbhishek Agarwal
01/11/2023, 11:40 AMMichael Tsibelman
01/11/2023, 3:21 PMErin Monday
01/11/2023, 4:50 PMErin Monday
01/11/2023, 5:31 PMMark Herrera
01/11/2023, 5:56 PMSergio Ferragut
01/11/2023, 6:37 PMREPLACE OVERWRITE WHERE __time = <hash value>.... While unlikely, one possible concern here is hash collisions such that you would need to include all data for any matching hash values when running the REPLACE.Michael Tsibelman
01/11/2023, 6:49 PMSergio Ferragut
01/11/2023, 7:36 PMREPLACE .... CLUSTERED BY and cluster on the most commonly used filtering dimension(s). If you can get 500MB segment sizes (or so) that would give you ~500 segments which are plenty to take advantage of parallelism and when filtering on the clustered dimensions, you'll also get segment pruning. Very doable.Sergio Ferragut
01/11/2023, 7:37 PMMichael Tsibelman
01/11/2023, 7:41 PM