Slackbot
10/10/2023, 8:59 PMClint Wylie
10/11/2023, 5:58 PMClint Wylie
10/11/2023, 5:59 PMUNNEST, and longer term doing a more efficient internal representation of arrays of objectsClint Wylie
10/11/2023, 6:00 PMClint Wylie
10/11/2023, 6:02 PMjson_value expressions which have nearly the same performance as regular flat druid columns because we store nested columns for each field)Jianshu Chi
10/11/2023, 6:07 PMJianshu Chi
10/11/2023, 6:11 PMJianshu Chi
10/11/2023, 6:12 PMClint Wylie
10/11/2023, 6:17 PMClint Wylie
10/11/2023, 6:17 PMJianshu Chi
10/11/2023, 6:18 PMUNNEST? Assuming UNNEST is scanning the original data then unpack to get a unnest table, and then doing the actual aggregation, vs the customized version can unpack and aggregate while the scanning of original data .Clint Wylie
10/11/2023, 6:20 PMClint Wylie
10/11/2023, 6:22 PMweight or other paths from the nested arrays of objects and spit out like ARRAY<LONG> or whatever type the path is, and then that could be unnest to aggregateClint Wylie
10/11/2023, 6:23 PMClint Wylie
10/11/2023, 6:23 PMClint Wylie
10/11/2023, 6:25 PMClint Wylie
10/11/2023, 6:28 PM$[0].weight is a column and $[1].weight is another column), which is kind of sad and inefficient in many ways, but otoh if you know what you want, it still should be relatively efficient to be able to get the values out of them for those specific fieldsClint Wylie
10/11/2023, 6:29 PM$[0].weight is basically a regular long columnJianshu Chi
10/11/2023, 6:29 PMonce i get some more stuff done i imagine using a wildcard paths should be pretty efficient,What is the time line for this?
Clint Wylie
10/11/2023, 6:30 PMClint Wylie
10/11/2023, 6:32 PMJianshu Chi
10/11/2023, 6:32 PMMara Gordon
10/12/2023, 11:00 PM{"driver":"Alice", "age":30, "vehicles_weights":{"Honda Civic":2877, "Toyata prius":3105}, "state": "CA"}
{"driver":"Bob", "age":27, "vehicles_weights":{"Honda CRV":3586, "Honda Civic":2877, "Toyata rav4":3640}, "state": "MA"}
{"driver":"Charlie", "age":40, "vehicles_weights":{"Honda CRV":3586, "Toyata prius":3005}, "state": "MI"}
{"driver":"Susan", "age":29, "vehicles_weights":{"Honda CRV":3586, "Toyata prius":3015}, "state": "MI"}
...}
Can you somehow sum all the weights across vehicles or do you still need to know all the keys and then use each key in json_value expression in this case?
Also Jianshu found that says with druid 24, there are significant impacts on storage and ingestion performance for nested data. Is that still true with druid 25?