Hi everyone, I’m trying to implement a new query a...
# pinot-dev
d
Hi everyone, I’m trying to implement a new query aggregation function for Apache Datasketches CPC sketches. Looking at the implementation for HLL, I came across a dictionary encoded expression. Can anybody give an example of what this is and how it is used? I’d like to ensure that I test it properly, if I need to support it. Thanks
x
Thanks for the contribution! This is a good example, it covers all the scenarios like approximation unique counting on any column to build hll object on the fly or directly aggregate the raw hll objects(bytes) if data are already prepared.
d
Thanks Xiang - I saw that you added integration tests for the Tuple sketch which I’ll add in my PR for this sketch. My colleague will be adding UltraLogLog support soon as well. Do you have any information about dictionary encoded expressions when used with aggregation functions? I’d like to understand whether this use case applies.
x
ack, dictionary is the pinot underly encoding mechanisms, if the column has a dictionary, you will need to read the integer dictionary id first, then look up the corresponding value from the dictionary.
note that a column may be dictionary encoded, or may not be, hence the implementation needs to handle two scenarios.
basically if the function only works on bytes sketch, then likely you don’t need to implement the dictionary encoding, however if you want to run this on any data type, then you need to follow this distinctCountHllAggregationFunction
You can refer to
IntegerTupleSketchAggregationFunction
d
Thanks, this makes sense. I did have a look at the Tuple example and based my integration on all the steps that were followed in prior PRs. One notable comment on the Tuple Aggregation Function is that it only appears to support serialised bytes at the moment, and not dictionaries. This is an understandable design decision as doing otherwise would require access to multiple columns at aggregation time. In terms of the dictionary support, I would like to add it. I was a bit confused about whether the values would have been eagerly materialised at the point this is executed.
x
ic, if you only have serialized sketch, then it has to be non-dictionary encoded byte array, then you can follow IntegerTupleSketchAggregationFunction. Don’t need to implement the dictionary encoding part.
d
I’ll give it a go and update my PR with any improvements. As a future enhancement, we might want to split the functions into different sub-packages or extract out the common bits. Ideally, a plugin model for extension might make it easier, but this is probably on your long term roadmap.
x
agreed, adding new func will involve whole code base/component update, e.g. query parser at pinot broker to query execution on server.
👍 1
ThetaSketch is a good example to support multiple implementations which you can refer to the abstraction it did.
d
The ThetaSketch integration is an interesting example and seems to return all sketches to the broker for post-aggregation. In my case, I might be able to offset some of the costs of expensive merges by also aggregating into lists and only merging once the list exceeds some size threshold. But, to keep things simple, I might just use lists without intermediate merges. These sketches are not nearly as big as Theta sketches so it shouldn’t put as much pressure on the broker.
x
Sounds good, as you said, I would also suggest to implement it first then optimize once you have a baseline.