Hi dev-team, I have moved this discussion from the...
# pinot-dev
d
Hi dev-team, I have moved this discussion from the general channel but I wanted to follow up on this issue discussion - particularly about the amount of effort involved in supporting Apache Datasketches in Pinot from storage, aggregation, indexing and query. Our use case relies on us being able to ingest pre-built sketches as well as create sketches from raw data during pre-aggregation (rollup in Druid). As a newcomer, I am willing to help with this should the changes be approachable.
We can continue the discussion on Github but initially I wanted to get some idea of the effort to support this feature - we are currently evaluating Druid which supports this feature out the box but I’ve been interested in some of the features that Pinot offers for user-facing analytics.
x
cc: @Mayank
k
are you thinking of replacing existing implementation with apache datasketches or adding new aggregation functions
d
I was asking mainly about broad support for different sketch types - my understanding of the current implementation is that it is computed at query time?
To be more specific about our use case - we have a lot of queries that involve distinct counts, and the identifier concerned is very high cardinality. In our current database, and potentially with Druid, we would precompute the sketches and combine them at query time. Is this antithetical to the Pinot usage model? Ie. would it be better to store raw rows and could an index actually store the sketch itself?
m
You can most certainly precompute sketches and store in Pinot. Additionally, with StarTree index, you can even pre materialize some of the cubes with sketches to make your queries even faster
d
You can most certainly precompute sketches and store in Pinot
This is good, but comes at a cost. Would this be supported currently by the Query Language? My understanding of the current sketch support is to compute them on the fly, whereas if sketches were stored it would be reading the raw sketch data and combining them at query time.
Additionally, with StarTree index, you can even pre materialize some of the cubes with sketches to make your queries even faster
I’m interested in this model for Pinot. Precomputing sketches seems to be at odds with real-time use cases and by definition involves pre-aggregation. Would a star tree index support creating a sketch off a high-cardinality dimension such as an identifier, with some secondary dimensions such as browser, country etc?
Sorry if these questions do not belong here - I have been evaluating Druid but believe that Pinot is the better choice for its ability to constrain query latency. Perhaps my use case does not apply to Pinot and I need to shift my mindset to realtime?
m
No, these are great questions and thanks for asking
Pinot supports both pre computed sketches as well as computing sketches at read time.
y
Hi @Mayank I just found this thread, which is 9 months old, yet I really need to figure out how to compute sketches at read/ingestion time. I have not found any documentation on how to perform something like this, yet I really need it. My use case is pre-computing n-percentile using KLL Sketches, and maing the percentile computation over a large dataset that much faster. Do you have any examples laying around which would point me in the right direction ?
m
Essentially you create sketch column in your input data. Only requirement is that the sketch binary is serialized using the same library that Pinot will use to deserialze