Hey folks, I’m curious if anyone has ever tried to...
# pinot-dev
e
Hey folks, I’m curious if anyone has ever tried to build or thought of building about a minion task (or similar) that could help serve the use case of mass data export from Pinot. I know this isn’t a primary focus of Pinot (understandably, with the focus on real-time ingestion and low latency querying instead), but I have a use case where mass export would be very useful. There are alternative options, but in particular when
upsert
is employed and data consistency across data export + aggregate results is important, serving results from the same source of data (i.e. Pinot) would be ideal. The high-level concept would be something like: given a SQL query without any aggregations, generate minion tasks in the form of segment name + ID_SET of matching doc IDs based on the provided SQL query; each task would then have minions download the segment from server/deepstore, pluck out the matching documents based on the segments corresponding ID_SET, apply transformations from the SQL query for provided projects, and then write those resulting rows back to deepstore in CSV form (or parquet, or configurable form, whatever haha). This approach would place a lot of the heavy-lifting of disk seeks/long-running queries related to mass export onto minions that might otherwise be concerning for servers to handle while also handling other queries. Is there anything similar to this today? Is there any reason this would be a very bad idea?
m
Thank @Evan Galpin for bringing this up. Async reporting queries use cases are something they keep coming. Having a solution (minion based or otherwise) would be really helpful. Would love to brainstorm on a proposal
🚀 1
s
+1, good idea. Can you create an issue and start a discussion on that?
🙌 1
e
Will do!