Evan Galpin
01/19/2024, 11:30 PMupsert is employed and data consistency across data export + aggregate results is important, serving results from the same source of data (i.e. Pinot) would be ideal.
The high-level concept would be something like: given a SQL query without any aggregations, generate minion tasks in the form of segment name + ID_SET of matching doc IDs based on the provided SQL query; each task would then have minions download the segment from server/deepstore, pluck out the matching documents based on the segments corresponding ID_SET, apply transformations from the SQL query for provided projects, and then write those resulting rows back to deepstore in CSV form (or parquet, or configurable form, whatever haha). This approach would place a lot of the heavy-lifting of disk seeks/long-running queries related to mass export onto minions that might otherwise be concerning for servers to handle while also handling other queries.
Is there anything similar to this today? Is there any reason this would be a very bad idea?Mayank
Subbu Subramaniam
01/21/2024, 6:17 PMEvan Galpin
01/22/2024, 6:01 PM