Slackbot
01/03/2024, 1:50 PMJohn Kowtko
01/03/2024, 2:20 PMTim Rannikko
01/03/2024, 2:55 PMlatest(), and it's something I can use. InfluxDB offers a similar approach for handling duplicates.
However, the main challenge for us seems to be updating pre-aggregates – it's a critical issue. Am I correct in assuming that there isn't a straightforward solution for updating older data on ingestion time rollups across different segments without resorting to application code?
I quite like the concept of ingestion time rollups, If there were just a way to automatically track and update segments when older data is added, that would eliminate our need for materialized views.John Kowtko
01/03/2024, 3:01 PMTim Rannikko
01/03/2024, 4:03 PMTimestamp | Insert Time | RId | Value | Status
----------------------------------------------------------------------
2024-01-01T00:00:00Z | 2024-01-01T00:00:00Z | R1 | 1.0 | Estimated
2024-01-01T00:15:00Z | 2024-01-01T00:15:00Z | R1 | 4.0 | OK
...
2024-02-01T00:00:00Z | 2024-02-01T00:00:00Z | R1 | 1.0 | OK
2024-02-01T00:15:00Z | 2024-02-01T00:15:00Z | R1 | 2.0 | OK
However, in February (Insert Time 2024-02-01T001500Z), we receive a duplicate value for January, needing a recalculation of the January rollup. So in this scenario I assume I need re-ingest the entire January raw data using the latest values or is there other way to periodically re-compute the time slice for monthly rollup accordingly?
Timestamp | Insert Time | RId | Value | Status
----------------------------------------------------------------------
2024-01-01T00:00:00Z | 2024-01-01T00:00:00Z | R1 | 1.0 | Estimated
2024-01-01T00:00:00Z | 2024-02-01T00:15:00Z | R1 | 2.2 | OK <-- NEW DUPLICATE ROW
2024-01-01T00:15:00Z | 2024-01-01T00:15:00Z | R1 | 4.0 | OK
...
2024-02-01T00:00:00Z | 2024-02-01T00:00:00Z | R1 | 1.0 | OK
2024-02-01T00:15:00Z | 2024-02-01T00:15:00Z | R1 | 2.0 | OKJohn Kowtko
01/03/2024, 6:16 PMTim Rannikko
01/04/2024, 12:09 AMJohn Kowtko
01/08/2024, 2:14 PM__time interval, and replace an interval in its entirety ... so if one record changes that affects a given rollup interval, unless you want to have the queries always resolving duplicates in rollups then you would need to reingest/overwrite that interval with a complete set of updated data.
Since your __time value represents the original date of the record being affected, you would have to scan the entire detail table to look for records that have a new insert_time so you could process those rollup intervals. In order to avoid a full table scan you could set up a duplicate feed to go into a "rollup staging" table, put the current date/time in the __time field and carry forward the original __time value in another field ... this would provide you with a somewhat ordered queue of time intervals that require reprocessing.
As for driving the batch jobs, yes you are right Druid does not have an internal job schedule as yet, although it has been discussed internally so the need is known. For now I have heard of cron and Apache Airflow as popular job scheduling tools to use.
Thanks. John