Backfill question -- we have a large REALTIME tabl...
# troubleshooting
t
Backfill question -- we have a large REALTIME table (~900GB/day). Due to a configuration error (ZK heap size too low) we lost some data because the Kafka retention was less than the time to fix the bug. This has me thinking of way to fill in missing data in the future for disaster recovery. We have all the raw data sitting in Parquet files in our data lake. My initial thought was to regenerate the segments with missing data (they are east to identify). Is it possible to upload (refresh) REALTIME segments, assuming the event time range is correct (there would be more events in the replacement segment)? Or do I have to use a HYBRID table and either populate the OFFLINE segments myself or use Pinot managed Offline flows?
m
Right now, data push to realtime table is disabled, and needs a managed offline flow. But afaik, Uber team is working on backfill support for RT tables. Is this still the case @Yupeng Fu?
y
right. we are working on such backfill pipeline in flink