Hello everyone , I am working backfill data using...
# troubleshooting
d
Hello everyone , I am working backfill data using spark batch ingestion , can we handle duplicate data while backfill , so that it won’t get duplicated in OFFLINE table
m
Can you explain a bit more what you mean by duplication? Do you mean duplication between data that is already loaded into a realtime table and data that you're gonna backfill into an offline table?
m
If you push a segment with name that already exists in Pinot, then the existing segment is overwritten. As long as your daily and backfill jobs produce consistent segment names for a given time period, data will be overwritten.