Hello! Is there a way to configure the ingesting ...
# general
m
Hello! Is there a way to configure the ingesting process such way that ALL data/segments are deleted before data loading (like "truncate table" in postgresql)? I've seen the HTTP DELETE method for flagging unused records and the subsequent kill task, but this would require modifications to the back-end of the our application, which I would prefer to avoid if possible. Or perhaps there's a more efficient way to clear the table before loading data? Thx!
a
1. If you are using index_parallel jobs,
appendToExisting = false
with
dropExisting = true
and input spec interval as ETERNITY / the interval you want to replace would overwrite all the data in that interval. 2. Similarly you could use MSQ with
REPLACE INTO <ds> OVERWRITE ALL ...
This would replace all the existing data with each ingestion You can issue kill tasks separately to clean data from deep storage
m
@Amatya Avadhanula In this case, do I need to know the interval in advance, or can I leave it unspecified and then all intervals will be deleted?
a
How are you ingesting data?
If it's with MSQ, OVERWRITE ALL is the same as specifying eternity as the interval
m
Sorry, I forgot to specify - with index_parallel jobs
a
And you are not interested in appending data, only replacing, is that correct?
m
Correct. I have several scenarios for working with tables where I would need to append data or perform an upsert, but with these specific tables, I need to delete any data and completely overwrite it.
a
If you know that your data lies in a given interval, say
"1000/9999"
just use it in the intervals within the ingestion spec, set
appendToExisting = false
dropExisting = true
If your data is adhoc and you do not know the interval, use this in the intervals field:
Copy code
"-146136543-09-08T08:23:32.096Z/146140482-04-24T15:36:27.903Z"
Just out of curiosity, can't you use MSQ for this ingestion? Or is it because you already have native batch jobs automated?
m
Yes, we already have an implementation with index_parallel jobs which, unfortunately, does not meet the requirements of the technical specification
👍 1
This feature (dropExisting) is marked as experimental - is it possible to use it in a production environment?