Hi. We accidentaly removed deep storage data from ...
# general
i̇
Hi. We accidentaly removed deep storage data from hdfs. Is there any method to copy loaded data from historicals to deep storage again?
s
I don't think there is a procedure for that. Thinking through it...The segments are kept unzipped in the historicals. You would need to zip them, rebuild the folder structures in HDFS and copy them over. The druid_segments table on the metadata DB should provide the list of segments files and their expected locations on HDFS.
i̇
Thanks. I've tried to reingest from druid but it gives error because there is no segment on deep storage. Is there any workaround so that I can reingest data from druid itself? Is it possible with multistage queries? I don't have too much knowledge about that concept.
s
In theory it is possible to reingest as long as segments have not been rebalanced by the coordinator. What is the segment availability for the tables involved? I'm not sure how it will behave if some segments are not available. You should also turn off or slow down the coordinator cycle until you resolve this because rebalancing of segments could request that historical drop some segments . You can turn it off with dynamic coordinator property
pauseCoordination=true
For the reingest you could do:
Copy code
REPLACE INTO table OVERWRITE ALL
SELECT * FROM table
PARTITIONED BY <segment granularity that you already have>
This will use MSQ to read segments from historicals and publish them into the same table. If you have any secondary partitioning you should also specify that as in:
Copy code
REPLACE INTO table OVERWRITE ALL
SELECT * FROM table
CLUSTERED BY <secondary partitioning columns>
PARTITIONED BY <segment granularity that you already have>
i̇
Yes. Datasources seems fully available. I will try this. Thanks
l
The latter suggestion shouldnt work IMO. MSQ loads the segments from the deep storage, which would fail in this case.
👍 1
s
Oh…right!
a
I was thinking the same as well - that MSQ task would read from deep storage and not add load on the historicals. And this works fine as the segments are immutable. I have another idea. I haven't tried this and I know nothing about Druid internals so please treat it with caution. Can this be done using JDBC input source in batch ingestion? The docs say that MySql & Postgres is supported but I assume Avatica JDBC would be supported as well since the Avatica jars are on the classpath of all the processes. So theoretically, you could launch a batch ingestion job which reads from a JDBC input source and specify Druid's Avatica JDBC URL itself in the connection params. Would this ingestion task then query the data from the broker and ingest back into Druid? Druid Experts - would this work? Regards, AR.
s
Straight JDBC SQL input source came up in a different thread. Unfortunately, that extension doesn't exist yet. Currently only the supported metadata DBs have a SQL Input Source.