Hi all, when we ingest the metadata of `hive`, if ...
# ingestion
g
Hi all, when we ingest the metadata of
hive
, if it encounter an abnormality in some tables, the ingestion will be interrupted. Is there any way to skip these abnormal tables when errors occur?
g
I try this, but occur some error:
Copy code
Error: no such option: --continue-on-error
this is my command:
Copy code
ingest -c hive_archive.yaml --continue-on-error
s
currently it is not present. I was asking if this is what you wanted
sorry for the confusion
g
It’s okay bro, thanks for your help:)
i
A possible workaround is to iteratively adjust deny patterns on tables & schemas that fail and re-try until successful ingestion. It’s not pretty but it does work.
g
maybe it is feasible, but the error is hidden in a schema with a lot of tables, and it takes a long time to deny one😢
i
You can exclude that particular schema if needed. This allows you to ingest tables in other schemas.
m
@green-football-48146 what is the specific exception that gets thrown in your case?
g
@mammoth-bear-12532 Hi, an error similar to the following is an error that will interrupt data ingestion
Copy code
OperationalError: (pyhive.exc.OperationalError) TExecuteStatementResp(status=TStatus(statusCode=3, infoMessages=['*org.apache.hive.service.cli.HiveSQLException:Error while processing statement: FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Parquet does not support date. See HIVE-6384:17:16', 'org.apache.hive.service.cli.operation.Operation:toSQLException:Operation.java:400', 'org.apache.hive.service.cli.operation.SQLOperation:runQuery:SQLOperation.java:238', 'org.apache.hive.service.cli.operation.SQLOperation:runInternal:SQLOperation.java:274', 'org.apache.hive.service.cli.operation.Operation:run:Operation.java:337', 'org.apache.hive.service.cli.session.HiveSessionImpl:executeStatementInternal:HiveSessionImpl.java:439', 'org.apache.hive.service.cli.session.HiveSessionImpl:executeStatement:HiveSessionImpl.java:405', 'org.apache.hive.service.cli.CLIService:executeStatement:CLIService.java:257', 'org.apache.hive.service.cli.thrift.ThriftCLIService:ExecuteStatement:ThriftCLIService.java:501', 'org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement:getResult:TCLIService.java:1313', 'org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement:getResult:TCLIService.java:1298', 'org.apache.thrift.ProcessFunction:process:ProcessFunction.java:39', 'org.apache.thrift.TBaseProcessor:process:TBaseProcessor.java:39', 'org.apache.hive.service.auth.TSetIpAddressProcessor:process:TSetIpAddressProcessor.java:56', 'org.apache.thrift.server.TThreadPoolServer$WorkerProcess:run:TThreadPoolServer.java:286', 'java.util.concurrent.ThreadPoolExecutor:runWorker:ThreadPoolExecutor.java:1149', 'java.util.concurrent.ThreadPoolExecutor$Worker:run:ThreadPoolExecutor.java:624', 'java.lang.Thread:run:Thread.java:748', '*java.lang.UnsupportedOperationException:Parquet does not support date. See HIVE-6384:33:17', 'org.apache.hadoop.hive.ql.io.parquet.serde.ArrayWritableObjectInspector:getObjectInspector:ArrayWritableObjectInspector.java:108', 'org.apache.hadoop.hive.ql.io.parquet.serde.ArrayWritableObjectInspector:<init>:ArrayWritableObjectInspector.java:64', 'org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe:initialize:ParquetHiveSerDe.java:114', 'org.apache.hadoop.hive.serde2.AbstractSerDe:initialize:AbstractSerDe.java:57', 'org.apache.hadoop.hive.serde2.SerDeUtils:initializeSerDeWithoutErrorCheck:SerDeUtils.java:542', 'org.apache.hadoop.hive.metastore.MetaStoreUtils:getDeserializer:MetaStoreUtils.java:390', 'org.apache.hadoop.hive.ql.metadata.Table:getDeserializerFromMetaStore:Table.java:273', 'org.apache.hadoop.hive.ql.metadata.Table:getDeserializer:Table.java:266', 'org.apache.hadoop.hive.ql.exec.DDLTask:describeTable:DDLTask.java:3164', 'org.apache.hadoop.hive.ql.exec.DDLTask:execute:DDLTask.java:380', 'org.apache.hadoop.hive.ql.exec.Task:executeTask:Task.java:214', 'org.apache.hadoop.hive.ql.exec.TaskRunner:runSequential:TaskRunner.java:99', 'org.apache.hadoop.hive.ql.Driver:launchTask:Driver.java:2052', 'org.apache.hadoop.hive.ql.Driver:execute:Driver.java:1748', 'org.apache.hadoop.hive.ql.Driver:runInternal:Driver.java:1501', 'org.apache.hadoop.hive.ql.Driver:run:Driver.java:1285', 'org.apache.hadoop.hive.ql.Driver:run:Driver.java:1280', 'org.apache.hive.service.cli.operation.SQLOperation:runQuery:SQLOperation.java:236'], sqlState='08S01', errorCode=1, errorMessage='Error while processing statement: FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Parquet does not support date. See HIVE-6384'), operationHandle=None)
[SQL: DESCRIBE FORMATTED `model`.`community_dpd_order_feature`]
(Background on this error at: <http://sqlalche.me/e/13/e3q8>)