i’m trying to run a batch ingest job using aws emr...
# troubleshooting
n
i’m trying to run a batch ingest job using aws emr. i’m running a spark submit step like so
Copy code
spark-submit 
--master yarn
--class org.apache.pinot.tools.admin.command.LaunchDataIngestionJobCommand 
s3://<bucket_name>/lib/pinot-all-0.9.0-jar-with-dependencies.jar -jobSpecFile s3://<bucket_name>/jobs/<table_name>/job.yaml
but i’m getting a java.lang.NoSuchMethodException: org.apache.pinot.tools.admin.command.LaunchDataIngestionJobCommand.main
x
I think this command is recently removed, try this one:
Copy code
--class org.apache.pinot.tools.admin.PinotAdministrator
s3://<bucket_name>/lib/pinot-all-0.9.0-jar-with-dependencies.jar LaunchDataIngestionJob -jobSpecFile s3://<bucket_name>/jobs/<table_name>/job.yaml
c
we tried this… and we are getting an exception at
Copy code
Exception in thread "main" java.lang.ExceptionInInitializerError
	at org.apache.pinot.tools.admin.command.StartKafkaCommand.<init>(StartKafkaCommand.java:51)
	at org.apache.pinot.tools.admin.PinotAdministrator.<clinit>(PinotAdministrator.java:98)
	at java.base/java.lang.Class.forName0(Native Method)
	at java.base/java.lang.Class.forName(Class.java:398)
	at org.apache.spark.util.Utils$.classForName(Utils.scala:207)
	at <http://org.apache.spark.deploy.SparkSubmit.org|org.apache.spark.deploy.SparkSubmit.org>$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:924)
	at org.apache.spark.deploy.SparkSubmit.doRunMain$1(SparkSubmit.scala:180)
	at org.apache.spark.deploy.SparkSubmit.submit(SparkSubmit.scala:203)
	at org.apache.spark.deploy.SparkSubmit.doSubmit(SparkSubmit.scala:90)
	at org.apache.spark.deploy.SparkSubmit$$anon$2.doSubmit(SparkSubmit.scala:1047)
	at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:1056)
	at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)
Caused by: java.util.NoSuchElementException
	at java.base/java.util.ServiceLoader$2.next(ServiceLoader.java:1309)
	at java.base/java.util.ServiceLoader$2.next(ServiceLoader.java:1297)
	at java.base/java.util.ServiceLoader$3.next(ServiceLoader.java:1395)
	at org.apache.pinot.tools.utils.KafkaStarterUtils.getKafkaConnectorPackageName(KafkaStarterUtils.java:54)
	at org.apache.pinot.tools.utils.KafkaStarterUtils.<clinit>(KafkaStarterUtils.java:46)
we are not using any kafka config at the moment.. java version is Corretto-11.0.13.8.1… any help will be appreciated.
x
Hmm, which command are you running ?
c
Copy code
spark-submit --master "yarn"--class "org.apache.pinot.tools.admin.PinotAdministrator" s3://<bucket_name>/lib/pinot-all-0.9.0-jar-with-dependencies.jar LaunchDataIngestionJob -jobSpecFile /usr/local/pinot/jobs/<table_name>/job.yaml
a
@Chris Theodore Jayakumar: Just following up here. I am facing a similar issue. any pointers on how you resolved this. Thanks!
c
This was the error we faced when using EMR… we ended up with a dockerized solution in the end to get this working with a custom version of spark & hadoop
a
Gotcha. Thanks. So you never got it working with EMR ?
c
nope..
👍 1