2022/02/08 11:48:43.990 INFO [SecurityManager] [ma...
# troubleshooting
r
2022/02/08 114843.990 INFO [SecurityManager] [main] SecurityManager: authentication disabled; ui acls disabled; users with view permissions: Set(ravishankar); groups with view permissions: Set(); users with modify permissions: Set(ravishankar); groups with modify permissions: Set() Exception in thread "main" java.lang.NoSuchMethodException: org.apache.pinot.tools.admin.command.LaunchDataIngestionJobCommand.main([Ljava.lang.String;) at java.base/java.lang.Class.getMethod(Class.java:2108) at org.apache.spark.deploy.JavaMainApplication.start(SparkApplication.scala:42) at org.apache.spark.deploy.SparkSubmit.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:855) at org.apache.spark.deploy.SparkSubmit.doRunMain$1(SparkSubmit.scala:161) at org.apache.spark.deploy.SparkSubmit.submit(SparkSubmit.scala:184) at org.apache.spark.deploy.SparkSubmit.doSubmit(SparkSubmit.scala:86) at org.apache.spark.deploy.SparkSubmit$$anon$2.doSubmit(SparkSubmit.scala:930) at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:939) at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala) 2022/02/08 114844.084 INFO [ShutdownHookManager] [Thread-0] Shutdown hook called
a
Correct class name is LaunchDataIngestionJob
r
THanks a lor @Aditya. It works. However one small issue- can you guide why am I getting java.lang.RuntimeException java.io.IOException: Failed to create directory: pinot-plugins-dir-0/plugins. THanks a lot...
DO I need to disable ACLs in S3 ?
a
Not sure about this error. I only tried running spark job in local mode The job extracted all plugins in the working directory. It shouldn't be related to s3. Does the user id which is submitting this job have write access to the working directory?
r
@Aditya yes definitely. These segments are created in S3 right?
x
do you have more stacktrace
this seems to be the issue of plugin directory copy
r
@Xiang Fu Here you go:
at org.apache.spark.rdd.RDD.withScope(RDD.scala:385) at org.apache.spark.rdd.RDD.foreach(RDD.scala:970) at org.apache.spark.api.java.JavaRDDLike$class.foreach(JavaRDDLike.scala:351) at org.apache.spark.api.java.AbstractJavaRDDLike.foreach(JavaRDDLike.scala:45) at org.apache.pinot.plugin.ingestion.batch.spark.SparkSegmentGenerationJobRunner.run(SparkSegmentGenerationJobRunner.java:246) at org.apache.pinot.spi.ingestion.batch.IngestionJobLauncher.kickoffIngestionJob(IngestionJobLauncher.java:146) ... 25 more Caused by: java.lang.RuntimeException: java.io.IOException: Failed to create directory: pinot-plugins-dir-0/plugins at org.apache.pinot.plugin.ingestion.batch.spark.SparkSegmentGenerationJobRunner$1.call(SparkSegmentGenerationJobRunner.java:267) at org.apache.pinot.plugin.ingestion.batch.spark.SparkSegmentGenerationJobRunner$1.call(SparkSegmentGenerationJobRunner.java:246) at org.apache.spark.api.java.JavaRDDLike$$anonfun$foreach$1.apply(JavaRDDLike.scala:351) at org.apache.spark.api.java.JavaRDDLike$$anonfun$foreach$1.apply(JavaRDDLike.scala:351) at scala.collection.Iterator$class.foreach(Iterator.scala:891)
Some more:
at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:939) at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala) Caused by: org.apache.spark.SparkException: Job aborted due to stage failure: Task 2 in stage 0.0 failed 1 times, most recent failure: Lost task 2.0 in stage 0.0 (TID 2, localhost, executor driver): java.lang.RuntimeException: java.io.IOException: Failed to create directory: pinot-plugins-dir-0/plugins at org.apache.pinot.plugin.ingestion.batch.spark.SparkSegmentGenerationJobRunner$1.call(SparkSegmentGenerationJobRunner.java:267) at org.apache.pinot.plugin.ingestion.batch.spark.SparkSegmentGenerationJobRunner$1.call(SparkSegmentGenerationJobRunner.java:246) at org.apache.spark.api.java.JavaRDDLike$$anonfun$foreach$1.apply(JavaRDDLike.scala:351) at org.apache.spark.api.java.JavaRDDLike$$anonfun$foreach$1.apply(JavaRDDLike.scala:351) at scala.collection.Iterator$class.foreach(Iterator.scala:891) at org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)
x
I see. Pinot spark job internally requires to create a directory to load pinot jars, do you know if you spark worker has permission to operate on local disk?
Copy code
java.io.IOException: Failed to create directory: pinot-plugins-dir-0/plugins
r
Never needed so, as I have many spark code running on my local. Is there any staging directory I should configure? I saw a parameter in jobspec sample to configure staging
x
stagingDir is for temp data storage
can you try to put the plugin libs into your classpath and remove the pluginDir config when you launch the job
r
No luck, then the first error is coming StartKafkaCOmmand....
I can see that when I run the job, Spark is creating a directory, even if I run as sudo, no permission...
can I specify a user when running spark-submit?
at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128) at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628) at java.base/java.lang.Thread.run(Thread.java:834) Caused by: java.io.IOException: Failed to create directory: pinot-plugins-dir-0/plugins/pinot-minion-tasks at org.apache.pinot.common.utils.TarGzCompressionUtils.untar(TarGzCompressionUtils.java:133)
It must be from pinot.
if (entry.isDirectory()) { if (!outputFile.getCanonicalPath().startsWith(outputDirCanonicalPath)) { throw new IOException(String .format("Trying to create directory: %s outside of the output directory: %s", outputFile, outputDir)); } if (!outputFile.isDirectory() && !outputFile.mkdirs()) { throw new IOException(String.format("Failed to create directory: %s", outputFile)); }
TarGzCompressionUtils.java from Pinot
@Xiang Fu I see that you committed the code SegmentGenerationUtils which uses this directory.
x
yes, it’s using java file api
so pretty weird that this can fail
r
EVen I changed sudo chmod recursively and then re execute, still its failing
x
hmm
r
I am using Mac. DO I need to disable anything
x
oh? so you start a local spark cluster?
r
yes
Thats the issue?
x
not really
which spark version are you using?
r
2.4.8
x
r
Let me go through, thanks
I have seen similar one, another URL in Pinot site
x
This is the last time I ran the spark-submit:
Copy code
export PINOT_VERSION=0.7.0-SNAPSHOT
export PINOT_DISTRIBUTION_DIR=${PINOT_ROOT_DIR}/pinot-distribution/target/apache-pinot-incubating-${PINOT_VERSION}-bin/apache-pinot-incubating-${PINOT_VERSION}-bin
cd ${PINOT_DISTRIBUTION_DIR}
${SPARK_HOME}/bin/spark-submit 
  --class org.apache.pinot.tools.admin.command.LaunchDataIngestionJobCommand 
  --deploy-mode client 
  --conf "spark.driver.extraJavaOptions=-Dplugins.dir=${PINOT_DISTRIBUTION_DIR}/plugins -Dlog4j2.configurationFile=${PINOT_DISTRIBUTION_DIR}/conf/pinot-ingestion-job-log4j2.xml" 
  --conf "spark.driver.extraClassPath=${PINOT_DISTRIBUTION_DIR}/lib/pinot-all-${PINOT_VERSION}-jar-with-dependencies.jar" 
  local://${PINOT_DISTRIBUTION_DIR}/lib/pinot-all-${PINOT_VERSION}-jar-with-dependencies.jar 
  -jobSpecFile /Users/xiangfu/temp/pinot/pinot-s3-test/segmentMetadataSparkIngestionJobSpec.yaml
I think the spark submit is now:
Copy code
${SPARK_HOME}/bin/spark-submit 
  --class org.apache.pinot.tools.admin.PinotAdministrator \\
  --deploy-mode client \\
  --conf "spark.driver.extraJavaOptions=-Dplugins.dir=${PINOT_DISTRIBUTION_DIR}/plugins -Dlog4j2.configurationFile=${PINOT_DISTRIBUTION_DIR}/conf/pinot-ingestion-job-log4j2.xml" \\
  --conf "spark.driver.extraClassPath=${PINOT_DISTRIBUTION_DIR}/lib/pinot-all-${PINOT_VERSION}-jar-with-dependencies.jar" \\
  local://${PINOT_DISTRIBUTION_DIR}/lib/pinot-all-${PINOT_VERSION}-jar-with-dependencies.jar \\
  LaunchDataIngestionJobCommand \\
  -jobSpecFile /Users/xiangfu/temp/pinot/pinot-s3-test/segmentMetadataSparkIngestionJobSpec.yaml
I do see some issue with spark combining jdk11, scala, jackson, will take another look
r
Ohh, it will not work now?
Really, please help. SOmeone needs to look into this. It does not work with Spark latest versions. With Spark 2.4 also, version 0.9.3 is giving weird issues
OK, One clue: I manually untarred the plugin.tar.gz to pinot-plugins-dir-0, now it works. So you may want to check why the code itself is unable to untar.
@Xiang Fu After successful ingestion, when I query airlineStats table, no rows returned. Where did I make a mistake? Result: {"exceptions":[],"numServersQueried":0,"numServersResponded":0,"numSegmentsQueried":0,"numSegmentsProcessed":0,"numSegmentsMatched":0,"numConsumingSegmentsQueried":0,"numDocsScanned":0,"numEntriesScannedInFilter":0,"numEntriesScannedPostFilter":0,"numGroupsLimitReached":false,"totalDocs":0,"timeUsedMs":0,"offlineThreadCpuTimeNs":0,"realtimeThreadCpuTimeNs":0,"segmentStatistics":[],"traceInfo":{},"minConsumingFreshnessTimeMs":0,"numRowsResultSet":0}
x
Hmm, then I feel no data are really ingested
Since you are on local Mac. Can you just try with the standalone ingestion job ?
r
@Xiang Fu All perfect. DO you mind me writing an article in medium on this? Can you please review and give your comments? I would like @Aditya also to be part of it. Let the community be benefitted
x
Sounds good
r
Can I have your email pls
Or pls send a mail to ravishankar.nair@gmail.com