Hi team, This is regarding batch ingestion from HD...
# troubleshooting
a
Hi team, This is regarding batch ingestion from HDFS to Offline_Table. After running the following command. bin/pinot-ingestion-job.sh -jobSpecFile /root/hdfsBatchIngestionSpec1.yaml Getting the following logs, segments are not getting created.
Copy code
Trying to create instance for class org.apache.pinot.plugin.ingestion.batch.standalone.SegmentGenerationJobRunner
Initializing PinotFS for scheme hdfs, classname org.apache.pinot.plugin.filesystem.HadoopPinotFS
Unable to load native-hadoop library for your platform... using builtin-java classes where applicable
log4j:WARN No appenders could be found for logger (org.apache.htrace.core.Tracer).
log4j:WARN Please initialize the log4j system properly.
log4j:WARN See <http://logging.apache.org/log4j/1.2/faq.html#noconfig> for more info.
No unit for dfs.client.datanode-restart.timeout(30) assuming SECONDS
No unit for dfs.client.datanode-restart.timeout(30) assuming SECONDS
The short-circuit local reads feature cannot be used because libhadoop cannot be loaded.
successfully initialized HadoopPinotFS
Creating an executor service with 1 threads(Job parallelism: 0, available cores: 24.)
Submitting one Segment Generation Task for <hdfs://nameservice1/data/poc/pinot-ingestion/part-00000-a75dbdce-f8f4-469f-8f70-d412b02b59cb-c000.gz.parquet>
Using class: org.apache.pinot.plugin.inputformat.parquet.ParquetRecordReader to read segment, ignoring configured file format: AVRO
Trying to create instance for class org.apache.pinot.plugin.ingestion.batch.standalone.SegmentTarPushJobRunner
Initializing PinotFS for scheme hdfs, classname org.apache.pinot.plugin.filesystem.HadoopPinotFS
successfully initialized HadoopPinotFS
Start pushing segments: []... to locations: [org.apache.pinot.spi.ingestion.batch.spec.PinotClusterSpec@5d28bcd5] for table poc_test_table
d
Is it possible to share the hdfsBatchIngestionSpec1.yaml with us?
a
BatchIngestionSpec File:
Copy code
executionFrameworkSpec:
    name: 'standalone'
    segmentGenerationJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentGenerationJobRunner'
    segmentTarPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentTarPushJobRunner'
    segmentUriPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentUriPushJobRunner'
    segmentMetadataPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.standalone.SegmentMetadataPushJobRunner'

jobType: SegmentCreationAndTarPush

inputDirURI: '<hdfs://nameservice1/data/poc/pinot-ingestion/>'
outputDirURI: '<hdfs://nameservice1/data/poc/segments/>'

overwriteOutput: true

pinotFSSpecs:

  - scheme: hdfs
    className: org.apache.pinot.plugin.filesystem.HadoopPinotFS
    configs:
       hadoop.conf.path: '/root/hadoop-3.0.0/etc/hadoop/'

recordReaderSpec:
    dataFormat: 'parquet'
    className: 'org.apache.pinot.plugin.inputformat.parquet.ParquetRecordReader'

tableSpec:
    tableName: 'poc_test_table'
    schemaURI: '<http://controller:9000/tables/poc_test_table/schema>'
    tableConfigURI: '<http://controller:9000/tables/poc_test_table>'

pinotClusterSpecs:
    - controllerURI: '<http://controller:9000>'

pushJobSpec:
  pushParallelism: 2
  pushAttempts: 1
j
@Anish Nair Given `hadoop.conf.path`: ‘/root/hadoop-3.0.0/etc/hadoop/’ should contain hadoop XML configuration files such as hdfs-site.xml, core-site.xml. Can you recheck provided path contains config files or
/root/hadoop-3.0.0/etc/hadoop/conf/
is correct path
a
yes, its present. From the logs, it seems , script is able to connect to hadoop cluster. since it has listed the file.
k
I think you might be missing some important logging output, given the log4j warnings. Also, what version of Pinot are you running? Finally, what happens if you try running a job for just segment generation, and do it locally (download the Parquet file, and use local FS for input/output)?
a
Will check logging conf once. We are running Pinot 0.8. Haven't done ingestion with LocalFS. will try
j
Any logs related to Starting/Ending of Segment Index Creator
a
No logging of Segment Index creation.
@jagadesh here is the complete log
j
@Anish Nair From the logs , I can’t see any reference to the ParquetRecordReader reading records from the parquet input file(part-00000-a75dbdce-f8f4-469f-8f70-d412b02b59cb-c000.gz.parquet). If it reads you will see something like this:
Copy code
RecordReader initialized will read a total of xxx records.
The reason could be your input parquet(part-00000-a75dbdce-f8f4-469f-8f70-d412b02b59cb-c000.
gz
.parquet) is a compressed
zip
file. Can you unzip(part-00000-a75dbdce-f8f4-469f-8f70-d412b02b59cb-c000.
parquet
) and test again
a
Hey @jagadesh, we tried earlier with s3 as deep storage. it was able to read gz file and segment got created. Will try unzipping file and run ingestion script.
@jagadesh tried with unzipped parquet file. Same issue. NO change in logs. Also i have verified parquet file by reading it using spark.
j
Would it be possible to download parquet and use local fs for input and output for testing
a
@jagadesh tried with localFs, same logs.
Copy code
Trying to create instance for class org.apache.pinot.plugin.ingestion.batch.standalone.SegmentGenerationJobRunner
Initializing PinotFS for scheme file, classname org.apache.pinot.spi.filesystem.LocalPinotFS
Creating an executor service with 1 threads(Job parallelism: 0, available cores: 24.)
Submitting one Segment Generation Task for file:/root/2021091800/part-00000-a75dbdce-f8f4-469f-8f70-d412b02b59cb-c000.gz.parquet
Using class: org.apache.pinot.plugin.inputformat.parquet.ParquetRecordReader to read segment, ignoring configured file format: AVRO
Unable to load native-hadoop library for your platform... using builtin-java classes where applicable
log4j:WARN No appenders could be found for logger (org.apache.htrace.core.Tracer).
log4j:WARN Please initialize the log4j system properly.
log4j:WARN See <http://logging.apache.org/log4j/1.2/faq.html#noconfig> for more info.
Trying to create instance for class org.apache.pinot.plugin.ingestion.batch.standalone.SegmentMetadataPushJobRunner
Initializing PinotFS for scheme file, classname org.apache.pinot.spi.filesystem.LocalPinotFS
Start pushing segment metadata: {} to locations: [org.apache.pinot.spi.ingestion.batch.spec.PinotClusterSpec@1f86099a] for table poc_test_table
@Mayank @Kamal Chavda any known troubleshoot?
k
Hi @Anish Nair, I've only set-up the ingestion job going to/from s3. Sorry haven't come across this error before. Please do share your resolution. I've kinda started over a few times by recreating the containers and starting everything fresh and it's fixed issues 🤷‍♂️
k
@Anish Nair - my guess is that the Hadoop-based Parquet code is failing, but there’s no log4j.properties file on the classpath, so you don’t get those error messages.