I'm having some of my snappy compressed parquet fi...
# troubleshooting
j
I'm having some of my snappy compressed parquet files failing to convert to segments. It seems to succeed on most of the files, but some are hitting a null pointer exception. Any ideas on what would cause this? I'm not sure if this is a bug, an issue with my underlying data, or an issue with my schema. The stack trace is below:
Copy code
Failed to generate Pinot segment for file - <s3://REDACTED.snappy.parquet>
java.lang.NullPointerException: null
	at shaded.com.google.common.base.Preconditions.checkNotNull(Preconditions.java:770) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
	at org.apache.pinot.segment.local.utils.CrcUtils.getAllNormalFiles(CrcUtils.java:63) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
	at org.apache.pinot.segment.local.utils.CrcUtils.forAllFilesInFolder(CrcUtils.java:52) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
	at org.apache.pinot.segment.local.segment.creator.impl.SegmentIndexCreationDriverImpl.handlePostCreation(SegmentIndexCreationDriverImpl.java:314) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
	at org.apache.pinot.segment.local.segment.creator.impl.SegmentIndexCreationDriverImpl.build(SegmentIndexCreationDriverImpl.java:258) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
	at org.apache.pinot.plugin.ingestion.batch.common.SegmentGenerationTaskRunner.run(SegmentGenerationTaskRunner.java:119) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
	at org.apache.pinot.plugin.ingestion.batch.standalone.SegmentGenerationJobRunner.lambda$submitSegmentGenTask$1(SegmentGenerationJobRunner.java:263) ~[pinot-batch-ingestion-standalone-0.9.3-shaded.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
	at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) [?:?]
	at java.util.concurrent.FutureTask.run(FutureTask.java:264) [?:?]
	at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) [?:?]
	at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) [?:?]
	at java.lang.Thread.run(Thread.java:833) [?:?]
A similar, but different null pointer exception:
Copy code
java.lang.NullPointerException: Cannot read the array length because "<local7>" is null
        at org.apache.pinot.common.utils.TarGzCompressionUtils.addFileToTarGz(TarGzCompressionUtils.java:91) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
        at org.apache.pinot.common.utils.TarGzCompressionUtils.addFileToTarGz(TarGzCompressionUtils.java:92) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
        at org.apache.pinot.common.utils.TarGzCompressionUtils.addFileToTarGz(TarGzCompressionUtils.java:92) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
        at org.apache.pinot.common.utils.TarGzCompressionUtils.createTarGzFile(TarGzCompressionUtils.java:67) ~[pinot-all-0.9.3-jar-with-dependencies.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
        at org.apache.pinot.plugin.ingestion.batch.standalone.SegmentGenerationJobRunner.lambda$submitSegmentGenTask$1(SegmentGenerationJobRunner.java:269) ~[pinot-batch-ingestion-standalone-0.9.3-shaded.jar:0.9.3-e23f213cf0d16b1e9e086174d734a4db868542cb]
        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) [?:?]
        at java.util.concurrent.FutureTask.run(FutureTask.java:264) [?:?]
        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) [?:?]
        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) [?:?]
        at java.lang.Thread.run(Thread.java:833) [?:?]
It also seems like there may be a race condition. I'm using an EBS volume for my scratch space and I'm seeing file not found exceptions. If it helps, I do have
segmentCreationJobParallelism
set to 64.
k
do you have a time column and the value for that column is null in the input?
j
No, I produced my time column from the partition names (hive partitions in this case), making it impossible to be null.
m
@James Mnatzaganian Which parquet record reader are you using? Thereโ€™s a Native one as well.
j
I'm using
org.apache.pinot.plugin.inputformat.parquet.ParquetRecordReader
m
Can you try
ParquetNativeRecordReader
?
j
If my current attempt fails, I'll give that a shot, thanks! I had success doing a smaller amount of data (the same set failed in the larger run), so I've manually chunked the work and I'm trying again ๐Ÿคž
m
I see, let me know.
๐Ÿ‘ 1
j
It's interesting, even running all of the small jobs at one (24 jobs, 1 job per hour for 1 day's worth of data) they seem to be processing just fine (at least the ones that have finished).
m
The NPE above means the output directory was empty.
๐Ÿ‘ 1
j
Here's another strange one -- I had a standalone job "successfully" complete, but there are missing files in S3. Looking at the logs, I see no errors and even see the copy for each file, but for some reason segment 0 is not in S3, but the others are. Edit: Ignore - I think this was an S3 UI bug, waiting a few minutes and refreshing showed the object
Looks like it worked with the multiple smaller jobs ๐ŸŽ‰
m
Thanks for confirming
๐Ÿ‘ 1