This message was deleted.
# troubleshooting
s
This message was deleted.
s
Have you configured the
druid.indexer.runner.javaOptsArray
property? It is what determines the heap size for the tasks. Perhaps increase the heap
-Xms
-Xmx
settings. Here are the recommendations for Task configuration. How big are the rows? Are you using lookups? Lookups live on heap. And yes, lowering
maxRowsInMemory
should also help.
k
The indexer peon java property is following:
Copy code
["-Xms2g", "-Xmx2g", "-XX:MaxDirectMemorySize=2g", "-Dlog4j.configurationFile=/opt/druid/var/peon-log4j2.xml"]
one row is around 800bytes. But in druid segment, the compression ratio is around 1:8. So 5million row per segment is around 500Mbytes.
No look ups.
I am trying ingesting 50G file per peon/task. Is this file too large? Do we have a ball park idea how long does it take? @Sergio Ferragut
s
I am not familiar with how the file size impacts memory usage. But that could be the culprit. It is worth a test.. One large file can only be processed by one task. Splitting into multiple files will also help because you can increase the maxNumConcurrentSubTasks in your spec and process the ingestion in parallel and speed it up.
k
Reducing to 500,000 rows in memory failed again with GC error after close to 30 mins. Here, I see this log:
Copy code
2022-10-13T00:28:42,652 WARN [task-runner-0-priority-0] org.apache.druid.segment.realtime.appenderator.AppenderatorImpl - Ingestion was throttled for [67,647] millis because persists were pending.
2022-10-13T00:28:42,652 INFO [task-runner-0-priority-0] org.apache.druid.segment.realtime.appenderator.AppenderatorImpl - Persisted rows[212,672] and (estimated) bytes[357,732,600]
2022-10-13T00:28:43,921 WARN [task-runner-0-priority-0] com.amazonaws.services.s3.internal.S3AbortableInputStream - Not all bytes were read from the S3ObjectInputStream, aborting HTTP connection. This is likely an error and may result in sub-optimal behavior. Request only the bytes you need via a ranged GET or drain the input stream after use.
so 357,732,600 * 6 is roughly 2G. The bound is hit by
maxBytesInMemory
which is 1/6 of total memory of 2G. Here we have only 212k rows. So I will try maybe 100k rows to see if it would work.
Also in the log, I saw 10 such persist before OOM. That also probably means 3.5G files would just be ok.