Slackbot
10/12/2022, 11:16 PMSergio Ferragut
10/13/2022, 12:09 AMdruid.indexer.runner.javaOptsArray property? It is what determines the heap size for the tasks. Perhaps increase the heap -Xms -Xmx settings.
Here are the recommendations for Task configuration.
How big are the rows?
Are you using lookups? Lookups live on heap.
And yes, lowering maxRowsInMemory should also help.Kai Sun
10/13/2022, 12:16 AM["-Xms2g", "-Xmx2g", "-XX:MaxDirectMemorySize=2g", "-Dlog4j.configurationFile=/opt/druid/var/peon-log4j2.xml"]Kai Sun
10/13/2022, 12:17 AMKai Sun
10/13/2022, 12:17 AMKai Sun
10/13/2022, 12:22 AMSergio Ferragut
10/13/2022, 12:30 AMKai Sun
10/13/2022, 12:38 AM2022-10-13T00:28:42,652 WARN [task-runner-0-priority-0] org.apache.druid.segment.realtime.appenderator.AppenderatorImpl - Ingestion was throttled for [67,647] millis because persists were pending.
2022-10-13T00:28:42,652 INFO [task-runner-0-priority-0] org.apache.druid.segment.realtime.appenderator.AppenderatorImpl - Persisted rows[212,672] and (estimated) bytes[357,732,600]
2022-10-13T00:28:43,921 WARN [task-runner-0-priority-0] com.amazonaws.services.s3.internal.S3AbortableInputStream - Not all bytes were read from the S3ObjectInputStream, aborting HTTP connection. This is likely an error and may result in sub-optimal behavior. Request only the bytes you need via a ranged GET or drain the input stream after use.
so 357,732,600 * 6 is roughly 2G. The bound is hit by maxBytesInMemory which is 1/6 of total memory of 2G. Here we have only 212k rows. So I will try maybe 100k rows to see if it would work.Kai Sun
10/13/2022, 12:40 AM