Slackbot
02/01/2024, 12:05 PMJohn Kowtko
02/01/2024, 12:42 PMmaxRowsPerSegment down to something much smaller, like 100k, and see if that works. Also drop maxRowsInMemory down to 50k as well.
Fyi I have seen cases where the rows are extremely wide an this setting has to be dropped very low in order to produce "reasonably sized" segments.
Let us know if that works for you. If it does work then let us know what the resulting segment size (in MB) and row counts are. If not, then please provide a task log that we can look at.
Thanks. JohnAshok Kumar Ragupathi
02/01/2024, 12:47 PM2024-02-01T12:38:45,470 ERROR [[index_kafka_ipos_tx_combine_c3c68f05fbc1e1f_jolfmfoh]-publish] com.google.common.util.concurrent.AggregateFuture - Input Future failed with Error
java.lang.OutOfMemoryError: Direct buffer memoryAshok Kumar Ragupathi
02/01/2024, 12:47 PM2024-02-01T12:38:45,480 ERROR [task-runner-0-priority-0] com.google.common.util.concurrent.AggregateFuture - Input Future failed with Error
java.lang.OutOfMemoryError: Direct buffer memoryAshok Kumar Ragupathi
02/01/2024, 12:48 PMCaused by: java.lang.OutOfMemoryError: Direct buffer memoryAshok Kumar Ragupathi
02/01/2024, 12:48 PM2024-02-01T12:38:45,532 INFO [task-runner-0-priority-0] org.apache.druid.indexing.worker.executor.ExecutorLifecycle - Task completed with status: {
"id" : "index_kafka_ipos_tx_combine_c3c68f05fbc1e1f_jolfmfoh",
"status" : "FAILED",
"duration" : 3629767,
"errorMsg" : "java.util.concurrent.ExecutionException: java.lang.OutOfMemoryError: Direct buffer memory\n\tat com.go...",
"location" : {
"host" : null,
"port" : -1,
"tlsPort" : -1
}
}John Kowtko
02/01/2024, 12:55 PMAshok Kumar Ragupathi
02/01/2024, 12:56 PMAshok Kumar Ragupathi
02/01/2024, 12:58 PMAshok Kumar Ragupathi
02/01/2024, 12:59 PMAshok Kumar Ragupathi
02/01/2024, 12:59 PMAshok Kumar Ragupathi
02/01/2024, 12:59 PMAshok Kumar Ragupathi
02/01/2024, 1:00 PMAshok Kumar Ragupathi
02/01/2024, 1:00 PMAshok Kumar Ragupathi
02/01/2024, 1:00 PMmaxMemory 564133888
totalMemory 564133888
freeMemory 383986040
usedMemory 180147848
directMemory 134217728Ashok Kumar Ragupathi
02/01/2024, 1:00 PMAshok Kumar Ragupathi
02/01/2024, 1:01 PMAshok Kumar Ragupathi
02/01/2024, 1:01 PMAshok Kumar Ragupathi
02/01/2024, 1:01 PMAshok Kumar Ragupathi
02/01/2024, 1:02 PMJohn Kowtko
02/01/2024, 1:03 PMrg.apache.druid.segment.realtime.appenderator.StreamAppenderator - Segment[datasourcename_2024-01-25T00:00:00.000Z_2024-01-26T00:00:00.000Z_2024-01-24T23:05:05.938Z_3694] of 167,576 bytes built from 11 incremental persist(s) in 972ms; pushed to deep storage in 599ms
do you have a log statement like the above?Ashok Kumar Ragupathi
02/01/2024, 1:04 PM2024-02-01T12:38:16,829 INFO [task-runner-0-priority-0] org.apache.druid.segment.realtime.appenderator.StreamAppenderator - Persisted rows[14,370] and (estimated) bytes[63,542,638]
2024-02-01T12:38:16,991 INFO [[index_kafka_ipos_tx_combine_c3c68f05fbc1e1f_jolfmfoh]-appenderator-persist] org.apache.druid.segment.realtime.appenderator.StreamAppenderator - Flushed in-memory data for segment[ipos_tx_combine_2023-08-01T00:00:00.000Z_2023-09-01T00:00:00.000Z_2024-01-31T14:34:35.179Z_2] spill[10] to disk in [162] ms (914 rows).
2024-02-01T12:38:17,084 INFO [[index_kafka_ipos_tx_combine_c3c68f05fbc1e1f_jolfmfoh]-appenderator-persist] org.apache.druid.segment.realtime.appenderator.StreamAppenderator - Flushed in-memory data for segment[ipos_tx_combine_2023-06-01T00:00:00.000Z_2023-07-01T00:00:00.000Z_2024-01-31T14:34:36.290Z_2] spill[9] to disk in [90] ms (383 rows).Ashok Kumar Ragupathi
02/01/2024, 1:04 PMAshok Kumar Ragupathi
02/01/2024, 1:05 PM2024-02-01T12:38:40,090 INFO [[index_kafka_ipos_tx_combine_c3c68f05fbc1e1f_jolfmfoh]-appenderator-merge] org.apache.druid.segment.realtime.appenderator.StreamAppenderator - Segment[ipos_tx_combine_2023-12-01T00:00:00.000Z_2024-01-01T00:00:00.000Z_2024-01-31T14:34:33.015Z_2] of 9,354,129 bytes built from 6 incremental persist(s) in 895ms; pushed to deep storage in 458ms. Load spec is: {"type":"hdfs","path":"file:/home/ubuntu/analytics/apache-druid-28.0.0/var/druid/segments/ipos_tx_combine/20231201T000000.000Z_20240101T000000.000Z/2024-01-31T14_34_33.015Z/2_5bfaa59c-778c-4ad8-bca4-5151330b3ca1_index.zip"}
2024-02-01T12:38:42,458 INFO [[index_kafka_ipos_tx_combine_c3c68f05fbc1e1f_jolfmfoh]-appenderator-merge] org.apache.druid.segment.realtime.appenderator.StreamAppenderator - Segment[ipos_tx_combine_2023-02-01T00:00:00.000Z_2023-03-01T00:00:00.000Z_2024-01-31T14:34:38.479Z_2] of 16,242,617 bytes built from 10 incremental persist(s) in 1,608ms; pushed to deep storage in 759ms. Load spec is: {"type":"hdfs","path":"file:/home/ubuntu/analytics/apache-druid-28.0.0/var/druid/segments/ipos_tx_combine/20230201T000000.000Z_20230301T000000.000Z/2024-01-31T14_34_38.479Z/2_2952b459-d0db-4e5f-b219-4241410ed160_index.zip"}
2024-02-01T12:38:44,358 INFO [[index_kafka_ipos_tx_combine_c3c68f05fbc1e1f_jolfmfoh]-appenderator-merge] org.apache.druid.segment.realtime.appenderator.StreamAppenderator - Segment[ipos_tx_combine_2023-07-01T00:00:00.000Z_2023-08-01T00:00:00.000Z_2024-01-31T14:34:35.734Z_2] of 13,196,730 bytes built from 8 incremental persist(s) in 1,281ms; pushed to deep storage in 618ms. Load spec is: {"type":"hdfs","path":"file:/home/ubuntu/analytics/apache-druid-28.0.0/var/druid/segments/ipos_tx_combine/20230701T000000.000Z_20230801T000000.000Z/2024-01-31T14_34_35.734Z/2_85112393-d9be-4eec-9a47-4a16679ecb59_index.zip"}
2024-02-01T12:38:45,465 ERROR [[index_kafka_ipos_tx_combine_c3c68f05fbc1e1f_jolfmfoh]-publish] org.apache.druid.indexing.seekablestream.SeekableStreamIndexTaskRunner - Error while publishing segments for sequenceNumber[SequenceMetadata{sequenceId=0, sequenceName='index_kafka_ipos_tx_combine_c3c68f05fbc1e1f_0', assignments=[], startOffsets={KafkaTopicPartition{partition=1, topic='null', multiTopicPartition=false}=0, KafkaTopicPartition{partition=0, topic='null', multiTopicPartition=false}=0, KafkaTopicPartition{partition=2, topic='null', multiTopicPartition=false}=0}, exclusiveStartPartitions=[], endOffsets={KafkaTopicPartition{partition=1, topic='null', multiTopicPartition=false}=194014, KafkaTopicPartition{partition=0, topic='null', multiTopicPartition=false}=199744, KafkaTopicPartition{partition=2, topic='null', multiTopicPartition=false}=217777}, sentinel=false, checkpointed=true}]
java.lang.OutOfMemoryError: Direct buffer memory
at java.nio.Bits.reserveMemory(Bits.java:175) ~[?:?]Ashok Kumar Ragupathi
02/01/2024, 1:07 PMAshok Kumar Ragupathi
02/01/2024, 1:07 PMAshok Kumar Ragupathi
02/01/2024, 1:08 PMJohn Kowtko
02/01/2024, 1:09 PMAshok Kumar Ragupathi
02/01/2024, 1:09 PMAshok Kumar Ragupathi
02/01/2024, 1:10 PMJohn Kowtko
02/01/2024, 1:11 PMAshok Kumar Ragupathi
02/01/2024, 1:12 PMAshok Kumar Ragupathi
02/01/2024, 1:15 PMAshok Kumar Ragupathi
02/01/2024, 1:15 PMAshok Kumar Ragupathi
02/01/2024, 2:51 PMAshok Kumar Ragupathi
02/01/2024, 2:52 PMAshok Kumar Ragupathi
02/01/2024, 2:52 PMAshok Kumar Ragupathi
02/01/2024, 2:52 PMJohn Kowtko
02/01/2024, 3:12 PMAshok Kumar Ragupathi
02/01/2024, 5:55 PMAshok Kumar Ragupathi
02/01/2024, 6:06 PMAshok Kumar Ragupathi
02/01/2024, 6:07 PMAshok Kumar Ragupathi
02/01/2024, 6:10 PMSergio Ferragut
02/01/2024, 6:25 PMJohn Kowtko
02/01/2024, 6:33 PMAshok Kumar Ragupathi
02/01/2024, 6:34 PMAshok Kumar Ragupathi
02/01/2024, 6:38 PMAshok Kumar Ragupathi
02/01/2024, 6:44 PMAshok Kumar Ragupathi
02/01/2024, 6:45 PMAshok Kumar Ragupathi
02/01/2024, 6:45 PMAshok Kumar Ragupathi
02/01/2024, 6:46 PMAshok Kumar Ragupathi
02/01/2024, 6:46 PMAshok Kumar Ragupathi
02/01/2024, 7:11 PMhttps://files.slack.com/files-pri/T0306CNUA90-F06GUUU3UTW/image.png▾
John Kowtko
02/01/2024, 7:16 PMAshok Kumar Ragupathi
02/02/2024, 5:36 AMAshok Kumar Ragupathi
02/02/2024, 5:36 AMAshok Kumar Ragupathi
02/02/2024, 5:36 AMAshok Kumar Ragupathi
02/02/2024, 6:21 AMAshok Kumar Ragupathi
02/02/2024, 6:30 AMAshok Kumar Ragupathi
02/02/2024, 6:40 AMAshok Kumar Ragupathi
02/02/2024, 8:34 AMmaxRowsPerSegment results in more frequent segment creation but smaller segments, potentially improving query performance. A larger value creates larger segments, which can reduce storage overhead but may affect query performance.
as per chatGPT looks like - having more frequent creation of smaller segment in size which is good to SQL query fetch performance ?? My thought is larger size segment with lesser number of segment creation is good !
which is right ?Ashok Kumar Ragupathi
02/02/2024, 8:35 AMJohn Kowtko
02/02/2024, 12:32 PMAshok Kumar Ragupathi
02/02/2024, 12:35 PMAshok Kumar Ragupathi
02/02/2024, 12:35 PMAshok Kumar Ragupathi
02/02/2024, 12:39 PMAshok Kumar Ragupathi
02/02/2024, 12:40 PMAshok Kumar Ragupathi
02/02/2024, 12:41 PMJohn Kowtko
02/02/2024, 12:43 PMAshok Kumar Ragupathi
02/02/2024, 12:44 PMAshok Kumar Ragupathi
02/02/2024, 12:44 PMAshok Kumar Ragupathi
02/02/2024, 12:45 PM