Hi team, we have been observing the issues with hi...
# general
s
Hi team, we have been observing the issues with historicals throwing out of memory errors and being unresponsive when we execute queries. we have almost 1.5M segments (trying to compact) in 14 historicals each holds 45 % of historical storage (each total storage as 8TB) - total volume of 110TB request you provide any recommendations around this. Below is the error:
Copy code
Error in DRUID historicals - [329987.237s][warning][gc,alloc] qtp1345611807-86[segmentMetadata_[data4g]_ef9299d7-8758-4da4-8845-3e2b3539cff2]: Retried waiting for GCLocker too often allocating 4 words
[329987.237s][warning][gc,alloc] qtp1345611807-99: Retried waiting for GCLocker too often allocating 7 words
2024-03-25T18:27:52,470 WARN [qtp1345611807-99] org.eclipse.jetty.util.thread.strategy.EatWhatYouKill -
java.lang.OutOfMemoryError: Java heap space
2024-03-25T18:27:52,464 ERROR [qtp1345611807-132] com.sun.jersey.spi.container.ContainerResponse - The exception contained within MappableContainerException could not be mapped to a response, re-throwing to the HTTP container
java.lang.OutOfMemoryError: Java heap space
Here is the historicals configuration:
Copy code
Below is the historical config:
historicals:
      nodeType: "historical"
      druid.port: 8088
      nodeConfigMountPath: "/opt/druid/conf/druid/cluster/data/historical"
      replicas: {{ .Values.druid.historical.replicas }}
      runtime.properties: |
        druid.service=druid/historical
        druid.segmentCache.locations=[{\"path\":\"/druid/data/segments\",\"maxSize\":8000000000000}]
        druid.server.maxSize=8000000000000
        druid.processing.buffer.sizeBytes=1000000000
        druid.server.tier={{ .Values.druid.historical.default_tier }}

        druid.historical.cache.useCache=true
        druid.historical.cache.populateCache=true
        druid.cache.sizeInBytes=256000000
        
        druid.query.groupBy.maxOnDiskStorage=30000000000
        
        druid.processing.numMergeBuffers=4
        druid.processing.numThreads=15
  

      extra.jvm.options: |-
        -server
        -Xmx32g
        -Xms32g
        -XX:+UseContainerSupport
        -XX:NewSize=1g
        -XX:MaxDirectMemorySize=80g
        -XX:+UseG1GC
        -Dlog4j.debug=INFO

      resources:
        {{ if .Values.druid.historical.resources }}
          {{ toYaml .Values.druid.historical.resources | nindent 8 }}
          {{ else }}
          limits:
            cpu: "16"
            memory: 96Gi
          requests:
            cpu: "16"
            memory: 96Gi
          {{ end }}

      volumeClaimTemplates:
        - metadata:
            name: data-volume
          spec:
            accessModes:
              - ReadWriteOnce
            resources:
              requests:
                storage: {{ .Values.druid.historical.pvc_storage }}
            storageClassName: netapp-no-snapshot
      volumeMounts:
        - mountPath: /druid/data/segments
          name: data-volume
k
With that many segments, you might need some more heap. I have seen historicals with lots more heap than that. I know https://druid.apache.org/docs/latest/operations/basic-cluster-tuning/#heap-sizing says 24GB but you can go a lot higher once you cross the 32GB mark. Any more memory laying around that you could give these pods, and the heap?
✅ 1