This message was deleted.
# troubleshooting
s
This message was deleted.
f
Kai, are you using Druid 24?
k
Couple of things to try: 1. Yes you could lower the maxRowsInMemory further 2. Try setting :
-XX:+HeapDumpOnOutOfMemoryError
in your
druid.indexer.runner.javaOpts
or
druid.indexer.runner.javaOptsArray
that will dump the heap on error, Then we can look at the heap to figure out which large objects are occupying the heap space and tune accordingly. 3. Or if you on 24.0 druid, you can start ingesting with the new sql based ingestion. https://druid.apache.org/docs/latest/multi-stage-query/index.html You could either use the druid console data loader view or could use the insertion sql like
Copy code
REPLACE INTO perf_data_test
OVERWRITE ALL
SELECT
  TIME_PARSE("timestamp") AS __time,
  dim1,
  dim2,
  dim3,
  dim4, 
  sum(measure1),
  sum(measure2),
  sum(measure3)
FROM TABLE(
    EXTERN(
      '{"type":"s3","prefixes:["<s3://monc-ra-common-lab1-monitoring-dev1-uswest2-s3/perf_data/>"]}',
      '{"type":"json"}',
      '[{"name":"timestamp","type":"string"},{"name":"dim1","type":"string"},{"name":"dim2","type":"string"},{"name":"dim3","type":"string"},{"name":"dim4","type":"string"},{"name":"measure1","type":"long"},{"name":"measure2","type":"long"},{"name":"measure3","type":"long"}}]'
    )
  )
group by dim1,dim2,dim3.dim4
PARTITIONED BY DAY
CLUSTERED BY dim1,dim2,dim3,dum4
A query like this also generates range partitioned segments which are super useful on query time as queries with filters on dim1, dim2. dim3, dim4 filter out segments on the broker itself.
k
Thanks for the info. Not using 2.4 yet.
f
@Kai Sun strongly recommend using Druid 24
there’s a reason why we upgraded the version from 0.23 to 24
there’s a major set of changes
k
Understood, I attended the Summit in September. For now, building a performance test environment. One of the purpose of this perf test bed is to certify new releases. It is in our roadmap.
By the way, just to let you know @Karan Kumar, by reducing further maxRowsInMemory, no more OOM at this time. Thx for the advice.
🙌 1
g
glad you got it fixed! I wanted to second the tip for SQL-based ingestion, btw. It isn't just good because it's SQL-based, it's also good because we've totally re-designed how it works to be faster and more robust. There's a lot less potential for OOM conditions than with native
👍 1