Slackbot
10/12/2022, 11:22 PMFJ
10/13/2022, 12:59 AMKaran Kumar
10/13/2022, 11:25 AM-XX:+HeapDumpOnOutOfMemoryError in your druid.indexer.runner.javaOpts or druid.indexer.runner.javaOptsArray that will dump the heap on error, Then we can look at the heap to figure out which large objects are occupying the heap space and tune accordingly.
3. Or if you on 24.0 druid, you can start ingesting with the new sql based ingestion. https://druid.apache.org/docs/latest/multi-stage-query/index.html You could either use the druid console data loader view or could use the insertion sql like
REPLACE INTO perf_data_test
OVERWRITE ALL
SELECT
TIME_PARSE("timestamp") AS __time,
dim1,
dim2,
dim3,
dim4,
sum(measure1),
sum(measure2),
sum(measure3)
FROM TABLE(
EXTERN(
'{"type":"s3","prefixes:["<s3://monc-ra-common-lab1-monitoring-dev1-uswest2-s3/perf_data/>"]}',
'{"type":"json"}',
'[{"name":"timestamp","type":"string"},{"name":"dim1","type":"string"},{"name":"dim2","type":"string"},{"name":"dim3","type":"string"},{"name":"dim4","type":"string"},{"name":"measure1","type":"long"},{"name":"measure2","type":"long"},{"name":"measure3","type":"long"}}]'
)
)
group by dim1,dim2,dim3.dim4
PARTITIONED BY DAY
CLUSTERED BY dim1,dim2,dim3,dum4
A query like this also generates range partitioned segments which are super useful on query time as queries with filters on dim1, dim2. dim3, dim4 filter out segments on the broker itself.Kai Sun
10/13/2022, 6:36 PMFJ
10/13/2022, 6:42 PMFJ
10/13/2022, 6:43 PMFJ
10/13/2022, 6:43 PMKai Sun
10/13/2022, 6:47 PMKai Sun
10/13/2022, 6:50 PMGian Merlino
10/14/2022, 10:46 PM