Slackbot
10/14/2022, 7:06 AMSourav Saraf
10/14/2022, 7:07 AM{
"type": "index_parallel",
"id": "index_parallel_DATASOURCE_NAME_fjmgkoad_2022-10-07T21:54:05.068Z",
"groupId": "index_parallel_DATASOURCE_NAME_fjmgkoad_2022-10-07T21:54:05.068Z",
"resource": {
"availabilityGroup": "index_parallel_DATASOURCE_NAME_fjmgkoad_2022-10-07T21:54:05.068Z",
"requiredCapacity": 1
},
"spec": {
"dataSchema": {
"dataSource": "DATASOURCE_NAME",
"timestampSpec": {
"column": "EVENT_TIMESTAMP",
"format": "auto",
"missingValue": null
},
"dimensionsSpec": {
"dimensions": [
{
"type": "string",
"name": "EVENT_TIMESTAMP",
"multiValueHandling": "SORTED_ARRAY",
"createBitmapIndex": true
},
{
"type": "string",
"name": "SOME_COLUMN_2",
"multiValueHandling": "SORTED_ARRAY",
"createBitmapIndex": true
},
{
"type": "string",
"name": "SOME_COLUMN_617",
"multiValueHandling": "SORTED_ARRAY",
"createBitmapIndex": true
}
],
"dimensionExclusions": [
"__time"
],
"includeAllDimensions": false
},
"metricsSpec": [],
"granularitySpec": {
"type": "uniform",
"segmentGranularity": "DAY",
"queryGranularity": {
"type": "none"
},
"rollup": false,
"intervals": []
},
"transformSpec": {
"filter": null,
"transforms": []
}
},
"ioConfig": {
"type": "index_parallel",
"inputSource": {
"type": "s3",
"uris": null,
"prefixes": [
"some_s3_location/2022-09-18/"
],
"objects": null,
"properties": null
},
"inputFormat": {
"type": "parquet",
"flattenSpec": null,
"binaryAsString": false
},
"appendToExisting": false,
"dropExisting": false
},
"tuningConfig": {
"type": "index_parallel",
"maxRowsPerSegment": 5000000,
"appendableIndexSpec": {
"type": "onheap",
"preserveExistingMetrics": false
},
"maxRowsInMemory": 1000000,
"maxBytesInMemory": 0,
"skipBytesInMemoryOverheadCheck": false,
"maxTotalRows": null,
"numShards": null,
"splitHintSpec": null,
"partitionsSpec": {
"type": "dynamic",
"maxRowsPerSegment": 5000000,
"maxTotalRows": null
},
"indexSpec": {
"bitmap": {
"type": "roaring",
"compressRunOnSerialization": true
},
"dimensionCompression": "lz4",
"metricCompression": "lz4",
"longEncoding": "longs",
"segmentLoader": null
},
"indexSpecForIntermediatePersists": {
"bitmap": {
"type": "roaring",
"compressRunOnSerialization": true
},
"dimensionCompression": "lz4",
"metricCompression": "lz4",
"longEncoding": "longs",
"segmentLoader": null
},
"maxPendingPersists": 0,
"forceGuaranteedRollup": false,
"reportParseExceptions": false,
"pushTimeout": 0,
"segmentWriteOutMediumFactory": null,
"maxNumConcurrentSubTasks": 4,
"maxRetry": 1,
"taskStatusCheckPeriodMs": 1000,
"chatHandlerTimeout": "PT10S",
"chatHandlerNumRetries": 5,
"maxNumSegmentsToMerge": 100,
"totalNumMergeTasks": 10,
"logParseExceptions": false,
"maxParseExceptions": 2147483647,
"maxSavedParseExceptions": 0,
"maxColumnsToMerge": -1,
"awaitSegmentAvailabilityTimeoutMillis": 0,
"maxAllowedLockCount": -1,
"partitionDimensions": []
}
},
"context": {
"forceTimeChunkLock": true,
"useLineageBasedSegmentAllocation": true
},
"dataSource": "DATASOURCE_NAME"
}Eric Tschetter
10/14/2022, 7:36 AM"splitHintSpec": {"type": "maxSize", "maxSplitSize": "500MiB", "maxNumFiles": 100}
And that will run more tasks.
That said, one common issue that people run into with parquet files coming from other systems is the question of what time ranges exist for data in those files. Other systems like to just randomly distribute rows of data across all of the files where Druid likes data to be a bit more orderly. Sometimes mis-matches in these preferences can cause ingestion to run slower than it needs to. You can control for this as well by either setting a segment granularity wide enough to cover your data (YEAR if crossing multiple years). Or, if you use a partitioning scheme, it will shuffle the data as part of the ingestion to get it all happy and organized for you. "hashed" partitioning doesn't require thought when using it, but "ranged" is almost always better in the long-term.
One last thing, with 617 columns, there's also a risk that you are running into poor memory estimations causing lots of spills to disk. If you look at the logs, you should see some lines talking about persisting due to some amount of memory being used. If you can share some of those logs, there might be hints in there.Sourav Saraf
10/14/2022, 9:16 AMGian Merlino
10/14/2022, 11:01 PMGian Merlino
10/14/2022, 11:01 PMEric Tschetter
10/17/2022, 5:06 AM2022-10-07T22:29:11,269 INFO [task-runner-0-priority-0] org.apache.parquet.hadoop.InternalParquetRecordReader - RecordReader initialized will read a total of 10000 records.
2022-10-07T22:29:11,269 INFO [task-runner-0-priority-0] org.apache.parquet.hadoop.InternalParquetRecordReader - at row 0. reading next block
2022-10-07T22:29:11,273 INFO [task-runner-0-priority-0] org.apache.parquet.hadoop.InternalParquetRecordReader - block read in memory in 4 ms. row count = 10000
Those 3 lines indicate the ParquetReader starting to read a file and then being done reading it, it takes all of 4ms. The next line is the next initialized read that happens after pulling down the next Parquet file. It seems like it is consistently taking ~5 seconds to pull down the next file.
There are other lines that indicate that it took 166 seconds to upload the segment file as well.
If you are not in the local region of your S3 bucket, try to move the data to an s3 bucket that is local. Otherwise, you really should probably switch around to ingest with the SQL-based MSQ engine. Additionally, it looks like you are maybe doing everything as String columns? If so, you might consider either doing some things as numerical and/or you might be interested in the nested column where you can just throw in nested data and Druid will figure out the types for you automatically.Sourav Saraf
10/18/2022, 12:16 PMSourav Saraf
10/18/2022, 12:20 PMdruid.storage.type=s3
druid.storage.bucket=some.aws.s3.location.druidstorage
druid.storage.baseKey=druid/segmentsGian Merlino
10/18/2022, 4:41 PMStill I see the whole data is cached in my data nodes(historical max cache is full ) along with copy in deep storage.yep, this is expected & you can read more here: https://druid.apache.org/docs/latest/design/architecture.html#deep-storage we're looking into options for querying data without having it all precached — this would be a storage requirement vs. performance tradeoff. (as the doc mentions, the precaching is for performance reasons.) if that tradeoff would be good for you then stay tuned for that in future releases 🙂
Sourav Saraf
10/19/2022, 7:14 AMGian Merlino
10/19/2022, 7:18 AMGian Merlino
10/19/2022, 7:19 AMSourav Saraf
10/20/2022, 5:17 AMGian Merlino
10/21/2022, 4:32 AM