This message was deleted.
# troubleshooting
s
This message was deleted.
s
spec file -
Copy code
{
  "type": "index_parallel",
  "id": "index_parallel_DATASOURCE_NAME_fjmgkoad_2022-10-07T21:54:05.068Z",
  "groupId": "index_parallel_DATASOURCE_NAME_fjmgkoad_2022-10-07T21:54:05.068Z",
  "resource": {
    "availabilityGroup": "index_parallel_DATASOURCE_NAME_fjmgkoad_2022-10-07T21:54:05.068Z",
    "requiredCapacity": 1
  },
  "spec": {
    "dataSchema": {
      "dataSource": "DATASOURCE_NAME",
      "timestampSpec": {
        "column": "EVENT_TIMESTAMP",
        "format": "auto",
        "missingValue": null
      },
      "dimensionsSpec": {
        "dimensions": [
          {
            "type": "string",
            "name": "EVENT_TIMESTAMP",
            "multiValueHandling": "SORTED_ARRAY",
            "createBitmapIndex": true
          },
          {
            "type": "string",
            "name": "SOME_COLUMN_2",
            "multiValueHandling": "SORTED_ARRAY",
            "createBitmapIndex": true
          },
          {
            "type": "string",
            "name": "SOME_COLUMN_617",
            "multiValueHandling": "SORTED_ARRAY",
            "createBitmapIndex": true
          }
        ],
        "dimensionExclusions": [
          "__time"
        ],
        "includeAllDimensions": false
      },
      "metricsSpec": [],
      "granularitySpec": {
        "type": "uniform",
        "segmentGranularity": "DAY",
        "queryGranularity": {
          "type": "none"
        },
        "rollup": false,
        "intervals": []
      },
      "transformSpec": {
        "filter": null,
        "transforms": []
      }
    },
    "ioConfig": {
      "type": "index_parallel",
      "inputSource": {
        "type": "s3",
        "uris": null,
        "prefixes": [
          "some_s3_location/2022-09-18/"
        ],
        "objects": null,
        "properties": null
      },
      "inputFormat": {
        "type": "parquet",
        "flattenSpec": null,
        "binaryAsString": false
      },
      "appendToExisting": false,
      "dropExisting": false
    },
    "tuningConfig": {
      "type": "index_parallel",
      "maxRowsPerSegment": 5000000,
      "appendableIndexSpec": {
        "type": "onheap",
        "preserveExistingMetrics": false
      },
      "maxRowsInMemory": 1000000,
      "maxBytesInMemory": 0,
      "skipBytesInMemoryOverheadCheck": false,
      "maxTotalRows": null,
      "numShards": null,
      "splitHintSpec": null,
      "partitionsSpec": {
        "type": "dynamic",
        "maxRowsPerSegment": 5000000,
        "maxTotalRows": null
      },
      "indexSpec": {
        "bitmap": {
          "type": "roaring",
          "compressRunOnSerialization": true
        },
        "dimensionCompression": "lz4",
        "metricCompression": "lz4",
        "longEncoding": "longs",
        "segmentLoader": null
      },
      "indexSpecForIntermediatePersists": {
        "bitmap": {
          "type": "roaring",
          "compressRunOnSerialization": true
        },
        "dimensionCompression": "lz4",
        "metricCompression": "lz4",
        "longEncoding": "longs",
        "segmentLoader": null
      },
      "maxPendingPersists": 0,
      "forceGuaranteedRollup": false,
      "reportParseExceptions": false,
      "pushTimeout": 0,
      "segmentWriteOutMediumFactory": null,
      "maxNumConcurrentSubTasks": 4,
      "maxRetry": 1,
      "taskStatusCheckPeriodMs": 1000,
      "chatHandlerTimeout": "PT10S",
      "chatHandlerNumRetries": 5,
      "maxNumSegmentsToMerge": 100,
      "totalNumMergeTasks": 10,
      "logParseExceptions": false,
      "maxParseExceptions": 2147483647,
      "maxSavedParseExceptions": 0,
      "maxColumnsToMerge": -1,
      "awaitSegmentAvailabilityTimeoutMillis": 0,
      "maxAllowedLockCount": -1,
      "partitionDimensions": []
    }
  },
  "context": {
    "forceTimeChunkLock": true,
    "useLineageBasedSegmentAllocation": true
  },
  "dataSource": "DATASOURCE_NAME"
}
e
Firstly, with 24.0.0 there's a SQL-based ingestion mechanism called MSQ that you can use to do ingestion. Doing things with that will work around a lot of the suggestions that I'm going to give you below. So, the recommendation would be to use that. If you cannot use that for some reason, there are some configs on the native jobs that can potentially be tuned to improve this stuff... There is a "splitHintSpec" config that controls the number of files given to a specific task. The defaults for that are 1GiB or 1000 files, whichever limit is hit first. Some computation for that is likely what's causing it to only use 2 tasks. You can try to parallelize it more by setting a spec like
Copy code
"splitHintSpec": {"type": "maxSize", "maxSplitSize": "500MiB", "maxNumFiles": 100}
And that will run more tasks. That said, one common issue that people run into with parquet files coming from other systems is the question of what time ranges exist for data in those files. Other systems like to just randomly distribute rows of data across all of the files where Druid likes data to be a bit more orderly. Sometimes mis-matches in these preferences can cause ingestion to run slower than it needs to. You can control for this as well by either setting a segment granularity wide enough to cover your data (YEAR if crossing multiple years). Or, if you use a partitioning scheme, it will shuffle the data as part of the ingestion to get it all happy and organized for you. "hashed" partitioning doesn't require thought when using it, but "ranged" is almost always better in the long-term. One last thing, with 617 columns, there's also a risk that you are running into poor memory estimations causing lots of spills to disk. If you look at the logs, you should see some lines talking about persisting due to some amount of memory being used. If you can share some of those logs, there might be hints in there.
s
Thanks Eric. Below is the log from one of the subtask.
g
Wanted to chime in to second the tip for using SQL-based ingestion. It isn't just good because it's SQL-based, it's also good because we've totally re-designed how it works to be faster and more robust. If you're on Druid 24, you can use the spec converter tool in the web console (https://druid.apache.org/docs/latest/tutorials/tutorial-msq-convert-spec.html) to get started with trying SQL versions of your classic ingest specs
would love your feedback if you try this!
e
Just reiterating, but you should definitely try out the MSQ based ingestion. From the logs themselves, I'm left wondering if you are running in the same region as your S3 bucket? It looks as if you are spending a good chunk of time just in fetching the parquet files. Specifically, these lines
Copy code
2022-10-07T22:29:11,269 INFO [task-runner-0-priority-0] org.apache.parquet.hadoop.InternalParquetRecordReader - RecordReader initialized will read a total of 10000 records.
2022-10-07T22:29:11,269 INFO [task-runner-0-priority-0] org.apache.parquet.hadoop.InternalParquetRecordReader - at row 0. reading next block
2022-10-07T22:29:11,273 INFO [task-runner-0-priority-0] org.apache.parquet.hadoop.InternalParquetRecordReader - block read in memory in 4 ms. row count = 10000
Those 3 lines indicate the ParquetReader starting to read a file and then being done reading it, it takes all of 4ms. The next line is the next initialized read that happens after pulling down the next Parquet file. It seems like it is consistently taking ~5 seconds to pull down the next file. There are other lines that indicate that it took 166 seconds to upload the segment file as well. If you are not in the local region of your S3 bucket, try to move the data to an s3 bucket that is local. Otherwise, you really should probably switch around to ingest with the SQL-based MSQ engine. Additionally, it looks like you are maybe doing everything as String columns? If so, you might consider either doing some things as numerical and/or you might be interested in the nested column where you can just throw in nested data and Druid will figure out the types for you automatically.
s
My s3 bucket and druid is present in the same location on AWS..
One more query - I have set my druid storage config in common.runtime.properties in all servers as below. Still I see the whole data is cached in my data nodes(historical max cache is full ) along with copy in deep storage. Though I want to query on my whole dataset, does that mean the whole data needs to resides on my local data node instances along with deep storage or is there a way druid can have suppose recent 1 week data in local for faster querying and get other data from deep storage to local cache on runtime ?
Copy code
druid.storage.type=s3
druid.storage.bucket=some.aws.s3.location.druidstorage
druid.storage.baseKey=druid/segments
g
Still I see the whole data is cached in my data nodes(historical max cache is full ) along with copy in deep storage.
yep, this is expected & you can read more here: https://druid.apache.org/docs/latest/design/architecture.html#deep-storage we're looking into options for querying data without having it all precached — this would be a storage requirement vs. performance tradeoff. (as the doc mentions, the precaching is for performance reasons.) if that tradeoff would be good for you then stay tuned for that in future releases 🙂
👍 1
s
So, suppose I need to change my cache location for one local disk to another, in that case, can I just copy the data from the existing cache location to new location on another mounted disk or I can also copy it from deep storage to new location on another mounted disk in case my current cache data is lost for some reason?
g
you can copy it from the existing location to the new one, just be sure to do it while the historical process is not running
you can also not copy it and just start up the historical with an empty cache dir. in that case it will fetch segment files from deep storage automatically
s
So it will fetch everything from deep storage or based on metadata it will pull all the segments of all the available datasources right ?
g
it will be assigned specific segments to pull from the coordinator