Hello Was there any change to `ingestFromFile` end...
# troubleshooting
j
Hello Was there any change to
ingestFromFile
endpoint between 0.8.0 and 0.9.1 ? We're now getting HTTP 500 with: Maybe an issue related to null support ?
m
are there nulls for the time column in that table?
Copy code
@Override
  public String generateSegmentName(int sequenceId, @Nullable Object minTimeValue, @Nullable Object maxTimeValue) {
    Preconditions.checkArgument(
        minTimeValue == null || isValidSegmentName(minTimeValue.toString()));
    Preconditions.checkArgument(
        maxTimeValue == null || isValidSegmentName(maxTimeValue.toString()));
    return JOINER
        .join(_segmentNamePrefix, minTimeValue, maxTimeValue, _segmentNamePostfix, sequenceId >= 0 ? sequenceId : null);
  }
line 53 is the
minTimeValue
check
j
Thanks for the quick answer @Mark Needham, checking !
Nope
Dates are all in 2020
Copy code
"dateTimeFieldSpecs": [
    {
      "dataType": "STRING",
      "format": "1:SECONDS:SIMPLE_DATE_FORMAT:yyyy-MM-dd'T'HH:mm:ss.SSSZ",
      "granularity": "1:SECONDS",
      "name": "timeString"
    },
    {
      "dataType": "STRING",
      "format": "1:SECONDS:SIMPLE_DATE_FORMAT:yyyy-MM-dd'T'HH:mm:ss.SSSZ",
      "granularity": "1:SECONDS",
      "name": "eventTimeString"
    }
  ],
m
oh I've seen that bit of the code do weird stuff with date strings. Let me see if I can find my example where that was happening
thankyou 1
j
Okay !
j
By default, the ingestion job spec will try to generate a segment name using the SimpleSegmentNameGenerator and expects minTimeValue to be a number. In your case minValue will be yyyy-MM-dd'T'HHmmss.SSSZ and will therefore fail the precondition check. Try adding the following:
segmentNameGeneratorSpec:
type: normalizedDate
configs:
segment.name.prefix: 'mytablename'
exclude.sequence.id: true
j
@Jeff Moszuti Thank you very much, I'll try that now
Do you think it affect REALTIME tables too ?
j
https://docs.pinot.apache.org/configuration-reference/job-specification#segment-name-generator-spec.
normalizedDate
- use this type when the time column in your data is in the String format instead of epoch time
👍 1
j
mmm... now that I think of it, we've seen some batch ingestion work (but not via API) with 0.9.1 without this change
j
Not sure, I haven't use REALTIME tables yet!
😉 1
j
But I'll try it 🙂
Oh, since I'm ingesting via the API (https://docs.pinot.apache.org/basics/data-import/batch-ingestion#ingestfromfile), I don't think we've got a jobSpecification per se I've tried
batchConfigMapStr={{"inputFormat":"json", "segmentNameGeneratorSpec.type": "normalizedDate"}}
Doesn't seem to change anything
m
I think for real time the following should do the same:
Copy code
{
  "tableName": "crimes",
  "tableType": "OFFLINE",
  "segmentsConfig": {
    "replication": 1,
    "timeColumnName": "Date",
    "timeType": "DAYS",
    "retentionTimeUnit": "DAYS",
    "retentionTimeValue": 365
  },
  "tenants": {
    "broker": "DefaultTenant",
    "server": "DefaultTenant"
  },
  "tableIndexConfig": {
    "loadMode": "MMAP"
  },
  "ingestionConfig": {
    "batchIngestionConfig": {
      "segmentIngestionType": "APPEND",
      "segmentIngestionFrequency": "DAILY",
      "batchConfigMaps": [
        {
          "segmentNameGenerator.type": "normalizedDate"
        }
      ]
    }
  },
  "metadata": {}
}
this bit:
Copy code
"batchConfigMaps": [
          {"segmentNameGenerator.type":  "normalizedDate"}
        ]
🙌 2
j
Interesting, thank for looking that up
m
that will have it use a different segment name generator
I had the same issue a few weeks ago!
😄 1
and IIRC this was how I worked around it
j
Yeeeah
That worked !
Thank you @Mark Needham & @Jeff Moszuti This community rocks 😄