Ashish Kumar
12/21/2022, 4:17 AMAshish Kumar
12/21/2022, 4:17 AM{
"schemaName": "geohashAreaMapDim",
"dimensionFieldSpecs": [
{
"name": "geohash",
"dataType": "STRING"
},
{
"name": "area",
"dataType": "STRING"
},
{
"name": "cityId",
"dataType": "LONG"
},
{
"name": "countryId",
"dataType": "LONG"
}
],
"primaryKeyColumns": [
"cityId",
"geohash"
]
}
Table:
{
"OFFLINE": {
"tableName": "geohashAreaMapDimOffline_OFFLINE",
"tableType": "OFFLINE",
"segmentsConfig": {
"schemaName": "geohashAreaMapDim",
"replication": "1",
"segmentPushType": "REFRESH",
"minimizeDataMovement": false
},
"tenants": {
"broker": "DefaultTenant",
"server": "DefaultTenant"
},
"tableIndexConfig": {
"invertedIndexColumns": [],
"rangeIndexVersion": 2,
"autoGeneratedInvertedIndex": false,
"createInvertedIndexDuringSegmentGeneration": false,
"loadMode": "MMAP",
"enableDefaultStarTree": false,
"enableDynamicStarTreeCreation": false,
"aggregateMetrics": false,
"nullHandlingEnabled": false,
"optimizeDictionaryForMetrics": false,
"noDictionarySizeRatioThreshold": 0
},
"metadata": {},
"quota": {
"storage": "200M"
},
"ingestionConfig": {
"batchIngestionConfig": {
"segmentIngestionType": "REFRESH",
"segmentIngestionFrequency": "DAILY"
},
"transformConfigs": [
{
"columnName": "cityId",
"transformFunction": "city_id"
},
{
"columnName": "countryId",
"transformFunction": "country_id"
}
]
},
"isDimTable": true
}
}
Batch Ingestion Config:
executionFrameworkSpec:
name: 'spark'
segmentGenerationJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.spark3.SparkSegmentGenerationJobRunner'
segmentTarPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.spark3.SparkSegmentTarPushJobRunner'
segmentUriPushJobRunnerClassName: 'org.apache.pinot.plugin.ingestion.batch.spark3.SparkSegmentUriPushJobRunner'
extraConfigs:
stagingDir: '<s3://bucket-name/midas-offline-staging/dimension/geohash.area_map>'
jobType: SegmentCreationAndTarPush
inputDirURI: '<s3://bucket-name/midas-offline-parquet/dimension/geohash.area_map/>'
includeFileNamePattern: 'glob:**/*.parquet'
outputDirURI: '<s3://bucket-name/midas-offline/dimension/geohash.area_map/>' # '/pinot/temp'
#auth-token
authToken: '<something>'
# overwriteOutput: Overwrite output segments if existed.
overwriteOutput: true
pinotFSSpecs:
- scheme: s3
className: org.apache.pinot.plugin.filesystem.S3PinotFS
configs:
region: 'ap-southeast-1'
recordReaderSpec:
dataFormat: 'parquet'
className: 'org.apache.pinot.plugin.inputformat.parquet.ParquetNativeRecordReader'
tableSpec:
tableName: 'geohashAreaMapDimOffline'
schemaURI: '<http://stg-mimic-pinot.coban.stg.com:9000/tables/geohashAreaMapDimOffline/schema>'
tableConfigURI: '<http://stg-mimic-pinot.coban.stg.com:9000/tables/geohashAreaMapDimOffline>'
segmentNameGeneratorSpec:
type: normalizedDate
configs:
segment.name.prefix: 'geohashAreaMapDimOffline_batch'
exclude.sequence.id: true
pinotClusterSpecs:
- controllerURI: '<http://stg-mimic-pinot.coban.stg.com:9000>'
pushJobSpec:
pushParallelism: 2
pushAttempts: 2
pushRetryIntervalMillis: 1000Ashish Kumar
12/21/2022, 4:18 AMAshish Kumar
12/21/2022, 4:18 AMAshish Kumar
12/21/2022, 4:20 AMAshish Kumar
12/21/2022, 4:20 AMSeunghyun
12/21/2022, 4:31 AMSeunghyun
12/21/2022, 4:35 AMgeohashAreaMapDimOffline ?Seunghyun
12/21/2022, 4:36 AM1. change schema geohashAreaMapDim -> geohashAreaMapDimOffline
2. change the "schema" field in table config to geohashAreaMapDimOfflineAshish Kumar
12/21/2022, 4:44 AMhow many files did you have for the input?2 parquet files in the input.
Can you double check if you have schema named asNo, geohashAreaMapDim existsgeohashAreaMapDimOffline
Seunghyun
12/21/2022, 4:45 AMSeunghyun
12/21/2022, 4:46 AMAshish Kumar
12/21/2022, 4:46 AMSeunghyun
12/21/2022, 5:18 AMSeunghyun
12/21/2022, 5:18 AMAshish Kumar
12/21/2022, 5:19 AMsegmentNameGeneratorSpec:
type: normalizedDate
configs:
segment.name.prefix: 'geohashAreaMapDimOffline_batch'
exclude.sequence.id: trueSeunghyun
12/21/2022, 5:19 AMexclude.sequence.id: true -> falseSeunghyun
12/21/2022, 5:19 AMAshish Kumar
12/21/2022, 5:22 AMSeunghyun
12/21/2022, 6:49 AMAshish Kumar
12/21/2022, 7:02 AMSeunghyun
12/21/2022, 10:25 PMold_0, old_1, old_2 -> new_0, new_1 . But this has not been added yet because the new protocol will now assume the time bucket based replacement (i.e. replace data for 1day at a time).
Do you expect to backfill data frequently and the number of files for the original data can change frequently?Ashish Kumar
12/22/2022, 3:51 AMDo you expect to backfill data frequently and the number of files for the original data can change frequently?yes, for now I am thinking of automating it using Pinot REST APIs... something like for backfills/original data changes.. (1) Delete all the existing segments with prefix! (2) Run the ingestion which will create new segments with new Data