Hi All . We are couple of new users from a NA base...
# troubleshooting
v
Hi All . We are couple of new users from a NA based Investment Bank . We are pocing heavily on Pinot and getting some road blocks and need some guidance
m
Hi @Vibhor Jaiswal welcome to the community, happy to assist.
v
Hi Mayank . Thanks Man .
Here is the use case . We are using Pinot 0.9.0 . Realtime
When we do the 2 million load and upserts with 5 batches of 2 million per batch .
Everything goes well . We are able to query per second with consistency .
However when we do a initial load of 10 million plus upserts of 10 batch of 10 million each , the things starts falling off
the segment starts rolling off in the middle and broker can not connect to the server for some time and server goes off the grid .
What we are doing wrong ?
Below is the table config -{  "REALTIME": {    "tableName": "sometable",    "tableType": "REALTIME",    "segmentsConfig": {      "schemaName": "someschema",      "timeColumnName": "AuditDateTimeUTC",      "allowNullTimeValue": false,      "replication": "1",      "replicasPerPartition": "1"    },    "tenants": {      "broker": "DefaultTenant",      "server": "DefaultTenant",      "tagOverrideConfig": {}    },    "tableIndexConfig": {      "invertedIndexColumns": [],      "rangeIndexColumns": [],      "rangeIndexVersion": 1,      "autoGeneratedInvertedIndex": false,      "createInvertedIndexDuringSegmentGeneration": false,      "sortedColumn": [],      "bloomFilterColumns": [],      "loadMode": "MMAP",      "noDictionaryColumns": [],      "onHeapDictionaryColumns": [],      "varLengthDictionaryColumns": [],      "enableDefaultStarTree": false,      "enableDynamicStarTreeCreation": false,      "aggregateMetrics": false,      "nullHandlingEnabled": false,      "streamConfigs": {        "streamType": "kafka",        "stream.kafka.topic.name": "sometopic",        "stream.kafka.broker.list": "host:9092",        "stream.kafka.consumer.type": "lowlevel",        "stream.kafka.hlc.bootstrap.server": "host:9092",        "stream.kafka.consumer.prop.auto.offset.reset": "largest",        "stream.kafka.consumer.factory.class.name": "org.apache.pinot.plugin.stream.kafka20.KafkaConsumerFactory",        "stream.kafka.decoder.class.name": "org.apache.pinot.plugin.stream.kafka.KafkaJSONMessageDecoder",        "realtime.segment.flush.threshold.rows": "0",        "realtime.segment.flush.threshold.size": "0",        "realtime.segment.flush.autotune.initialRows": "10000000",        "realtime.segment.flush.threshold.time": "24h",        "realtime.segment.flush.threshold.segment.size": "10G"      }    },    "metadata": {},    "quota": {},    "routing": {      "instanceSelectorType": "strictReplicaGroup"    },    "query": {},    "upsertConfig": {      "mode": "FULL",      "comparisonColumn": "AuditDateTimeUTC",      "hashFunction": "NONE"    },    "ingestionConfig": {},    "isDimTable": false  } }
m
Have a couple of questions, let me dm
Right off the bat, 10G is too big a segment size