aditi
07/20/2023, 2:14 PMcurl "<http://ip:9000/tables/table_name/size?detailed=true>" 2> /dev/null | jq -c
{"tableName":"table_name","reportedSizeInBytes":0,"estimatedSizeInBytes":0,"offlineSegments":null,"realtimeSegments":{"reportedSizeInBytes":0,"estimatedSizeInBytes":0,"missingSegments":0,"segments":{}}}
kafka topic has around GB's of data but no segments in pinot...
When i run producer to add records in pinot , i do see increase in number of records via pinot query console (ran count(*))...
what could be the issue here, why segments are not getting created?aditi
07/20/2023, 2:15 PMaditi
07/20/2023, 2:26 PM{
"segment.realtime.numReplicas": "2",
"segment.creation.time": "1689846621631",
"segment.flush.threshold.size": "100000",
"segment.realtime.startOffset": "0",
"segment.realtime.status": "IN_PROGRESS"
}Johan Adami
07/20/2023, 2:29 PMaditi
07/20/2023, 2:33 PMrealtime.segment.flush.threshold.size": "16M"
in table json...
Its been 5-6 hours and still in consuming state....Johan Adami
07/20/2023, 2:38 PMsize thresholdJohan Adami
07/20/2023, 2:39 PMJohan Adami
07/20/2023, 2:39 PMrealtime.segment.flush.threshold.rows=0, otherwise i think pinot waits for 100k rowsaditi
07/20/2023, 2:40 PMaditi
07/20/2023, 2:41 PMrealtime.segment.flush.threshold.rows to some value?Johan Adami
07/20/2023, 2:41 PMaditi
07/20/2023, 2:42 PMaditi
07/20/2023, 2:44 PM{
"REALTIME": {
"tableName": "table_name",
"tableType": "REALTIME",
"segmentsConfig": {
"timeType": "MILLISECONDS",
"schemaName": "table_name",
"replicasPerPartition": "2",
"timeColumnName": "timestamp",
"peerSegmentDownloadScheme": "http",
"minimizeDataMovement": false
},
"tenants": {
"broker": "DefaultTenant",
"server": "DefaultTenant"
},
"tableIndexConfig": {
"streamConfigs": {
"streamType": "kafka",
"stream.kafka.consumer.type": "simple",
"stream.kafka.topic.name": "kafka-topic",
"stream.kafka.decoder.class.name": "org.apache.pinot.plugin.stream.kafka.KafkaJSONMessageDecoder",
"stream.kafka.consumer.factory.class.name": "org.apache.pinot.plugin.stream.kafka20.KafkaConsumerFactory",
"stream.kafka.broker.list": "broker_list",
"realtime.segment.flush.threshold.rows": "0",
"realtime.segment.flush.threshold.segment.size": "16M",
"realtime.segment.serverUploadToDeepStore": "true",
"stream.kafka.consumer.prop.auto.offset.reset": "smallest"
},
"rangeIndexVersion": 2,
"autoGeneratedInvertedIndex": false,
"createInvertedIndexDuringSegmentGeneration": false,
"loadMode": "MMAP",
"enableDefaultStarTree": false,
"enableDynamicStarTreeCreation": false,
"aggregateMetrics": false,
"nullHandlingEnabled": false,
"optimizeDictionary": false,
"optimizeDictionaryForMetrics": false,
"noDictionarySizeRatioThreshold": 0
},
"metadata": {
"customConfigs": {}
},
"isDimTable": false
}
}Johan Adami
07/20/2023, 2:54 PMaditi
07/20/2023, 2:57 PMaditi
07/20/2023, 2:58 PMJohan Adami
07/20/2023, 2:59 PMJohan Adami
07/20/2023, 3:00 PMaditi
07/20/2023, 3:01 PMaditi
07/20/2023, 3:29 PMJohan Adami
07/20/2023, 3:37 PMJohan Adami
07/20/2023, 3:38 PMaditi
07/20/2023, 3:56 PMkafka-log-dirs.sh --describe --bootstrap-server servers --topic-list kafka-topic-1
It has around 2.5 GB in each partition, there are 32 partitions so 80 GB of data.
You are saying duplicate records does not have impact on pinot and then i am not sure why pinot is not showing the data in segments....aditi
07/20/2023, 4:33 PMbin/kafka-run-class.sh kafka.tools.GetOffsetShell --broker-list localhost:9092
--topic kafka-topic-1 | awk -F ":" '{sum += $3} END {print "Result: "sum}'
and
select count(*) from table_name;
gave same result.