Hi team, I’ve a question about 434 rows data lost....
# troubleshooting
a
Hi team, I’ve a question about 434 rows data lost. In my case, Flink is used to do data transformation, and then send data to Kafka, and transaction is enabled. Therefore, ‘read_commit’ is enabled in our table config. During data validation, we found some records were not ingested by pinot. There’s no error in server log. In order to trace the root cause, we created a new table with the same table config and found those missed records were in this new table. So, my question is, what’s the possible reason for the data missing in the first time?
Copy code
2022/07/24 22:38:22.836 INFO [LLRealtimeSegmentDataManager_table_name__1__390__20220724T2139Z] [table_name__1__390__20220724T2139Z] Consumed 0 events from (rate:0.0/s), currentOffset=508692985, numRowsConsumedSoFar=567388, numRowsIndexedSoFar=550671
2022/07/24 22:39:24.041 INFO [LLRealtimeSegmentDataManager_table_name__1__390__20220724T2139Z] [table_name__1__390__20220724T2139Z] Consumed 0 events from (rate:0.0/s), currentOffset=508692985, numRowsConsumedSoFar=567388, numRowsIndexedSoFar=550671
2022/07/24 22:39:57.515 INFO [LLRealtimeSegmentDataManager_table_name__1__390__20220724T2139Z] [table_name__1__390__20220724T2139Z] Stopping consumption due to time limit start=1658698797360 now=1658702397515 numRowsConsumed=567388 numRowsIndexed=550671
2022/07/24 22:39:57.515 INFO [LLRealtimeSegmentDataManager_table_name__1__390__20220724T2139Z] [table_name__1__390__20220724T2139Z] Stopping consumption due to time limit start=1658698797360 now=1658702397515 numRowsConsumed=567388 numRowsIndexed=550671
2022/07/24 22:39:57.600 INFO [HttpClient] [table_name__20220724T2139Z] Sending request: <http://pinot-controller-1.pinot-controller-headless.pinot.svc.cluster.local:9000/segmentConsumed?reason=timeLimit&streamPartitionMsgOffset=508693423&instance=Server_pinot-server-1.pinot-server-headless.pinot.svc.cluster.local_8098&offset=-1&name=table_name__1__390__20220724T2139Z&rowCount=567388&memoryUsedBytes=159996334> to controller: pinot-controller-1.pinot-controller-headless.pinot.svc.cluster.local, version: Unknown
2022/07/24 22:39:57.600 INFO [ServerSegmentCompletionProtocolHandler] [table_name__1__390__20220724T2139Z] Controller response {"offset":508693423,"streamPartitionMsgOffset":"508693423","buildTimeSec":126,"isSplitCommitType":true,"controllerVipUrl":"<http://pinot-controller-1.pinot-controller-headless.pinot.svc.cluster.local:9000>","status":"COMMIT"} for <http://pinot-controller-1.pinot-controller-headless.pinot.svc.cluster.local:9000/segmentConsumed?reason=timeLimit&streamPartitionMsgOffset=508693423&instance=Server_pinot-server-1.pinot-server-headless.pinot.svc.cluster.local_8098&offset=-1&name=table_name__1__390__20220724T2139Z&rowCount=567388&memoryUsedBytes=159996334>
2022/07/24 22:39:57.600 INFO [Metrics] [table_name__1__390__20220724T2139Z] Metrics scheduler closed
Above is the log where data discrepancy occurred. Before pinot server stopped consuming, the offset is 508692985. When pinot server stopped consuming, controller response offset was 508693423, while numRowsConsumed was unchanged 567388. Based on our data comparision, about 434 records were lost.
Data missing chance is very low. About 1 out of 138 segments has this data missing issue.
s
possible you may be hitting this issue? - https://github.com/apache/pinot/issues/9091
a
Got it. Thanks. @Stuart Coleman
BTW, I’m using Pinot 0.11.