hey guys, we are ingesting data from a kafka; it ...
# troubleshooting
p
hey guys, we are ingesting data from a kafka; it seeems pinot is stuck in a loop; it complains about not finding the right offset and does not consume the data that is sitting in kafka.
m
Seems like upstream Kafka got changed (as in offset completely changed)?
p
we had small retention so possibly that data got removed
$ /opt/bitnami/kafka/bin/kafka-run-class.sh kafka.tools.GetOffsetShell --broker-list :9092 --topic kf_metrics_topic --time -2 kf_metrics_topic02031 kf_metrics_topic11926 kf_metrics_topic21918
shouldn’t pinot move forward and start consuming from the latest offset?
m
Yeah, so pinot checkpoints the offset it has committed to. If that offset doesn't match on what is seen at kafka side, then there is no way for Pinot to recover from that scenario. Consuming from smalles/ latest only applies when there are no offset checkpoints.
You'll have to delete and recreate the table.
p
even if table streaming is set as lowlevel?
“stream.kafka.consumer.type”: “lowlevel”,
m
Yes
p
does doing disable/enable on the web console for the table do anything to help with or we have to delete/recreate the table?
m
The issue is that the checkpointed offset in Pinot doesn't corresponding to upstream as your upstream completely changed. No way for Pinot to guarantee data consistency after this point. So delete/recreate is the only option
p
so offset checkpoint management is always done; why would one pick lowlevel v/s highlevel consumer type?
till now we had always picked lowlevel assuming that there is no offset management
m
May I ask what really changed on Kafka side?
Offset management is independent of consumer type right?
p
nothing really; i think our retention is set to really small value; 1/2GB. no segment was sealed and i restarted pinot;
m
Pinot has to checkpoint offsets to prevent data loss.
Oh, so segments are committed yet?
p
no segments were comitted; i restarted and i was hoping it will read the kafka queue.. last time i did it there was a few hours gap
so I am guessing kafka retention kicked in
m
Setting retention based on size might cause issue right? What if retention kicks in before Pinot could consume the events (likely that's what happened here)
Is this a local setup, or are you planing to run int his configuration in production env?
p
trying to tune this setup for production.. k we will increase the data retention. Thanks for your help
m
In production, I'd recommend time based retention.
s
@Pankaj Thakkar did you try running the realtime validation manually? When kafka consumers cannot find an offset, the segments go into OFFLINE state. Periodic task realtime validator comes around to restart consumption from the current offsets available. Not sure how often you run this task in your environment (I think default is evry hour), but you can also trigger it manually to correct things
p
no didn’t know about that. The way I had run it was to run some workload to populate kafka; then stopped the workload; pinot created segments but the data was not enough so no segments were created; then i restarted pinot after a long time;
How do you run it?
s
try the swagger apis under /periodictask