From what I understand, pinot does not have abilit...
# troubleshooting
j
From what I understand, pinot does not have ability to recover from this condition itself.
m
Typically, pinot will take checkpoints several times a day. If your Kafka retention is < 1 day, then you can run into this issue. But most practical deployments have a much longer retention in kafka.
l
i think the retention for this topic in particular is 1 day can you run me into the scenario where this may happen?
j
Makes sense. We changed our retention to 7 days after.
l
also nothing is getting consumed at all
in short, pinot is trying to consume an older offset that may have been deleted by kafka already?
m
That should not happen. Pinot typically will save offset every few hours.
l
do you have any other advice to troubleshoot consuming problems?
when you say save offset you mean the offset it last saved yes?
so that if we restart or anything it will pick it up from there
m
yes
Saving offset -> committing segment
l
it saves the offset when it commits the segment?
r
Hi Mayank: for "Pinot typically will save offset every few hours." --> isn't this configurable by tableConfig.streamConfig ? e.g.
realtime.segment.flush.threshold.time
would it be possible that the flush didn't get trigger because the time threshold is higher then kafka retention period (and no other flush criteria met either)
m
@Rong R yes technically that is possibles what I am pointing out is that in practice, you want to flush segments at least once per day, and you want to have Kafka retention at least a few days. If you configure Pinot to not flush segments for several days, that is not a good design choice.
Perhaps you can help document this in ingestion Faq ?
👍 1
r
yes. will do.
l
😮 Rong R that’s interesting
i have the options as this
Copy code
"realtime.segment.flush.threshold.rows": "0",
"realtime.segment.flush.threshold.time": "24h",
"realtime.segment.flush.segment.size": "250M"
but this has been working so far for our setup is the first time i have been seeing this
Fetch position FetchPosition
 
is out of range for partition
 
resetting offset
i guess i don’t understand the relation between it
r
reduce
"realtime.segment.flush.threshold.time": "6h",
or increase your kafka retention will resolve this issue going forward. however you wont be able to recover what’s lost since data is gone already
m
@Luis Fernandez what’s your retention in Kafka?
Also do you know how often are segments flushed in your case?
l
they flush every day we have 8 kafka partitions so i see 8 being created every day
omg i just came to know our default retention for our dev cluster is 1h 😄 sooooo yeaa maybe something funky was happening
😬 1
so hear this out and let me know
if i had my server running all well and then i decided say, to restart it, it would try to go to the last commited offset it has an idea of from the last segment buut because our retention is only 1h then it would try to keep on fetch that and get this error that we have and wouldn’t be possible to read that and would halt consumption
m
Yes, if your kafka retention is so small that messages not in pinot’s checkpoint can be deleted, you will run into this issue.
l
so how do i recover from this? so i want it to ingest data again
could i change this somehow and get it to ingest again
"stream.kafka.consumer.prop.auto.offset.reset"
bumping this thread
m
But from what I understand, your Kafka retention is one hour? And so there is no data in Kafka to recover right?
l
right no data to recover but I want to resume consumption cause consumption is halted
r
yes you can change it to largest and get it to consume again. but you will lose data
m
You can manually change the offset in zk
@Neha Pawar any better suggestions? If not, we should add a way (api?) to resume consumption.
j
May I suggest that this become a configuration option for consumption. Like, is offset is missing in kafka then resume from latest. Sort of on par with Kafka consumer.
➕ 1
👍 1
l
so my configuration already has largest what does that mean just upload the same config again and something will happen?
n
What's the exact complete log line? Somehow I'm unable to find this in the code. Also this means what rong and Mayank said. But you don't need to do anything to recover from it. The message is telling that it has reset the offset
l
let me see what the state is right now, well thing is that consumption is just halted and when i look into the pinot-server logs i just see that being spammed
Copy code
[Consumer clientId=consumer-null-23, groupId=null] Fetch position FetchPosition{offset=5063422319, offsetEpoch=Optional.empty, currentLeader=LeaderAndEpoch{leader=Optional[xx..xx..x (id: 4, epoch=5}} is out of range for partition beacon-ads-listing-stats-7, resetting offset
[Consumer clientId=consumer-null-23, groupId=null] Resetting offset for partition beacon-ads-listing-stats-7 to position FetchPosition{offset=553642866, offsetEpoch=Optional.empty, currentLeader=LeaderAndEpoch{leader=Optional[xx..x..xx (id: 4 )], epoch=5}}.
and it just keeps doing that over and ver again
that’s weird that null in the clientId tho
Copy code
Consumed 0 events from (rate:0.0/s), currentOffset=5063422318, numRowsConsumedSoFar=0, numRowsIndexedSoFar=0
n
oh this is from kafka. so yes, looks like there’s no way to recover. we should fix this. for now, what Mayank said is the only way. You’ll need to manually change startOffset in the consuming segment zk metadata and restart
l
can i find that in the zk browser? have to manually update but just wanna see where do i see it in the zookeeper browser in the UI
n
under PROPERTYSTORE SEGMENTS
l
sorry may have to get many questions here, i do see all the segments here, how do i find the consuming one? i guess i found it just by look by clicking around
Copy code
"segment.realtime.status": "IN_PROGRESS"
that means is the current yes
n
Yes
l
so there’s a huge list there how can i find it quickly, i guess no easy way?
n
if this is a dev table, would it be easier to drop the table and restart consumption, and set the time threshold to something low like 30m or increase kafka retention?
l
yea i guess i just wanted to have a scenario and steps that hey if you ever run into this and don’t want to lose all data this may be what you want to do
we already increased our retention in the topics to 2 days in dev and 5 days in prod, but we don’t have a pinot consumer in prod yet
n
we will think about it some more. I’m wondering why kafka side, ifit is saying the offset was reset, why the consumption didn’t continue from the new offset it set. so will check if it’s returning any exception that we can catch at Pinot side and reset the offset. Do you mind filing an issue on github?
l
cool i will file the issue
👍 1
s
@Luis Fernandez if your afka retention is 24h, then you should not set
realtime.segment.flush.threshold.time
to be
24h
. It should be less than kafka retention time, preferably 1/3 or even 1/4. Secondly, if a consumer tries to consume an offset that does not exist in Kafka, the Kafka driver is expected to throw an exception. If that is not happening, then it is a bug in Kafka. whichKafka plugin are you using? Thirdly, if the driver indeed throws an exception, Pinot background jobs automatically re-create new segments that start to consume from earliest available offset. If there is data loss, a metric is emitted (from the controller) as well as a log warning so. cc: @Sajjad Moradi
l
2.0
and yea topics retention time has been increased
s
So, you are saying the 2.0 version (in open source pinot trunk) does not throw exception if the offset is in the past? Please add that observation to the issue, and we know that the fix is in that area and not in core pinot.
l
cool, will get back to you
updated