Hi Pinot experts, we had a sudden ingestion lag (&...
# troubleshooting
p
Hi Pinot experts, we had a sudden ingestion lag (> 45m) in one of our tables today without any increase in incoming throughput. This happened suddenly to just that one table (the rest of the tables had very little lag in seconds). We had to scale up our Kafka partitions for it to catch up. What might be the possible reason behind this sudden lag? @Tanmay Movva
m
Seems like the issue was on the Kafka side then?
t
Metrics related to kafka were normal. And all the tables are consuming from the same cluster. We saw this issue only on one of the tables.
Any pointers on how we can confirm/dismiss if this was a kafka related issue?
We also didn’t see anything abnormal in pinot’s metrics, but the ingestion slowed down. We increased partitions and scaled up pinot relatime servers to mitigate lag and it has helped.
m
What did you change on the Pinot side
t
Nothing changed on pinot and kafka, since atleast 2 weeks.
m
Oh you did increase pinot serves
t
Yes, we did after we started observing lag in pinot table.
m
Did the ingestion on pinot side slow down or was there periods of stopped consumption for some partitions
Also did read qps increase for those servers
t
No. QPS was very low.
CPU utilisation for those servers was also normal. (60-70% like usual).
This is the ingestion rate graph for the table which had lag. The spikes at the right end are when we restarted the servers after increasing kafka topic partitions and scaling realtime servers from 2 to 4.
m
Were a lot of segments being generated? Also, was there GC?
a
Usually with Kafka, there are problems with rebalancing. Have you checked out how often you have extra brokers being spun off? Spinning off extra brokers means that each broker also has to replicate the topic partitions and replicas which can be time-consuming. This could be causing a bottleneck as well.