has anyone gotten this issue before? I’m getting t...
# troubleshooting
l
has anyone gotten this issue before? I’m getting the following exception in one of the brokers:
Copy code
2021-12-13 11:31:40	
java.lang.OutOfMemoryError: Direct buffer memory
2021-12-13 11:31:40	
Caught exception while handling response from server: pinot-server-1_R
we currently have 2 brokers, currently doing a lot of garbage collection i’m unaware as to why. latency from broker to server has been severed by a lot but I’m not sure what happened as to we haven’t been touching the pinot cluster lately, we did stop one of our apps from streaming but that doesn’t line up with the spikes on response times.
r
broker OOM?
l
time spent on GC on one of the brokers, the other broker seems to not have the issue but I’m unsure why
that’s for G1 Old Generation
r
this could relate to something I've seen before, a netty contributor warned me about OOM
this isn't related to heap memory, so the GC metrics are a red herring
how much direct memory have you given your brokers? (
-XX:MaxDirectMemorySize
)
l
Copy code
value: "-Xms2G -Xmx2G -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -Xlog:gc*:file=/opt/pinot/gc-pinot-broker.log -Dlog4j2.configurationFile=/opt/pinot/conf/log4j2.xml -Dplugins.dir=/opt/pinot/plugins -javaagent:/opt/pinot/etc/jmx_prometheus_javaagent/jmx_prometheus_javaagent-0.12.0.jar=8008:/opt/pinot/etc/jmx_prometheus_javaagent/configs/pinot.yml"
image.png
seems like i haven’t set that up
but it’s not using all the heap i’m confused
that’s jvm heap ^
r
ok, so unless you set it, it should default to 2G because that's how you set Xmx
Do you have off heap memory metrics?
l
image.png
so that’s the pod memory
def more than 2g 😄
but the jvm metics show something else why is off heap
so not the best but i restarted the pod and then we are good
but i’m concerned that i don’t know how and what happened 😄
r
no it's not the best, let me figure out if there is a broker metric which shed more light in to what happened
m
I see a couple of possibilities - a) Large responses at high throughput returned from server to broker, b) We have seen OOM in netty when netty version in server/broker mis-match. This shouldn't be the case if you are using an official release.
For a), there's probably a metric on server response size (I know we log it in the broker). For b) can you
unzip -l
the jars to check the netty version on broker/server?
l
just more context don’t know if it helps we haven’t really touch pinot for a long time it has been ingesting and shadowing mysql requests and this happened to one of the brokers out of nowhere
sorry for b) which jars are we talking about?