This message was deleted.
# ask-for-help
s
This message was deleted.
v
like how would we know if the gpu is running out of memory?
j
are you able to have any inference, or 503 only happens when you are sending alot of request? for gpu, you can perhaps use
nvidia-smi
to check on the gpu memory usagE?
v
is there no way to get the error that bentoml is seeing before returning Service Busy?
j
could you try to run
bentoml serve --debug
and see if you could find some logs there?
v
I just see stuff like this
Copy code
bentoml._internal.marshal.dispatcher - Dynamic batching cork released, batch size: 3
But nothing specific around the 503
j
503 typically happens when server is overloaded. Could you monitor the memory usage of the runner? @Aaron Pham do we know what are the other possible cases that might trigger 503 from the runner server?
a
This might happen when the dispatcher is overloaded, and it will reject incoming requests
you can try increaes timeout
v
Will increase the API_SERVER_TRAFFIC_TIMEOUT to 300s from 60s, but you can see from the logs that the 503 errors are barely 1s. The cpu usage is like really small (less than 1 core) and memory usage is constant at around 10GB but we have plenty more. Like the pods aren’t running out of memory and restarting or anything.
j
i think what aaron meant was to increase the
max_latency
for the batching configuration, what is your current settings for batching?
as well as the timeout of the
runner
, because from the logs it is the runner that is returning 503 which is propagated to the api server
v
The
runner.traffic.timeout
is set to 300s and the
RUNNERS_BATCHING_MAX_LATENCY_MS
is 500ms
My understanding was the max_latency for batching is how long it would wait before sending in a batch for processing though, not the max latency of the api, runner, or entire request.
j
max_latency is the upper limit of latency on the runner. if you set it to 500ms and the runner' dispatcher predicts that a new incoming request will violate the limit, it will reject the request directly. So you should increase this to a limit that you are able to accept, ideally around the 80-90% of your application's SLA
v
Ohh, okay gotcha. Just to confirm it’s this right?
runners.batching.max_latency
(link to entry in default config) I guess I was confused because this is namespaced under the
batching
key so I thought it has to be related to batching. If batching was disabled, you’d still use this
runners.batching.max_latency
as the upper limit of the runner?
j
yes! you are right for both. if batching is disabled, the dispatcher will just forward the request as it is
v
Thank you everyone! Seems like increasing the max_latency timeout fixed the 503 error issue. I’ll keep an eye on the logs if we get some more.
How is
runners.traffic.timeout
different from
runners.batching.max_latency
if the latter is “max_latency is the upper limit of latency on the runner”
Hi @Jian Shen Yap can you please clarify when you get the chance?