Slackbot
08/17/2023, 4:39 PMVisaal Ambalam
08/17/2023, 4:39 PMJian Shen Yap
08/17/2023, 5:57 PMnvidia-smi to check on the gpu memory usagE?Visaal Ambalam
08/17/2023, 6:03 PMJian Shen Yap
08/17/2023, 6:05 PMbentoml serve --debug and see if you could find some logs there?Visaal Ambalam
08/17/2023, 6:36 PMbentoml._internal.marshal.dispatcher - Dynamic batching cork released, batch size: 3Visaal Ambalam
08/17/2023, 6:36 PMJian Shen Yap
08/17/2023, 7:07 PMAaron Pham
08/17/2023, 7:07 PMAaron Pham
08/17/2023, 7:07 PMVisaal Ambalam
08/17/2023, 7:18 PMJian Shen Yap
08/17/2023, 7:22 PMmax_latency for the batching configuration, what is your current settings for batching?Jian Shen Yap
08/17/2023, 7:24 PMrunner , because from the logs it is the runner that is returning 503 which is propagated to the api serverVisaal Ambalam
08/17/2023, 7:25 PMrunner.traffic.timeout is set to 300s and the RUNNERS_BATCHING_MAX_LATENCY_MS is 500msVisaal Ambalam
08/17/2023, 7:26 PMJian Shen Yap
08/17/2023, 7:36 PMVisaal Ambalam
08/17/2023, 8:26 PMrunners.batching.max_latency (link to entry in default config)
I guess I was confused because this is namespaced under the batching key so I thought it has to be related to batching. If batching was disabled, you’d still use this runners.batching.max_latency as the upper limit of the runner?Jian Shen Yap
08/17/2023, 10:04 PMVisaal Ambalam
08/18/2023, 7:41 PMVisaal Ambalam
09/19/2023, 4:07 PMrunners.traffic.timeout different from runners.batching.max_latency if the latter is “max_latency is the upper limit of latency on the runner”Visaal Ambalam
09/25/2023, 4:49 PM