This message was deleted.
# ask-for-help
s
This message was deleted.
c
It depends on the runner implementation and model architecture. Most of the time, it will be 1. Could you share a bit more about your use cases?
v
We’re trying to understand when bentoml spawns more runner workers because our pods are running OOMing and were wondering if that’s a potential cause. Here’s the profile, somewhere in the
bentoml/io_descriptors/json.py
it’s allocating 54GB of memory. So maybe it’s loading all the requests in the backlog and that’s what is causing the OOM? So maybe reduce
api_server.backlog
or something?