~Hello, in my production environment we are runnin...
# ask-for-help
m
~Hello, in my production environment we are running ML models on the edge device. For A/B testing reasons we would like to be able to run two instances of bentoml server with TF models, however we only have 1 GPU. So far this approach worked for our own server implementation as we've used
soft_device_placement = True
and
set_memory_growth = True
. I've tried doing that in bentoml server, but it does not work. One server allocates all of the memory, and the second one faile with CUDA_ERROR_OUT_OF_MEMORY. Is there an option in bentoml where this can be configured? If not, is there a good place to patch it in? Do you have any other recommendations? Best regards Michał~