Ah, the problem in my case was caused by bentoml s...
# ask-for-help
m
Ah, the problem in my case was caused by bentoml spawning an unbounded number of api workers after almost every request (even though i didn't do concurrent requests at all), each of which took some GPU memory Setting api_workers in config didn't help.