Ah, the problem in my case was caused by bentoml spawning an unbounded number of api workers after almost every request (even though i didn't do concurrent requests at all), each of which took some GPU memory
Setting api_workers in config didn't help.