Slackbot
11/17/2022, 3:35 PMJiang
11/18/2022, 3:21 AMthe container still seems to receive multiple requests simultaneously.It's an uncommon demand. Why do you want it to only receive single request? Then let me answer the first question. The API server is an async server, each worker can handle many requests at the same time.
Jiang
11/18/2022, 3:26 AMasyncio.gather the run multiple runner calls at same time rather than one-by-one if possible. What actual error did you encounter?Jiang
11/18/2022, 3:35 AMrunners:
pytorch_mnist:
resources:
<http://nvidia.com/gpu|nvidia.com/gpu>: [0, 2]
If not specified, bentoml is having a optimistic policy. It will deploy every runner to each GPU (if the runner support), and also making use of all the CPU cores. It guaranteed all the resources is fully used as possible.Sangeon Yong
11/18/2022, 7:28 AM