This message was deleted.
# ask-for-help
s
This message was deleted.
j
the container still seems to receive multiple requests simultaneously.
It's an uncommon demand. Why do you want it to only receive single request? Then let me answer the first question. The API server is an async server, each worker can handle many requests at the same time.
For Q2, that is supported. But we suggest to use
asyncio.gather
the run multiple runner calls at same time rather than one-by-one if possible. What actual error did you encounter?
For Q3, I head to the main doc and find that we still not having a scheduling doc. At first, the scheduling of GPU and CPU can be controlled during deployment stage. https://docs.bentoml.org/en/latest/guides/configuration.html#configuration See the configuration: In short, if you want one of your runner takes GPU0 and GPU2, you just need to specify it in the config file:
Copy code
runners:
 pytorch_mnist:
   resources:
     <http://nvidia.com/gpu|nvidia.com/gpu>: [0, 2]
If not specified, bentoml is having a optimistic policy. It will deploy every runner to each GPU (if the runner support), and also making use of all the CPU cores. It guaranteed all the resources is fully used as possible.
s
Thank you for your answer! At first, I wanted to try a single worker because of the error in Q2. Actually, it is not for actual usage. In Q2, I just encountered the error that the tensor dimension mismatch is inside the model, and this does not happen when we do not send multiple API calls simultaneously. Currently, I found that using async_run() with await instead of run() does not cause the error, but I don’t know that it is a recommended solution.