This message was deleted.
# ask-for-help
s
This message was deleted.
🏁 1
l
Hi Pavel, BentoML currently does not isolate gpu resource per runner. if you can guarantee that model B will use max 500MB vram, then you can set
workers_per_resource=4
to have 4 instances of model B runner on a single GPU (and if you assign 2 gpus then you will have 8 instances of model B runner). ref: https://docs.bentoml.org/en/latest/guides/scheduling.html
p
great thanks I think it solves my problem
l
welcome
p
I set
workers_per_resource: 4
for model B and
workers_per_resource:1
for model A, I need to run the server with
--api-workers 5
?
l
api-workers deciedes how many instances of api server you will spin up. When you writing the service endpoint function, all codes run inside api server except the
runner.<method_name>.run/async_run
call, which is executed in runner server. For you use case, you can put cpu-only pre-processing/post-processing codes in api-server and scale up to many instances because cpus are cheap. On the other hand, the heavy computation happens inside runner server which will utilize the GPU. If you are not sure, just don’t set
—api-workers
and we will spin up n api server instance where n = cpu number.
this graph may give you a visual idea of our server architecture
(just ignore the dispatcher part
p
ok, and final question. if for instance I have a string paragraph as an input. and I split it into sentences, I want to process every sentence in parallel. I should use a runner for the processing with some workers or I can use
asyncio.grather
on the api server ?
l
Hi Pavel, does you model accept batch input?
p
no
the model is not built for batch input