Slackbot
07/04/2023, 8:11 AMlarme (shenyang)
07/04/2023, 4:28 PMworkers_per_resource=4 to have 4 instances of model B runner on a single GPU (and if you assign 2 gpus then you will have 8 instances of model B runner).
ref: https://docs.bentoml.org/en/latest/guides/scheduling.htmlPavel Schudel
07/05/2023, 8:45 AMlarme (shenyang)
07/05/2023, 8:46 AMPavel Schudel
07/05/2023, 9:18 AMworkers_per_resource: 4 for model B and workers_per_resource:1 for model A, I need to run the server with --api-workers 5 ?larme (shenyang)
07/05/2023, 9:25 AMrunner.<method_name>.run/async_run call, which is executed in runner server. For you use case, you can put cpu-only pre-processing/post-processing codes in api-server and scale up to many instances because cpus are cheap. On the other hand, the heavy computation happens inside runner server which will utilize the GPU.
If you are not sure, just don’t set —api-workers and we will spin up n api server instance where n = cpu number.larme (shenyang)
07/05/2023, 9:26 AMlarme (shenyang)
07/05/2023, 9:27 AMPavel Schudel
07/05/2023, 9:34 AMasyncio.grather on the api server ?larme (shenyang)
07/05/2023, 3:55 PMPavel Schudel
07/05/2023, 8:07 PMPavel Schudel
07/05/2023, 8:08 PM