Slackbot
12/29/2021, 6:59 PMChaoyu
01/06/2022, 3:21 PMmodel_b runner, the Runner’s implementation will decide how many workers will be started - it is customizable in the Runner’s implementationChaoyu
01/06/2022, 3:22 PMChaoyu
01/06/2022, 3:22 PMdef num_concurrency_per_replica(self) -> int:
if self.resource_quota.on_gpu:
return 1
nthreads = self._booster_params.get("nthread", -1) # type: int
if nthreads == -1:
return int(round(self.resource_quota.cpu))
return nthreads
@property
def num_replica(self) -> int:
if self.resource_quota.on_gpu:
return len(self.resource_quota.gpus)
return 1
https://github.com/bentoml/BentoML/blob/main/bentoml/_internal/frameworks/xgboost.py#L209-L221Yilun Zhang
01/12/2022, 10:07 PMwill I still need to create multiple bentos (of the same model) in order to host on multiple GPUs like I was doing in pre-1.0 era?Since this will largely reduce the complexity of the whole pipeline for me.
Chaoyu
01/20/2022, 10:19 PMwill I still need to create multiple bentos (of the same model) in order to host on multiple GPUs like I was doing in pre-1.0 era?@Yilun Zhang the answer is no, you dont’ need to do that with BentoML 1.0
Chaoyu
01/20/2022, 10:20 PMYilun Zhang
01/20/2022, 10:29 PMYilun Zhang
01/20/2022, 10:29 PM