This message was deleted.
# announcements
s
This message was deleted.
c
Hi Yilun, it actually means the GPU 1 and 2 will be assigned to
model_b
runner, the Runner’s implementation will decide how many workers will be started - it is customizable in the Runner’s implementation
We are still working on documentation for 1.0, however you can find an example of this from the built-in Runner implementation for XGBoost model here:
Copy code
def num_concurrency_per_replica(self) -> int:
        if self.resource_quota.on_gpu:
            return 1
        nthreads = self._booster_params.get("nthread", -1)  # type: int
        if nthreads == -1:
            return int(round(self.resource_quota.cpu))
        return nthreads


    @property
    def num_replica(self) -> int:
        if self.resource_quota.on_gpu:
            return len(self.resource_quota.gpus)
        return 1
https://github.com/bentoml/BentoML/blob/main/bentoml/_internal/frameworks/xgboost.py#L209-L221
y
Thanks Chadoyu! (Apology for late reply, I didn’t have notification for this slack channel). I guess my main question would be:
will I still need to create multiple bentos (of the same model) in order to host on multiple GPUs like I was doing in pre-1.0 era?
Since this will largely reduce the complexity of the whole pipeline for me.
c
will I still need to create multiple bentos (of the same model) in order to host on multiple GPUs like  I was doing in pre-1.0 era?
@Yilun Zhang the answer is no, you dont’ need to do that with BentoML 1.0
both for single machine with multiple GPU and GPU cluster, BentoML and yatai on kubernetes will handle that for the user
y
Amazing! Thanks!
Need to ramp my head around bentoml 1.0 then.