Slackbot
02/11/2021, 10:24 PMBo
02/11/2021, 11:30 PMbackend for OnnxModelArtifact to onnxruntime-gpu
@bentoml.artifacts([OnnxModelArtifact('model_name', backend='onnxruntime-gpu')])
class MyService(bentoml.BentoService):
...
You can make sure it runs on gpu by:
>>> import onnxruntime as ort
>>> ort.get_device()
'GPU'
If you are using docker, there is a community workaround for the gpu support. We are working on the bring the solution to BentoML.
If you are using BentoML as api server, you will able to take advantage of micro batching feature from BentoML. You can read more about the architecture and design at https://docs.bentoml.org/en/latest/guides/micro_batching.html#the-overall-architecture-of-bentoml-s-micro-batching-server
There are two parameters you can update for the best performance for your prediction. mb_max_batch_size and mb_max_latency . You can read about these in the link I provided previously
class MyService(BentoService):
@api(input=DataframeInput(), mb_max_latency=3000, mb_max_batch_size=10000, batch=True)
def predict(self, df):
....Raphael
02/21/2021, 9:41 PM