This message was deleted.
# announcements
s
This message was deleted.
b
hi @Raphael Great to hear. Love to chat more to understand your use-cases better. Can you tell me more about your production setup? Are you running the service in docker? For running onnx model on GPU, make sure you set
backend
for
OnnxModelArtifact
to
onnxruntime-gpu
Copy code
@bentoml.artifacts([OnnxModelArtifact('model_name', backend='onnxruntime-gpu')])
class MyService(bentoml.BentoService):
    ...
You can make sure it runs on gpu by:
Copy code
>>> import onnxruntime as ort
>>> ort.get_device()
'GPU'
If you are using docker, there is a community workaround for the gpu support. We are working on the bring the solution to BentoML. If you are using BentoML as api server, you will able to take advantage of micro batching feature from BentoML. You can read more about the architecture and design at https://docs.bentoml.org/en/latest/guides/micro_batching.html#the-overall-architecture-of-bentoml-s-micro-batching-server There are two parameters you can update for the best performance for your prediction.
mb_max_batch_size
and
mb_max_latency
. You can read about these in the link I provided previously
Copy code
class MyService(BentoService):
    @api(input=DataframeInput(), mb_max_latency=3000, mb_max_batch_size=10000, batch=True)
    def predict(self, df):
        ....
r
Thank you so much, I am going to run some benchmarks for now and evaluate
👍 1