This message was deleted.
# announcements
s
This message was deleted.
c
BentoML has its own async layer built with aiohttp which is optimized for ML model serving
The flask part is the backend model server, that we could replace with something else later on and does not affect the overall performance that much
There’s a benchmark project under the bentoml GitHub org page
w
yeah I've checked the bentoml/benchmark before, and I'm just thinking maybe a contrast matrix or something like that might make it more attractive. bentoml has made model to service much simpler, I like the idea of packing model with the service code and generating docker image, but some users could hesitate on performance, an intuitive show-out might do some help.
👍🏼 1
by the way, last few days I'm working on how to make my transformers model run on bentoml with gpu, I solve that problem with a customize bento artifact(found some sample code on slack), but I'm wondering if there is an easier way like changing some setting in bento service?
Copy code
# now i do like this
class MyTMA(TransformersModelArtifact):
    def pack(self, model):
        super().pack(model)
        self._model["model"].cuda()



@bentoml.env(infer_pip_packages=True)
@bentoml.artifacts([MyTMA("albert")])
class TransformerService(bentoml.BentoService):
    @bentoml.api(input=JsonInput(), batch=False)
    def predict(self, parsed_json):
        pass
j
optimized for ML model serving
Just curious, beside batching algorithm, what are the optimization that you guys have made?
j
@Jian Shen Yap Micro-batching is the major approach. We also made other optimizations like Dataframe merging (calling
pandas.read_json
in a loop is too slow), backpressure control (a common problem for async apps).
Yeah, appling
fastAPI
or
sanic
to the model server (bentoml also has an asyncio batching server in the front) wouldn’t make it faster. Normal web servers are IO-intensive, but inferences of ML models are CPU-intensive.
👍 3
The exception here is the batching server. It holds thousands of connections from clients, merges inputs to batch, and dispatches to model servers. They are IO-bound. That’s why it's implemented with asyncio. In addition to https://github.com/bentoml/benchmark, there is a user in the channel who just made some benchmarks comparing with the tensorflow-serving. Bentoml and tf-serving have similar latency and throughput under high concurrent requests.
j
@Jiang if you are referring to me, yes, I did a benchmark comparing with tf-serving. on low load, few request at a time, tf serving will be able to achieve very low latency. I supposed this is due to tf serving not needing to have python overhead. but at high load thousands RPS, bentoML are able to match the throughput and beat the latency of tf serving. for some reason the response time for bentoML is very stable, while tfserving has higher variation in response time (difference between p5, p50 , p95 response time is huge)
🤣 1
j
hi @Jiang, I compare bert with bentoml and tf-serving. And I get a huge gap of the performance. Betoml will get lots of 429 response and slow responses while tf-serving can success fast and completely. Do you have any ideas how to improve the service?
j
@Jackson Hsieh I also tried BERT with bentoml. There are two possible reasons for your problem: • Can you find many lines like
triggered tf.function retracing
in the logs? The restracing is very expensive. Here is the solution: https://github.com/tensorflow/tensorflow/issues/38561#issuecomment-613894395 .
@tf.function(experimental_relax_shapes=True)
• When running with micro-batching enabled, bentoml requires about five requests to preheat the model and the batching layer. During this the micro-batching will not take effect. So at the beginning, some 429's are normal. The performance will gradually reach the best.
j
@Jiang I test my service and find most of the micro-batch will have size 1; only few times will have size larger one. Do I need to tune mb_max_latency or mb_max_batch_size ?
Also, when I want to increase our RPS but I will got lots of 429. Do I need to increase the mb_max_latency?