Slackbot
12/04/2020, 9:11 AMChaoyu
12/04/2020, 9:17 AMChaoyu
12/04/2020, 9:18 AMChaoyu
12/04/2020, 9:18 AMWang
12/04/2020, 9:40 AMWang
12/04/2020, 9:47 AMWang
12/04/2020, 10:06 AM# now i do like this
class MyTMA(TransformersModelArtifact):
def pack(self, model):
super().pack(model)
self._model["model"].cuda()
@bentoml.env(infer_pip_packages=True)
@bentoml.artifacts([MyTMA("albert")])
class TransformerService(bentoml.BentoService):
@bentoml.api(input=JsonInput(), batch=False)
def predict(self, parsed_json):
passJian Shen Yap
12/04/2020, 3:40 PMoptimized for ML model serving
Just curious, beside batching algorithm, what are the optimization that you guys have made?Jiang
12/05/2020, 3:00 AMpandas.read_json in a loop is too slow), backpressure control (a common problem for async apps).Jiang
12/16/2020, 9:38 AMfastAPI or sanic to the model server (bentoml also has an asyncio batching server in the front) wouldn’t make it faster. Normal web servers are IO-intensive, but inferences of ML models are CPU-intensive.Jiang
12/16/2020, 9:45 AMJian Shen Yap
12/17/2020, 3:38 AMJackson Hsieh
01/14/2021, 3:16 AMJiang
01/14/2021, 5:06 AMtriggered tf.function retracing in the logs? The restracing is very expensive. Here is the solution: https://github.com/tensorflow/tensorflow/issues/38561#issuecomment-613894395 . @tf.function(experimental_relax_shapes=True)
• When running with micro-batching enabled, bentoml requires about five requests to preheat the model and the batching layer. During this the micro-batching will not take effect. So at the beginning, some 429's are normal. The performance will gradually reach the best.Jackson Hsieh
01/14/2021, 7:28 AMJackson Hsieh
01/14/2021, 10:05 AM