Slackbot
08/29/2023, 2:52 PMAaron Pham
08/29/2023, 3:05 PMAaron Pham
08/29/2023, 3:09 PMIury Alves de Souza
08/29/2023, 3:20 PMAaron Pham
08/29/2023, 3:20 PMIury Alves de Souza
08/29/2023, 3:21 PMIury Alves de Souza
08/29/2023, 3:21 PMAaron Pham
08/29/2023, 3:22 PMbentoml serve you will need to load the model into memory
In this quickstart example, we are running the transformers model, sshleifer/distilbart-cnn-12-6, and depending on the # of replicas you setup, hence the RAM usageIury Alves de Souza
08/29/2023, 3:26 PMThe reason for the ram usage is withyou will need to load the model into memorybentoml serve
In this quickstart example, we are running the transformers model, sshleifer/distilbart-cnn-12-6, and depending on the # of replicas you setup, hence the RAM usageGot it! Thanks!
Iury Alves de Souza
08/29/2023, 3:27 PMAaron Pham
08/29/2023, 4:33 PMIury Alves de Souza
08/30/2023, 3:19 PMAaron Pham
08/30/2023, 3:28 PMHongyi Liu
09/08/2023, 1:21 PMAaron Pham
09/11/2023, 7:29 PMJian Shen Yap
09/11/2023, 7:53 PMHongyi Liu
09/12/2023, 8:14 AMTF_RUN_EAGER_OP_AS_FUNCTION to False. It helped (less memory increase per request) but still the container memory only increase never decrease.Jian Shen Yap
09/14/2023, 10:15 AMHongyi Liu
09/14/2023, 11:01 AMJian Shen Yap
09/14/2023, 11:05 AMHongyi Liu
09/14/2023, 11:15 AMghz -n 20000 --rps=5 --proto /path_to_your_local_proto_file/proto.proto --insecure localhost:1997 --call bentoml.grpc.v1.BentoService/Call -d '{
"apiName": "classify",
"json": {
"photo_ref": "path_to_an_image"
}
}'
2 we don’t know, simply deploy and there is a memory limit, if pod reach the limit it will restart
3 we fix the tensorflow problem mentioned above, but only partly helped (less memory increase each time) still there are leakage…Hongyi Liu
09/14/2023, 11:16 AMJian Shen Yap
09/14/2023, 5:35 PMimport gc; gc.collect()
to see if it cleans up the memory leak issue?Hongyi Liu
09/15/2023, 11:19 AM