This message was deleted.
# ask-for-help
s
This message was deleted.
a
Hey there, can u explain abit more ur usecase as well as your code snippet for using this attribute from prometheus?
Idk if these default collector are thread-safe
i
Hi Aron, We are experimenting with BentoML for one of our ML services. We noticed that the Bento application requires 5Gb of ram memory to start. That is without any load (The service is only deployed in a dev environment for now). We would like to get a better understanding to why BentoML is consuming so much memory without any load.
a
Can you provide more information, such as service.py, types of models you are using?
i
I don’t have any specific snippet where a configure those metrics. We are running BentoML based on this quickstart: https://github.com/bentoml/quickstart
I will do, just a sec
a
The reason for the ram usage is with
bentoml serve
you will need to load the model into memory In this quickstart example, we are running the transformers model, sshleifer/distilbart-cnn-12-6, and depending on the # of replicas you setup, hence the RAM usage
i
The reason for the ram usage is with
bentoml serve
you will need to load the model into memory
In this quickstart example, we are running the transformers model, sshleifer/distilbart-cnn-12-6, and depending on the # of replicas you setup, hence the RAM usage
Got it! Thanks!
I asked my coworker for permission to share the model. I am a platform engineer, so I am not sure if I can share the model or not. I will get back to you once I get a response. Thank you for your quick reply!
a
Sounds good lmk when you get the hold of the model name
💯 1
i
Hi Aaron, I created a sample project that mimics the setup my company has: https://github.com/IuryAlves/sample-ml-project I couldn’t use the same model that we use in our service, but I added another one that behaves similarly and also has the same size: https://tfhub.dev/google/imagenet/efficientnet_v2_imagenet1k_s/feature_vector/2
👍 1
a
Ths for the jnfo i will try this out promptly
🙌 1
h
@Aaron Pham hi thanks for the help, I am colleague of @Iury Alves de Souza who is on vacation, wondering is there any new updates on the memory issue?
a
Oh sorry for the late reply, I will ask @Jian Shen Yap to help me with this
👍 1
j
I'll take over
h
Thanks @Jian Shen Yap looking forward for solutions if there is any 😃. We found out that in our bentoml built container, the memory usage always increase linearly with the number of grpc request to it. We initially thought it might be a tensorflow issue since we use TF 2.10, hence, fix following this page and set
TF_RUN_EAGER_OP_AS_FUNCTION
to
False
. It helped (less memory increase per request) but still the container memory only increase never decrease.
j
Hey @Hongyi Liu, let me try to understand the problem. Are you currently trying to tackle the potential memory leak of the model? or is the issue related to the prometheus metrics?
h
@Jian Shen Yap yes, we are trying to tackle the potential memory leak problem 😃
j
Right, let's try to dive into it. I have a few ask: 1. with the example code that you provide, what is the running instruction to reproduce the situation? 2. how are you currently handling the memory leak? 3. what are your current findings regarding this problem? i.e. what have you discovered or potential suspect of the error?
h
1. we simply call the grpc service once per second, similar as load test, and found out that the memory only increase and not going down. you can use tool like ghz doing it. Example code:
Copy code
ghz  -n 20000 --rps=5 --proto /path_to_your_local_proto_file/proto.proto --insecure   localhost:1997 --call bentoml.grpc.v1.BentoService/Call -d '{
   "apiName": "classify",
   "json": {
      "photo_ref": "path_to_an_image"
   }
}'
2 we don’t know, simply deploy and there is a memory limit, if pod reach the limit it will restart 3 we fix the tensorflow problem mentioned above, but only partly helped (less memory increase each time) still there are leakage…
For the code you need to install `ghz`and do some setting, it is the same as sending grpc call manually/script on schedule
j
Hey @Hongyi Liu, do you mind running
import gc; gc.collect()
to see if it cleans up the memory leak issue?
h
@Jian Shen Yap thanks! we tried it before, but it didn’t help…