can i set a max memory allowed for a runner? if so, what is the expected impact on performance?
I have a naive model ( a python function that reads from a large json file and outputs the corresponding values during an inference call) the packaged bentoml model’s size is only a 1mb (packaged as a bentoml.picklable_model) , but when its used in the bento service, the memory usage seems to bloat up to 6Gb . i have large deep learning models which have bigger sizes that only take up 2Gb of memory at most. So i am wondering why this is the case. Is there a memory leak? or is this expected behaviour?