This message was deleted.
# ask-for-help
s
This message was deleted.
c
the default API worker count probably didn’t make sense in your case with a host that have big number of cores
lazily import modules in the service definition file can help as well - try avoid importing large libraries like PyTorch in
service.py
and only do that in the Runner class may help with reducing the memory used by each API server worker as well
m
I only import
Copy code
import bentoml
from <http://bentoml.io|bentoml.io> import Multipart, Image as BImage, JSON

from PIL import Image
c
API workers shouldn’t be taking GPU memory then 🤔
happy to help take a look at your service definition file, feel free to DM me the code if it’s sensitive
m
Maybe bentoml.transformers.get causes implicit imports
👍 1
https://paste.sr.ht/~hummer12007/55d9fbcc2268d6004c0e1b01b9cf3449de588585 (changed save_model to load a quantized 8-bit version since)
Thanks a lot!
c
yes you’re right,
bentoml.transformers.get
does import
transformers
library. We can probably optimize the code to avoid that behavior and only import it in the Runner workers.
I will discuss this more with the team and benchmarks the cost of importing this type of libraries in api workers. For now, I’d suggest limit the number of api workers using either the configuration file or the CLI argument
m
Thanks!