This message was deleted.
# ask-for-help
s
This message was deleted.
o
Decided to use my head and search this slack for answers. I have some paths to explore now, looks like lazy loading models is supported out of the box. Thanks!
j
What model do you want to use? @Aaron Pham maybe you could help with some suggestions too
o
I’m starting with
databricks/dolly-v2-3b
I’m open to using other models though, I’m not sure which one will be best for my use case yet
To explain a little more, I want to deploy my bento container to a gpu accelerated kubernetes node group, and load the model at runtime via and EFS volume. This way I can reduce the image size
c
One option is to configure the node with dependencies and run openllm start command directly, without packing all dependencies and model files into the image. Fast scaling of large container images on GPU cluster needs a stack of complicated optimizations and that’s one of the reasons we built BentoCloud :)