This message was deleted.
# ask-for-help
s
This message was deleted.
j
Hey Kyle! Honestly that is a great question. The short answer is, the LLM ecosystem is still at its infant stage so any practices are yet to stand against the test of time. But we are leaning towards to latter.
load the model on startup
We always load the model into memory on startup, I guess you are referring to "downlaoding" the model on startup instead of packaging inside a docker container. Currently BentoML's default workflow is to package to model into the container itself, but that's a design that was made pre-LLM era. We quickly realised its inefficient to pass massive images across the wire and we are currently working on a workflow that separates the model from the image itself. Curious to hear your opinion on this!
k
Hi Jian! Thanks for the reply. Yes, I'm referring to the 'downloading' step. I'm curious to hear more about this new workflow. Is there a GH issue or doc I could read?
j
it has been discussed internally, and we are still in the midst of finalizing one. so, i'm sorry as of now we don't have any official docs on it yet 😞