Slackbot
09/09/2023, 4:50 PMJian Shen Yap
09/10/2023, 4:38 PMload the model on startupWe always load the model into memory on startup, I guess you are referring to "downlaoding" the model on startup instead of packaging inside a docker container. Currently BentoML's default workflow is to package to model into the container itself, but that's a design that was made pre-LLM era. We quickly realised its inefficient to pass massive images across the wire and we are currently working on a workflow that separates the model from the image itself. Curious to hear your opinion on this!
Kyle White
09/11/2023, 3:38 PMJian Shen Yap
09/11/2023, 6:19 PM