Hey! We're running into an issue when attempting t...
# ask-for-help
d
Hey! We're running into an issue when attempting to build a Bento on our machines to deploy to our Yatai infrastructure and a HuggingFace model (not on the default models list).
Copy code
$ bentoml build
Downloading (…)lve/main/config.json: 100%|█████████████████████████████████████████████| 837/837 [00:00<00:00, 8.18MB/s]
Downloading model.safetensors: 100%|███████████████████████████████████████████████| 7.26G/7.26G [18:18<00:00, 6.61MB/s]
Killed
We're guessing this is an OOM error, since our machines don't have enough resources to actually load the model after the download is finished. Is there a way to build Bentos and have the models and everything be downloaded on the Yatai cluster and skip the part where the local machine has to have everything on it first? We're trying to launch a GPTQ quantized LLM that is hosted on HuggingFace, so a workaround or method for doing this would be really insightful! 🙂