Daniel Holler
08/26/2023, 5:32 PM$ bentoml build
Downloading (…)lve/main/config.json: 100%|█████████████████████████████████████████████| 837/837 [00:00<00:00, 8.18MB/s]
Downloading model.safetensors: 100%|███████████████████████████████████████████████| 7.26G/7.26G [18:18<00:00, 6.61MB/s]
Killed
We're guessing this is an OOM error, since our machines don't have enough resources to actually load the model after the download is finished.
Is there a way to build Bentos and have the models and everything be downloaded on the Yatai cluster and skip the part where the local machine has to have everything on it first? We're trying to launch a GPTQ quantized LLM that is hosted on HuggingFace, so a workaround or method for doing this would be really insightful! 🙂