This message was deleted.
# ask-for-help
s
This message was deleted.
j
We also noticed that our pods were getting killed due to OOM. We had to increase resources quite a bit, specifically memory. Our pods run at a constant 4.6 GB with no load. And our model was much smaller, ~80 MB. We have only just started experimenting with Bento, so we are still new.
d
Hey there! That's interesting. Have you been able at all to load larger models like a few GB, or does it break for your team as well on instances that can usually run it? What types of models do you use? We're trying to run quantized LLM models, so we might have similar issues if you're doing LLMs too?
j
No, we have not used it for larger models. We are not using LLMs, we are using Bento with a couple of Tensorflow classification models.
j
Hey all, we are currently looking into this! will circle back once we have some clarity
d
Thank you very much @Jian Shen Yap 👍👍
m
@Jian Shen Yap & team, any update on the above ? 🙂
j
Hey guys,
bentoml build
should technically be a step that is not memory intensive. It is initializing
service.py
in the build step, perhaps you are loading the model in the global scope?
@matthew hamerton do you mind sharing where you are running into problem?
d
Hello @Jian Shen Yap, this is the exact code we are attempting to run, and it gets "Killed" every time, after it occupies all memory of the computer after downloading.
Copy code
import transformers
import bentoml

model= "TheBloke/WizardLM-1.0-Uncensored-Llama2-13B-GPTQ"
task = "text-generation"

bentoml.transformers.save_model(
    task,
    transformers.pipeline(task, model=model),
    metadata=dict(model_name=model),
)
+
Copy code
bentoml build
j
@Daniel Holler can you share your
service.py
file?
d
Yes, one second
This is service.py, but we can't run this since download_model can't finish:
Copy code
from __future__ import annotations

import bentoml

llm_runner = bentoml.transformers.get("text-generation:latest").to_runner()

svc = bentoml.Service(name="llm-test-service", runners=[llm_runner])

@svc.api(input=bentoml.io.Text(), output=bentoml.io.Text())
async def prompt(input_text: str) -> str:
    answer = await llm_runner.generate.async_run(input_text)
    return answer[0]["generated_text"]
When downloading the model to local env, it downloads the 7.9GB model, and then once that is done downloading, it just says "Killed". When I retry, I see in task manager that it's taking up all my memory (32 GB ram) and right as it maxes out, this is when it says "Killed". I may have been mistaken, we are running "python download_model.py", the code being the first of the two snippets I just sent
j
ah okay, so your steps stops at
python download_model.py
instead of
bentoml build
, if I understood it correctly?
d
Yes you are correct. We had to do this, since we had some issues getting a new Runner instance in service.py to detect the HuggingFace model id we gave it
Basically we're trying to combine two examples on the BentoML docs to get a demo LLM from HuggingFace up and running: https://docs.bentoml.org/en/latest/quickstarts/deploy-a-transformer-model-with-bentoml.html + https://docs.bentoml.org/en/latest/quickstarts/deploy-an-iris-classification-model-with-bentoml.html We're using the transformers module since that's relevant for the LLM, but we're downloading the model from HF, like in the Iris classification example.
j
got it. let me try to reproduce it, downloading stage should not eat up the memory, otherwise it should be a bug
d
Okay, thank you! It's only the download part that doesn't work. It works totally fine for the Iris classification example, and all parts after the successful download works. We're working with LLMs, so the one in my code snippet takes up almost 8 GB, as opposed to the Iris example <100MB where the download part works fine.
j
had a 64gb machine and it indeed took up all the memory. will circle back in abit and see what causes the memory leak
d
Gotcha, thank you @Jian Shen Yap. Looking forward 🙂
m
Thanks Jian and Daniel. Just saw this - @Jian Shen Yap would OpenLLM be production ready ? Didn't expect to have this issue here
j
@matthew hamerton we already have users that is using OpenLLM for production. We are going to release a blog/guide on how to deploy Llama-2 with BentoCloud.
m
Cheers - wouldn't the above OOM issue affect them too ?
j
we are currently stilll investigating, we suspect its a upstream bug from HF. looking at the model artifacts that is downloaded, it downloaded the original pytorhc version which is way larger
Hi @Daniel Holler, circling back to this,its apparently a huggingface issue where that particular model is not implemented correctly with
pipeline
. Just running this will results you the same OOM
Copy code
import transformers
model= "TheBloke/WizardLM-1.0-Uncensored-Llama2-13B-GPTQ"
task = "text-generation"

pipeline = transformers.pipeline(task, model=model)
I looked a lil deeper and found that its actually pulling the original pytorch llama2 model, which is 40+ GB. This model will be available in OpenLLM soon!
m
Thanks @Jian Shen Yap - what would be the next step for this ? How would this model or similar be loaded ? How could we contribute to fixing it ?
Hey @Jian Shen Yap - appreciate you're busy, but would you have any update on the above ? Would be of help for the community
j
hey, for the
pipeline
issue, you might have to raise an issue to huggingface since it is an upstream bug. another choice is not to use pipeline but saving the pretrained model directly. https://docs.bentoml.org/en/latest/frameworks/transformers.html#pre-trained-models bentoml.transformers.save_model takes in either a pretrained transformer object or a pipeline object. do you mind giving it a try?