Slackbot
08/26/2023, 5:38 PMJesus Galvan
08/27/2023, 7:03 AMDaniel Holler
08/27/2023, 7:35 AMJesus Galvan
08/28/2023, 6:31 AMJian Shen Yap
08/28/2023, 6:39 AMDaniel Holler
08/28/2023, 6:56 AMmatthew hamerton
08/29/2023, 2:46 PMJian Shen Yap
08/29/2023, 4:20 PMbentoml build should technically be a step that is not memory intensive. It is initializing service.py in the build step, perhaps you are loading the model in the global scope?Jian Shen Yap
08/29/2023, 4:33 PMDaniel Holler
08/29/2023, 4:36 PMimport transformers
import bentoml
model= "TheBloke/WizardLM-1.0-Uncensored-Llama2-13B-GPTQ"
task = "text-generation"
bentoml.transformers.save_model(
task,
transformers.pipeline(task, model=model),
metadata=dict(model_name=model),
)
+
bentoml buildJian Shen Yap
08/29/2023, 4:36 PMservice.py file?Daniel Holler
08/29/2023, 4:37 PMDaniel Holler
08/29/2023, 4:41 PMfrom __future__ import annotations
import bentoml
llm_runner = bentoml.transformers.get("text-generation:latest").to_runner()
svc = bentoml.Service(name="llm-test-service", runners=[llm_runner])
@svc.api(input=bentoml.io.Text(), output=bentoml.io.Text())
async def prompt(input_text: str) -> str:
answer = await llm_runner.generate.async_run(input_text)
return answer[0]["generated_text"]Daniel Holler
08/29/2023, 4:43 PMJian Shen Yap
08/29/2023, 4:44 PMpython download_model.py instead of bentoml build , if I understood it correctly?Daniel Holler
08/29/2023, 4:45 PMDaniel Holler
08/29/2023, 4:47 PMJian Shen Yap
08/29/2023, 4:49 PMDaniel Holler
08/29/2023, 4:51 PMJian Shen Yap
08/29/2023, 9:45 PMDaniel Holler
08/30/2023, 6:10 AMmatthew hamerton
08/30/2023, 1:03 PMJian Shen Yap
08/30/2023, 2:42 PMmatthew hamerton
08/30/2023, 3:02 PMJian Shen Yap
08/30/2023, 3:48 PMJian Shen Yap
08/31/2023, 5:53 PMpipeline.
Just running this will results you the same OOM
import transformers
model= "TheBloke/WizardLM-1.0-Uncensored-Llama2-13B-GPTQ"
task = "text-generation"
pipeline = transformers.pipeline(task, model=model)
I looked a lil deeper and found that its actually pulling the original pytorch llama2 model, which is 40+ GB.
This model will be available in OpenLLM soon!matthew hamerton
09/01/2023, 10:49 AMmatthew hamerton
09/04/2023, 9:00 AMJian Shen Yap
09/04/2023, 2:45 PMpipeline issue, you might have to raise an issue to huggingface since it is an upstream bug. another choice is not to use pipeline but saving the pretrained model directly. https://docs.bentoml.org/en/latest/frameworks/transformers.html#pre-trained-models
bentoml.transformers.save_model takes in either a pretrained transformer object or a pipeline object. do you mind giving it a try?