Slackbot
10/10/2023, 1:26 PMAmit Gelber
10/10/2023, 3:25 PMupscaling_scheduler = EulerDiscreteScheduler.from_pretrained("stabilityai/stable-diffusion-x4-upscaler", subfolder="scheduler")
pipe = StableDiffusionUpscalePipeline.from_pretrained(
"stabilityai/stable-diffusion-x4-upscaler",
torch_dtype=torch.float16,
scheduler=upscaling_scheduler
)
bentoml.diffusers.save_model("upscale_model", pipe)Jian Shen Yap
10/10/2023, 4:39 PMAmit Gelber
10/11/2023, 1:12 PMAmit Gelber
10/11/2023, 2:05 PMbentoml.diffusersload_model("upscale_model",pipeline_class=StableDiffusionUpscalePipeline,
torch_dtype=torch.float16,
device_id="cuda")larme (shenyang)
10/11/2023, 6:14 PMbento_model = bentoml.diffusers.get("upscale_model:latest")
runner = bento_model.to_runner()
then the runner will automatically do optimization like using half precision and put model to GPU. It can be used as computation unit. You can test it locally by
runner.init_local()
runner.run(prompt="a bento box", ...)
Later you can use the runner inside a bentoml service, this will run the model in a separated process. So you can have multiple api workers handling image pre-processing and post-processing utilizing CPU resource, while runner process utilizing GPU resource for image generation/upscaling
one example may related to your usage is at:
https://github.com/bentoml/diffusers-examples/blob/main/sd2_with_upscaler/service.py