Slackbot
07/07/2023, 2:53 AMFog Dong
07/08/2023, 9:06 AMlarme (shenyang)
07/08/2023, 2:19 PM@svc.api(input=JSON(pydantic_model=SDArgs), output=Image())
async def txt2img(input_data):
kwargs = input_data.dict()
res = await sd21_runner.async_run(**kwargs)
images = res[0]
return images[0]
Only the line res = await sd21_runner.async_run(**kwargs) is doing computation in runner server, all codes around this line is doing computation in api server. So you can put any preprocessing and postprocessing codes there and scale up the api server to remove preprocessing/postprocessing bottlenecklarme (shenyang)
07/08/2023, 2:20 PMbentoml serve service:svc, to scale up the API server you can simply do nothing because in that way BentoML will spin up n API servers where n = number of CPU cores. You can also use bentoml serve service:svc --api-workers=m to spin up m api serversJanson Liew
08/17/2023, 8:20 AMJanson Liew
08/17/2023, 8:35 AMlarme (shenyang)
08/17/2023, 8:53 AMJanson Liew
08/17/2023, 9:29 AMFog Dong
08/17/2023, 9:44 AMrequest , so it is possible that Kubernetes schedule the workload to a node only have 1 cpu if you go with Case 1. If you want to make sure the workload have 4 CPUs, you can go with Case2.Janson Liew
08/17/2023, 9:53 AM