This message was deleted.
# ask-for-help
s
This message was deleted.
f
Any suggestions? cc @Jiang @larme (shenyang)
l
Hi Jason, assuming your model is on GPU and postprocessing code is on CPU, then I'd like to suggest following idea 1 and putting postprocessing code in api server and model inference in runner server. Then scale up the number of api server (because CPU is cheap). In that way if you have multiple requests, every image's postprocessing codes will be run in parallel. For a service endpoint code like below:
Copy code
@svc.api(input=JSON(pydantic_model=SDArgs), output=Image())
async def txt2img(input_data):
    kwargs = input_data.dict()
    res = await sd21_runner.async_run(**kwargs)
    images = res[0]
    return images[0]
Only the line
res = await sd21_runner.async_run(**kwargs)
is doing computation in runner server, all codes around this line is doing computation in api server. So you can put any preprocessing and postprocessing codes there and scale up the api server to remove preprocessing/postprocessing bottleneck
If you use
bentoml serve service:svc
, to scale up the API server you can simply do nothing because in that way BentoML will spin up n API servers where n = number of CPU cores. You can also use
bentoml serve service:svc --api-workers=m
to spin up m api servers
j
Hi, sorry for late response. My team has verified that this is working. Thanks!
💪 1
Hi @larme (shenyang), I have additional question about the bento. I am thinking of running a bento on a managed k8s cluster, thinking of having 4 workers. Should I set the deployment as: Case 1: Min CPU: 1 Max CPU: 4 Case 2: Min CPU: 4 Max CPU: 4 Based on my very limited understanding of k8s, so correct me if I am wrong. Assuming no resource overprovisioning, I am leaning more towards Case 2 as Case 1 might have issue when the bento is handling multiple incoming request, where there might be some delay in worker getting CPU from K8s. Is my concern unfounded? Side note: How does BentoML handle the worker situation when being assigned a non-integer CPU numbers, e.g. 0.5 - 2.5 CPU
l
Hi Janson, if your model support multithreading, BentoML need to setup thread number environment variable for this model, where the thread number is equal to CPU number you assigned to the worker. In that case you should not set CPU number to non-integer value. However if you are just talking about setting CPU numbers to non-integer to API workers which contains no model (the model should be in runner worker in most case), then I guess it will work fine. For k8s questions, @Fog Dong is in a better position to answer them.
j
Yeah, models are all on GPU, CPU is for API workers only
f
Kubernetes will schedule the workload according to the resource
request
, so it is possible that Kubernetes schedule the workload to a node only have 1 cpu if you go with Case 1. If you want to make sure the workload have 4 CPUs, you can go with Case2.
j
Assuming all nodes in my node pool have at least 8 cores, does this change the outcome? @Fog Dong