This message was deleted.
# ask-for-help
s
This message was deleted.
c
API server workers and runner workers are controlled differently in BentoML
What does your Runner do? Most of the ML frameworks/libraries can actually take advantage of multiple CPU cores
n
It's running inference but we have a limited memory budget per pod in Kubernetes and having 5 workers makes us hit that limit pretty frequently
Even though we can take advantage of the multithreading I would rather limit that number if possible
If I understand correctly the documentation you linked,
cpu
should define the number of processes to spawn?
-e BENTOML_CONFIG_OPTIONS='runners.resources.cpu=1'
doesn't appear to work either, I still see the messages shown multiple times
Is that what you had in mind @Chaoyu? (thanks for your help with this, much appreciated ๐Ÿ™ )
j
are you seeing 5 workers for runners or the server ?
n
Now that I looked closer at the logs, it appears to be multiple api servers?
Copy code
2023-09-07T16:26:59+0000 [INFO] [api_server:2] Service loaded from Bento directory: bentoml.Service(tag="pipeline-dr-regression-icdr:xd5lqpcnsk5h47qe", path="/home/bentoml/bento/")
2023-09-07T16:27:00+0000 [INFO] [api_server:4] Service loaded from Bento directory: bentoml.Service(tag="pipeline-dr-regression-icdr:xd5lqpcnsk5h47qe", path="/home/bentoml/bento/")
2023-09-07T16:27:00+0000 [INFO] [api_server:5] Service loaded from Bento directory: bentoml.Service(tag="pipeline-dr-regression-icdr:xd5lqpcnsk5h47qe", path="/home/bentoml/bento/")
2023-09-07T16:27:00+0000 [INFO] [api_server:1] Service loaded from Bento directory: bentoml.Service(tag="pipeline-dr-regression-icdr:xd5lqpcnsk5h47qe", path="/home/bentoml/bento/")
2023-09-07T16:27:00+0000 [INFO] [api_server:3] Service loaded from Bento directory: bentoml.Service(tag="pipeline-dr-regression-icdr:xd5lqpcnsk5h47qe", path="/home/bentoml/bento/")
Though I am setting the number of workers to 1 so not sure why it would spawn 5 of them?
j
-e BENTOML_CONFIG_OPTIONS='api_server.workers=1'
hmm. this should work though.
heres a local testing
Copy code
#############################
# ๐Ÿฟ๐Ÿฟ๐Ÿฟ๐Ÿฟ๐Ÿฟ๐Ÿฟ๐Ÿฟ๐Ÿฟ๐Ÿฟ Simple Service
#############################

import bentoml

svc = bentoml.Service("simple")

@svc.api(input=bentoml.io.Text(), output=bentoml.io.Text())
async def count(input_text:str) -> str:
    return "yo"
with a minimal service for testing
n
Let me give it a shot
This seems to work on my local machine, will now test on Kubernetes
Quick question - is the default to have N api server workers and M runner processes? It seems like if I only set one of them I still get multiple versions running concurrently
j
The default is api server process = number of cores. as for runner, it depends on the runner configuration. For For runner, you can see here https://docs.bentoml.com/en/latest/guides/scheduling.html
๐Ÿ™ 1
n
This worked, thank you! ๐Ÿ™