This message was deleted.
# ask-for-help
s
This message was deleted.
j
I was hoping to just have an Env var to set
Looks like I need to set
BENTOML_CONFIG_OPTIONS='runners.resources.cpu=1'
👍 2
c
Yes that’s right, are you using a built-in runner? Most of the built in runner should be default to one worker
j
Yes that’s right, are you using a built-in runner? Most of the built in runner should be default to one worker
Nope! Using a custom runner which does a feature store interaction as well as an inference call
How can I set the default runner number on a custom runner?
c
Typically we recommend out feature store fetching code in API server and leverage async. And in some cases, feature fetching code in its own runner to leverage batching
Putting feature fetching and model inference in the same runner is not recommended because one is IO intensive the other is compute intensive
You can use the configuration you linked above to change the number of workers, it’s also tied to the scheduling strategy
Did you set the support multi threading flag in runnable class?
j
Typically we recommend out feature store fetching code in API server and leverage async. And in some cases, feature fetching code in its own runner to leverage batching
I'm not sure that this is generally a good recommendation. The Runner is batched and the API worker is not.
If each request to BentoML would result in an API call to the Feature Store it seems clear to me that there is a benefit to batching together these API calls
Which fits nicely into the adaptive batching model
It is additionally worth noting that in our specific usecase out feature data is in an on-disk sqlite database and so the query we run is somewhat cpu bound too.
Did you set the support multi threading flag in runnable class?
SUPPORTS_CPU_MULTI_THREADING = False
j
interesting. embedded db can be seen as a model in some way