This message was deleted.
# ask-for-help
s
This message was deleted.
🍱 1
🏁 1
💾 1
a
Hi there, the BENTOML_CONFIG is different from the bento.yaml. The bento.yaml acts as a metadata of any given bento. BENTOML_CONFIG are the configuration of bentoml itself. You can take a look at the configuration from our docs https://docs.bentoml.org/en/latest/guides/configuration.html#configuration You can find the default configuration here https://github.com/bentoml/BentoML/blob/main/src/bentoml/_internal/configuration/default_configuration.yaml To mount the BENTOML_CONFIG to a docker container, you can provide with volume mount
Copy code
d run -e BENTOML_CONFIG=/home/bentoml/bentoml_configuration.yaml -v /path/to/bentoml_configuration.yaml:/home/bentoml/bentoml_configuration.yaml bento:tag ...
s
Also had a question about batching @Aaron Pham. I don't see any docs or sample examples to actually explain the difference between setting batchable= True and batchable=False. Do you have anything that can help?
a
s
I still have some issues understanding the batchable param as normally I write functions to consume and predict on a batch and I had been advised to keep batchable=False where the api was already receiving data in batch. What happens if the function is batch process and we also set batchable to true?
a
can you clarify what do you mean by
function is batch process
? Does this mean the function takes batch inputs? cc @Jiang
j
As mentioned here, it automatically merged data from different requests. When the
batchable=True
, if we call
runner.run([1])
and
runner.run([2])
on the api servers side, on the runner side may bentoml may just call
model.predict([1, 2])
because of the batching optimization.
s
so that means if the runner.run was built to take batch inputs like runner.run([1,2,3,4]) and runner.run([8,2,3,4]) then with batchable=True it would be called as model.predict([1,2,3,4,8,2,3,4]) right? Also From an api response perspective how are the automatically batched inputs then segregated back so the client receives proper response. like if there are 2 clients which make a request to the server at the same time and their inputs get batched , what is the identifier that sends back correct response to the correct client?
j
Yeah
BentoML adaptive batching is supposed to work as you mentioned.
s
makes sense but what is the logic for unmerging of adaptive batched request for client response?
j
It's not a challenge for the dispatcher.
👍 1
s
also for a adaptive batching , we need to mention the max tolerable latency for batching is for example 1000ms - does bentoml automatically calculate based on model inference time whether to batch inputs or does it try adaptive batching at first and incase of failure then makes separate requests? I'm a bit curious of how this works , if there's no documentation available can you point me to the part of source code where this happens. Would really help me out :)
j
The dispatch knows the length for each batch and holds the futures for every request/resonse
does bentoml automatically calculate based on model inference time whether to batch inputs
This one. It will track the RTT of model inferences
s
Makes sense ! Thanks for the clear explanations :)
j
We may have a blog post for the implementation details some day
🍱 1