This message was deleted.
# announcements
s
This message was deleted.
o
Even that overall model processing is mostly cpu/gpu intensive tasks it would be great to have async support as sometimes there is a need in quering many external services inside api functions for example.
t
Yes I agree with you Olek. One classic use-case we plan to query some feature store before doing the prediction. Flask 2.0 just got release with some asyncio support. Hopefully this could facilitate the support on BentoML. I would also be curious to hear if it is in the roadmap or others think it would be interesting feature!
👍 2
j
You can handle the query of feature store outside of BentoML. At least that is what I did for my architecture. I have a FastAPI service that retrieves feature from feature store, and send them to bentoML service. Because we use microbatching, every operation in bentoML should be batchable. querying external API is not. Hence i believe a better design is to have a layer before BentoML that does all the feature extractions.
o
Yeah, it is seems like the most reasonable one choice, but would be great to do it using provided toolset without adding another dependencies:)
t
Having another service in front of the inference service adds quite a lot of complexity and increase the response time. We do not want to do that 🙂 And why is querying the external API would not be batchable? you can query all the feature in your batch at the same time I suppose.
s
Another caveat of calling a downstream in the api function is that it could add complexity to the micro-batching logic. Micro-batching uses latency as a proxy for the busy-ness of api server. Latency from a separate downstream (e.g. feature store) could add noise.
Created https://github.com/bentoml/BentoML/issues/1684 @Oleh Kuchuk @Theodore Meynard We'd love to hear your thoughts there.
🙌 1
+ @Jian Shen Yap
j
Personally I think it is an unnecessary complication to the core bentoml logic. What I did was, have another layer that query feature store before hitting bentoML.
That would allow bentoML team to optimize what it does best, which is the microbatching logic