This message was deleted.
# ask-for-help
s
This message was deleted.
c
We recently did one internally! For serving single small model, the performance is quite similar - the only difference is that we need to manually tune a few things to make fastapi to fully utilize the CPU, whereas BentoML does that out of the box. BentoML really starts to show the performance benefits when you are dealing with multi stage model pipelines, running concurrent model ensembles, serving large model that supports batching, or if your business logic includes IO intensive operations, such as fetching data from a database, vector db or feature store or reading a file.
There are a bunch of other important features for AI/ML workloads and workflows that’s simply not available in traditional web framework - for example if you need to do CICD with your model training pipeline on MLflow or if you need to run batch inference