We recently did one internally! For serving single small model, the performance is quite similar - the only difference is that we need to manually tune a few things to make fastapi to fully utilize the CPU, whereas BentoML does that out of the box.
BentoML really starts to show the performance benefits when you are dealing with multi stage model pipelines, running concurrent model ensembles, serving large model that supports batching, or if your business logic includes IO intensive operations, such as fetching data from a database, vector db or feature store or reading a file.