This message was deleted.
# ask-for-help
s
This message was deleted.
c
Hi @jiewpeng - when doing smaller ensemble models, I’d recommend using something like
sklearn.ensemble.VotingClassifier
directly in one runner. The inference graph pattern is typically meant for more heavy models, e.g. using an LLM to augment the prompt for a stable diffusion model, or chaining multiple NLP models
And yes you’re right, the runner architecture is a distributed framework at core, and it comes with additional overhead. That’s why for smaller models, using one Runner for the entire ensemble would make more sense.
j
For my use case, it sometimes involves a larger model - my team does search and recommendations, where multiple scorers feeding to a final learn-to-rank model is becoming quite a common pattern. In this pattern, each model is not always small e.g. we might use an LLM as one of the scorers.
c
Got it, in that case you may need custom code to define how to combine the scores from downstream models, is that the right understanding?
Any suggestions how we may improve this API?
j
Got it, in that case you may need custom code to define how to combine the scores from downstream models, is that the right understanding?
yes that's right, and these models may not all use the same ml framework (e.g. some might use pytorch, others use sklearn etc.)
Any suggestions how we may improve this API?
Unfortunately no, that's why i started this thread to see if anyone else had done something like this and how they dealt with it.
c
However, if we take this self-contained model and deploy it as a bento, it is not consistent with the pattern of separate runners for each model, and in my experience, results in poorer prediction speed.
Do you mind explain what you mean it’s not consistent with the pattern of separate runners for each model?
If you use what’s suggested in the inference graph doc, each model are in fact scheduled separately from each other
This pattern is consistent across different BentoML deployment patterns too - in simple docker deployment, all model runners will be running in completely isolated processes. In Yatai/Kubernetes deployment or when deploying to BentoCloud, they will be orchestrated as separate micro-services.
j
Do you mind explain what you mean it’s not consistent with the pattern of separate runners for each model?
i meant that my model is deployed as a single runner, where the runner points to a model e.g. pytorch
nn.Module
which contains all the models required, and this has a
predict
method exposed to the runner that calls all the underlying models. This is different from what is suggested in the inference graph doc, where each model is deployed as a separate runner.