Slackbot
08/01/2023, 6:02 AMChaoyu
08/01/2023, 5:42 PMsklearn.ensemble.VotingClassifier directly in one runner. The inference graph pattern is typically meant for more heavy models, e.g. using an LLM to augment the prompt for a stable diffusion model, or chaining multiple NLP modelsChaoyu
08/01/2023, 5:43 PMjiewpeng
08/01/2023, 11:15 PMChaoyu
08/01/2023, 11:39 PMChaoyu
08/01/2023, 11:39 PMjiewpeng
08/02/2023, 12:59 AMGot it, in that case you may need custom code to define how to combine the scores from downstream models, is that the right understanding?yes that's right, and these models may not all use the same ml framework (e.g. some might use pytorch, others use sklearn etc.)
Any suggestions how we may improve this API?Unfortunately no, that's why i started this thread to see if anyone else had done something like this and how they dealt with it.
Chaoyu
08/02/2023, 1:33 AMHowever, if we take this self-contained model and deploy it as a bento, it is not consistent with the pattern of separate runners for each model, and in my experience, results in poorer prediction speed.Do you mind explain what you mean it’s not consistent with the pattern of separate runners for each model?
Chaoyu
08/02/2023, 1:34 AMChaoyu
08/02/2023, 1:36 AMjiewpeng
08/02/2023, 1:39 AMDo you mind explain what you mean it’s not consistent with the pattern of separate runners for each model?i meant that my model is deployed as a single runner, where the runner points to a model e.g. pytorch
nn.Module which contains all the models required, and this has a predict method exposed to the runner that calls all the underlying models. This is different from what is suggested in the inference graph doc, where each model is deployed as a separate runner.