Or is a single model shared across multiple runner workers
Visaal Ambalam
09/07/2023, 6:31 PM
Trying to see how much memory to allocate in our deployments, thanks in advance!
c
Chaoyu
09/07/2023, 6:32 PM
One model instance per runner replica! Sharing one instance across multiple python processes is possible but is very hacky and can lead to many production issues 😊
gratitude thank you 1
Chaoyu
09/07/2023, 6:33 PM
Btw this is just the default behavior, BentoML is flexible enough that users can actually customize runners’ scheduling strategy as well