And can I have the same fast speed with Bento deployment?
l
larme (shenyang)
06/27/2023, 4:33 PM
How does your model perform when using offline? 150 000 000 requests a month is roughly 60 req/s. So I guess for a moderate model we need a BentoCloud/yatai to utilize a k8s cluster.
o
Oleh Kopyl
06/27/2023, 11:42 PM
@larme (shenyang) not using it offline, why?
Oleh Kopyl
06/27/2023, 11:43 PM
@larme (shenyang) so what's the price and request response time going to be?
l
larme (shenyang)
06/28/2023, 3:16 AM
I think there's a lot of factors here. For example how the model performance when doing inference? A model taking 1ms to finish a inference and a model taking 30 seconds to do a inference will have complete different pricing because you need more ec2 (or other cloud provider) instances for the later model