This message was deleted.
# ask-for-help
s
This message was deleted.
👍 2
e
Some of the questions I think we'd discuss with a TAM:
• Our model registry scenario is pretty much a blank slate. Currently we train models wiith Metaflow and save weights/pickles in S3 with a run ID. But this has limitations. MLFlow allows you to tag models. I saw the MLFlow/BentoML docs. I'd love to talk model registry strategy with a knowledgable BentoML person
• we do have much need for batch inference. That's probably our main deployment category. Online inference is more the exception for us. I can see BentoML potentially helping for serialization/deserialization, but is there more? Maybe some fanciness with runnables? I've heard there's an offline mode coming which I'm super excited for. (This is totally a duplicate question I've asked here before so no need to spend time on it here 🙂 )
• Can we decouple model training from model serving? I've asked this here, too. But I didn't see the any indication of model requirements in the
model
artifact. They all seem to be declared in the
bentofile.yaml
which is the full
bento
. And you can't seem to have a bento without a
bento.Service
,. To me that means that the REST API code is tightly coupled to the model code. What if we wrote a really nice service with JWT authentication, fancy authorization, and then hooked it up with an image classifier? I could see that classifier being implmented in sklearn one day, pytorch the next, and tensorflow the day after that. Assuming we have an interface defined (we pass in an image and get a vector of length 10), I feel like this should be achievable. You'd definitely need to rebuild the service container if you grabbed a model with a different implementation--reason being that different models need different inference code. But besides that, Bento seems to make defininng a unified interface really easu. I'd ask a TAM if they've seen other teams solve this problem.
• Oh another one: we're hooking up model monitoring now--tracking our model inputs and outputs over time. Looking for missing data, values out of the range of training, that sort of thing. If an error surfaces from monitoring, we'll probably want to ad-hoc investigate it or trigger a rollback or rebuild. That's sort of an observability/DevOps question, but with nuances related to ML. A Q&A session with a TAM would be so helpful as we reason about how we should set up that CD system. Not that Bento needs to provide a monitoring solution, we know that's not the scope. But Bento is great at deployment so it's their opinion I'd really want in this situation.
I'm not sure how you'd classify these questions. Architecture? Strategy? Implementation? Basically all three 😄
t
Hi @Eric Riddoch! Great questions! We're actually working on most of that. New monitoring architecture is going to be released in about a week, batch is coming behind that but is close. Definitely have different model repository strategies (one of yatai's features is a model repository) Let me DM and we can grab a call to go through all of this 🙂
s
Love the questions! Many of these questions are relevant to me. I will share my 2 cents and experience here for reference.
How does BentoML (the company) make money?
Is it just through a hosted Yatai solution? We probably wouldn’t use that at BEN since we’re so ECS focused.
I don’t see Porch to adopt the full Yatai solution either. But, we are interested in model registry which is part of Yatai. I really like BentoML as a generic solution that can be dockerized and deployed to any platform. We are using Kubernetes (GKE). BentoML saves me a lot of time in model deployment and scaling. Because it integrates with our existing infrastructure and deployment orchestration platform, it allows other teams to contribute to our model services and scale them using their Kubernetes knowledge.
But I could absolutely see us paying for premium support
I was thinking along this line of reasoning too. Since we don’t use Yatai, can we pay for their consulting. I want BentoML to be around and continue to innovate 🙂
Can we decouple model training from model serving?
It is sort of decoupled for us. We are using a managed Flyte service (https://www.union.ai/) to train our models. We save the model artifacts (pickle or pytorch.bin) to a GCS bucket with the run id (this is where model registry can improve our current stack). On the CI/CD deployment side, we load the artifacts using the framework and then save it using bentoml python API to produce bento model artifacts. In the CI/CD we use bentoml CLI to build a bento, export the bento, and then dockerize it using our container script.
To me that means that the REST API code is tightly coupled to the model code. What if we wrote a really nice service with JWT authentication, fancy authorization, and then hooked it up with an image classifier?
In my opinion, the coupling is necessary in order to deliver the performance without us worrying about the implementations. But, we don’t have to couple our business logic to the framework. It is like any Web App frameworks where we have a good hexagonal design to separate our concerns. If we have a security module, it can be layered into BentoML as a service or library. There is no conflict. On AWS, that is to utilize API Gateway to implement all kinds of public facing rules and security that is totally decoupled from the BentoML service itself. I don’t think there is a conflict there, unless you have a unique use case that I am not aware of.
Oh another one: we’re hooking up model monitoring now
Ha! I made the same request to the BentoML community. I am glad to hear it is coming soon. Also, I requested A/B testing capability which is an important part in model deployment.
t
@Shihgian Lee awesome reply! With the monitoring features out so soon, I was going to reach out to sync up with you on the A/B testing design talks that we've been having. I'll dm you, would love to catch up over a call
❤️ 2