This message was deleted.
# announcements
s
This message was deleted.
a
I know you can load a MLflow model from a
run_id
(as shown in this notebook) but as I’m working with NLP models a lot, I tend to pack everything up into a pyfunc (model + tokenizer) and can simply load the model using
mlflow.pyfunc.load_model
and its
.predict()
method. It would be ideal to be able to wrap these pyfunc models around a BentoML service and use it for deployment. Technically I can create a custom MLflow artifact and service extracting the
.predict()
from the pyfunc but this is not ideal cause pyfunc forces you to define your predict that way
Copy code
class MyModel(mlflow.pyfunc.PythonModel):
    def load_context(self, context):
        # load your artifacts
    def predict(self, context, model_input):
        return my_predict(model_input.values)
and that
context
parameter is messing up with the
predict()
method I need for BentoML
👍 1
c
PyFunc support is already supported in the 1.0, we are still working on related documentation tho
❤️ 3
Btw here’s related inline documentation on how to work with MLFlow saved models as PyFunc flavor in BentoML https://github.com/bentoml/BentoML/blob/d55741b5932809f9ad8698eeaea5ab9f9abe3c22/bentoml/_internal/frameworks/mlflow.py#L86-L150
🙏 1
a
thanks! I found it and started playing around with the beta version I found an issue with the loading method that I just filed here in github
I’m curious though if I’m saving a pyfunc model with such structure
Copy code
class ModelPyfunc(mlflow.pyfunc.PythonModel):
    
    def load_context(self, context):
        self.model = clf
    
    def predict(self, context, model_input):
        return self.model.predict(model_input)
how can I create a bento service around it? the model itself already contains the predict method, so not sure what would be a working equivalent to this in your quickstart
c
You can import the entire PyFunc model as a BentoML model:
Copy code
bentoml.mlflow.import_from_uri(...)
And then load the PyFunc model back as a Runner and use it in your Service definition:
Copy code
runner = bentoml.mlflow.load_runner("name:version")

my_svc = bentoml.Service('my_service', runners=[runner])

@my_svc.api(...)
def predict(...):
    runner.run( .. )
Basically the
runner.run
will be sending input data to the
predict
function in your PyFunc model definition
and it will run in its own process, and input will be batched