This message was deleted.
# ask-for-help
s
This message was deleted.
a
in addition, i have a feeling like the
external_modules
is not associated with the “unpickling process” of the
scikit_pipeline
, but only help with unpickling the
custom_objects
. is that right? im asking cuz i was thinking that if i save the model with
external_modules=custom_step
, then it should solve the
CustomStep
import issues. but it doesnt seem to be the case…
e
Listening in, interested in the same problem
t
I'm experiencing the same issue. If this bug is able to be resolved, it would be helpful.
c
i have a feeling like the
external_modules
is not associated with the “unpickling process” of the
scikit_pipeline
, but only help with unpickling the
custom_objects
that is correct, external_modules is only applicable for custom_objects, not for the model itself
you can try include the
custom_step.py
file in your Bento, although this requires it has the exact same import path, as when the pipeline was saved
another workaround is to use
bentoml.picklable
module, which should be able to pickle the entire pipeline
e
I'm experiencing this regardless of the bento, it's a problem that appears when I use Bento models
even before constructing the service
so I create a bento model in one environment, and try to load it in another environment
i intend to build the service in the second environment
but I can't get the model to work
c
yes - basically you need the
custom_step.py
file to present, at the environment where you load the model
bentoml.picklable
is probably an easier way to workaround this issue
e
My goal was to separate the research environment, where the Bento model is created, from the serving environment. I (naively?) assumed that its possible to only pass the Bento Model from one environmetn to another and have it work
c
understood,
bentoml.picklable
will do that for you
e
how should I use
bentoml.picklable
in this context?
c
bentoml.picklable.save_model('my_pipeline', scikit_pipeline)
e
what I'm savign now is a skleanr pipeline
c
yes
e
so use the same pipeline, just save it with
bentoml.picklable
?
c
it pickles the additoinal classes it needs to deserialize it
yes, and
bentoml.picklable.load_model
to load it back
e
👍 let me see
c
we should probably bring this to
bentoml.sklearn
module as well
e
@Chaoyu it didn't work, we get the same error on the custom step
🥲
c
Sorry I didn’t follow up sooner, @Jian Shen Yap could you help take a look?
j
sure, let me take it from here @Chaoyu
e
any missing info from our side? the basic problem is that when we save a sklearn pipeline with a custom transformer (custom being our own sub-class of an sklearn transformer) as a Bento Model, we can't seem to be able to load it in a different environment
c
@Jian Shen Yap is working on a minimal reproducible example to experiment with it more
👀 1
The workaround is to include the Python file in your Bento. I understand this doesn’t allow you to reproduce the model without a Bento and other files.
e
we are under the contraint where in one environment we create the model (bento model is the only output of that environment) and the other environment creates the bento itself (bento service) I know with the bentofile (service) I can include more python files, but I don't have that option with bento models
I coulnd't find a way to bundle extra python files with the Bento Model
at first we figured out that maybe the
external_modules
is for that, but that didn't work
c
@Elior Cohen do you mind share with me the related code for me to reproduce on my end?
e
let me check
@Chaoyu here you go, this is the code.
Copy code
import bentoml
import re
import pandas as pd

from typing import List
from sklearn.calibration import LabelEncoder
from sklearn.base import TransformerMixin, BaseEstimator
from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics import classification_report
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from xgboost import XGBClassifier
from multiprocessing import Pool, cpu_count
from nltk.corpus import stopwords


class TextCleaner(TransformerMixin, BaseEstimator):
    """
    A class used to clean textual data.
    
    Attributes
    ----------
    lower : bool
        If True, converts all text to lowercase.
    remove_punctuation : bool
        If True, removes all punctuation from the text.
    """
    
    def __init__(self, lower=True, remove_punctuation=True):
        super().__init__()
        self.lower = lower
        self.remove_punctuation = remove_punctuation
     

    def clean_text(self, text: str) -> str:
        """
        Cleans the text.

        Parameters
        ----------
        text : str
            The text to clean.
            
        Returns
        -------
        str
            The cleaned text.
        """
        if type(text) is str:
            if self.lower:
                text = text.lower()
            
            # If True, punctuation is removed (i.e., any non-alphanumeric character or whitespace).
            if self.remove_punctuation:
                text = re.sub(r'[^\w\s]', ' ', text)
          
        return text

    def transform(self, texts: List[str]) -> List[str]:
        """
        Applies the clean_text function to a list of texts in parallel.

        Parameters
        ----------
        texts : List[str]
            A list of texts to clean.
        
        Returns
        -------
        List[str]
            A list of cleaned texts.        
        """
        with Pool(cpu_count() -1) as pool:
            return pool.starmap(self.clean_text, [(text,) for text in texts])


class TextCleanerWrapper(BaseEstimator, TransformerMixin):
    """
    A wrapper for the TextCleaner class to be used in a sklearn pipeline.
    Assumes that the provided config file contains the following keys:
    - custom_stopwords: a list of custom stopwords to be added to the default stopwords.
    - stopwords_languages: a list of languages to be used for the default stopwords.
    - hyperparams: a dictionary of hyperparameters to be used in the TextCleaner class. 
    """
    def __init__(self, **config):
        super().__init__()
        self.config = config
        self.cust_sw = set(config['custom_stopwords'])
        self.stopwords_languages = config['stopwords_languages']
        self._set_stopwords()
        self._text_cleaner = TextCleaner(lower=True, remove_punctuation=True)
        
    def _set_stopwords(self):
        """
        Sets the stopwords attribute.
        Assumptions:
        - In case a language is not installed, it is downloaded from nltk.
        - If a language is not supported by nltk, a ValueError is raised.
        """
        self.stopwords = set()
        for lang in self.stopwords_languages:
            try:
                if stopwords.words(lang) is None:
                    stopwords.download(lang)
                self.stopwords = self.stopwords.union(set(stopwords.words(lang)))
                
            except ValueError as e:
                raise ValueError(f"Language {lang} is not supported by nltk. Please choose a different language. \
                                   You can download the language by running the following command in a python console: \
                                   nltk.corpus.stopwords.download('lang') \
                                   For a list of supported languages, please visit <https://www.nltk.org/book/ch02.html>.") from e
        self.stopwords = self.stopwords.union(self.cust_sw)
        

    def fit(self, X, y=None):
        """
        Calls transform method.

        Parameters
        ----------
        X : pandas.DataFrame
            The input data.

        y : pandas.Series
            The target data.
        
        Returns
        -------
        pandas.Series
            the clean text column.
        """
        return self.transform(X)
        
        
    def transform(self, X):
        """
        Cleans the textual features in X as specified in the config file and returns the combined clean text column.

        Parameters
        ----------
        X : pandas.DataFrame
            The input data.

        Returns
        -------
        pandas.Series
            the clean text column.
        """
        for col in self.config['textual_features']:
            X[f"{self.config['col_prefix']}_{col}"] = self._text_cleaner.transform(X[col].values)
            if self.config['combined_text_col'] in X.columns:
                X[self.config['combined_text_col']] = X[self.config['combined_text_col']] + '\n' + X[f"{self.config['col_prefix']}_{col}"]
            else:
                X[self.config['combined_text_col']] = X[f"{self.config['col_prefix']}_{col}"]

        return X[self.config['combined_text_col']]
    
    def fit_transform(self, X, y=None):
        """
        Calls fit method.

        Parameters
        ----------
        X : pandas.DataFrame
            The input data.

        y : pandas.Series
            The target data.

        Returns
        -------
        pandas.Series
            the clean text column.
        """
        return self.fit(X, y)
    

if __name__ == '__main__':

    newsgroups_train = fetch_20newsgroups(subset='train', categories=['<http://sci.space|sci.space>', 'alt.atheism',])
    newsgroups_test = fetch_20newsgroups(subset='test', categories=['<http://sci.space|sci.space>', 'alt.atheism',])

    X = pd.DataFrame({'text': newsgroups_train.data, 'label': newsgroups_train.target})
    X['subject'] = X['text'].apply(lambda x: x.split('Subject: ')[1].split('Lines: ')[0])
    X['body'] = X['text'].apply(lambda x: x.split('Subject: ')[0])

    X_train, X_test, y_train, y_test = train_test_split(X[['subject', 'body']], X['label'], test_size=0.1, random_state=42)

    label_encoder = LabelEncoder()
    y_train_encoded = label_encoder.fit_transform(y_train)

    
    tcw_conf = {'stopwords_languages': ['hebrew', 'english'],
                'custom_stopwords': ['hellp', 'world', 'hi'],
                'col_prefix': 'clear', 'textual_features': ['subject', 'body'],  'combined_text_col': 'clear_text'}
    
    pipe_params ={'verbose': True}

    text_cleaner = TextCleanerWrapper(**tcw_conf)
    tfidf_vectorizer = TfidfVectorizer(ngram_range=(1, 2), max_features=200)
    xgb_classifier = XGBClassifier()


    pipeline = Pipeline(steps=[('text_cleaner', text_cleaner),
                               ('vectorizer', tfidf_vectorizer),
                               ('classifier', xgb_classifier)], **pipe_params)
    

    pipeline.fit(X_train, y_train_encoded)

    y_pred = pipeline.predict(X_test)
    y_pred_encoded = label_encoder.inverse_transform(y_pred)
    print(classification_report(y_test, y_pred_encoded))


    local_bento_model = bentoml.picklable_model.save_model("sk_pipeline", pipeline,
                                                       signatures={'transform': {"batchable": True, "batch_dim":0},
                                                                    'predict': {"batchable": True, "batch_dim":0},
                                                                    'predict_proba': {"batchable": True, "batch_dim":0}},         
                                                       custom_objects={'label_encoder': label_encoder},
                                                       )
j
@Elior Cohen is this the error that you are facing?
Copy code
PicklingError: Can't pickle : attribute lookup TextCleaner on __main__ failed
hey @Elior Cohen, I spent some hours looking into this. it looks like
Copy code
with Pool(cpu_count() -1) as pool:
            return pool.starmap(self.clean_text, [(text,) for text in texts])
multiprocessing doesn't really play well with pickling, you may want to consider not using multiprocessing here?
e
It's that or another pickling error Though I'm 99% sure we also went through this pool thing suspicion, we removed it and it still didn't work - right @Tom Landman? @Jian Shen Yap
Sorry for misleading you @Jian Shen Yap Seems like you are right, when using
bentoml.pickable_model
without the multiprocessing it indeed works. Sorry about the trouble 🙏 seems like this is not a bento issue
j
No worries @Elior Cohen! took me abit to play around too. the difference was, pickable model uses
cloudpickle
and sklearn integration uses
joblib
Joblib serializes external modules by reference so for sklearn integration will not even work because it needs the source code to be present. i raised a PR to address this https://github.com/bentoml/BentoML/pull/4168#pullrequestreview-1610063168 as for multiprocessing, it apparently just doesn't work will with pickling. There might be ways to make it work but i didn't dive in and try all the possibilities. maybe you could give it a try
for your job, i think the task should be simple enough that i think multiprocessing might introduce more overhead than just doing it in one process. perhaps you could benchmark that with your use case
e
We did, in research since we work with high loads of data it really makes the difference, especially because
re.sub
takes a lot of time... In production indeed it is not necessary, but we wanted to avoid having to maintain two separate logic flows
j
right, perhaps your team could do a few experiments for serializing multiprocessing code, maybe we might find something that works. for this i think you could maybe try multiprocessing on basic python class/object/function and try pickling with cloud pickle. if there's something that works then it will likely work with the sklearn pipeline too! I would be interested to hear your findings!
e
Will let you know if we find anything interesting