Slackbot
08/22/2023, 2:29 PMAsaf Alina
08/22/2023, 2:35 PMexternal_modules is not associated with the “unpickling process” of the scikit_pipeline, but only help with unpickling the custom_objects. is that right?
im asking cuz i was thinking that if i save the model with external_modules=custom_step , then it should solve the CustomStep import issues. but it doesnt seem to be the case…Elior Cohen
08/23/2023, 5:29 AMTom Landman
08/23/2023, 7:18 AMChaoyu
08/23/2023, 8:09 AMi have a feeling like thethat is correct, external_modules is only applicable for custom_objects, not for the model itselfis not associated with the “unpickling process” of theexternal_modules, but only help with unpickling thescikit_pipelinecustom_objects
Chaoyu
08/23/2023, 8:10 AMcustom_step.py file in your Bento, although this requires it has the exact same import path, as when the pipeline was savedChaoyu
08/23/2023, 8:11 AMbentoml.picklable module, which should be able to pickle the entire pipelineElior Cohen
08/23/2023, 8:11 AMElior Cohen
08/23/2023, 8:11 AMElior Cohen
08/23/2023, 8:11 AMElior Cohen
08/23/2023, 8:11 AMElior Cohen
08/23/2023, 8:11 AMChaoyu
08/23/2023, 8:12 AMcustom_step.py file to present, at the environment where you load the modelChaoyu
08/23/2023, 8:12 AMbentoml.picklable is probably an easier way to workaround this issueElior Cohen
08/23/2023, 8:12 AMChaoyu
08/23/2023, 8:13 AMbentoml.picklable will do that for youElior Cohen
08/23/2023, 8:13 AMbentoml.picklable in this context?Chaoyu
08/23/2023, 8:13 AMbentoml.picklable.save_model('my_pipeline', scikit_pipeline)Elior Cohen
08/23/2023, 8:13 AMChaoyu
08/23/2023, 8:14 AMElior Cohen
08/23/2023, 8:14 AMbentoml.picklable?Chaoyu
08/23/2023, 8:14 AMChaoyu
08/23/2023, 8:14 AMbentoml.picklable.load_model to load it backElior Cohen
08/23/2023, 8:14 AMChaoyu
08/23/2023, 8:15 AMbentoml.sklearn module as wellElior Cohen
08/23/2023, 10:17 AMElior Cohen
08/29/2023, 11:53 AMChaoyu
08/29/2023, 5:54 PMJian Shen Yap
08/29/2023, 5:59 PMElior Cohen
09/03/2023, 4:54 AMChaoyu
09/03/2023, 5:59 AMChaoyu
09/03/2023, 6:00 AMElior Cohen
09/03/2023, 6:01 AMElior Cohen
09/03/2023, 6:02 AMElior Cohen
09/03/2023, 6:02 AMexternal_modules is for that, but that didn't workChaoyu
09/03/2023, 7:33 AMElior Cohen
09/03/2023, 7:58 AMElior Cohen
09/03/2023, 10:26 AMElior Cohen
09/03/2023, 10:26 AMimport bentoml
import re
import pandas as pd
from typing import List
from sklearn.calibration import LabelEncoder
from sklearn.base import TransformerMixin, BaseEstimator
from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics import classification_report
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from xgboost import XGBClassifier
from multiprocessing import Pool, cpu_count
from nltk.corpus import stopwords
class TextCleaner(TransformerMixin, BaseEstimator):
"""
A class used to clean textual data.
Attributes
----------
lower : bool
If True, converts all text to lowercase.
remove_punctuation : bool
If True, removes all punctuation from the text.
"""
def __init__(self, lower=True, remove_punctuation=True):
super().__init__()
self.lower = lower
self.remove_punctuation = remove_punctuation
def clean_text(self, text: str) -> str:
"""
Cleans the text.
Parameters
----------
text : str
The text to clean.
Returns
-------
str
The cleaned text.
"""
if type(text) is str:
if self.lower:
text = text.lower()
# If True, punctuation is removed (i.e., any non-alphanumeric character or whitespace).
if self.remove_punctuation:
text = re.sub(r'[^\w\s]', ' ', text)
return text
def transform(self, texts: List[str]) -> List[str]:
"""
Applies the clean_text function to a list of texts in parallel.
Parameters
----------
texts : List[str]
A list of texts to clean.
Returns
-------
List[str]
A list of cleaned texts.
"""
with Pool(cpu_count() -1) as pool:
return pool.starmap(self.clean_text, [(text,) for text in texts])
class TextCleanerWrapper(BaseEstimator, TransformerMixin):
"""
A wrapper for the TextCleaner class to be used in a sklearn pipeline.
Assumes that the provided config file contains the following keys:
- custom_stopwords: a list of custom stopwords to be added to the default stopwords.
- stopwords_languages: a list of languages to be used for the default stopwords.
- hyperparams: a dictionary of hyperparameters to be used in the TextCleaner class.
"""
def __init__(self, **config):
super().__init__()
self.config = config
self.cust_sw = set(config['custom_stopwords'])
self.stopwords_languages = config['stopwords_languages']
self._set_stopwords()
self._text_cleaner = TextCleaner(lower=True, remove_punctuation=True)
def _set_stopwords(self):
"""
Sets the stopwords attribute.
Assumptions:
- In case a language is not installed, it is downloaded from nltk.
- If a language is not supported by nltk, a ValueError is raised.
"""
self.stopwords = set()
for lang in self.stopwords_languages:
try:
if stopwords.words(lang) is None:
stopwords.download(lang)
self.stopwords = self.stopwords.union(set(stopwords.words(lang)))
except ValueError as e:
raise ValueError(f"Language {lang} is not supported by nltk. Please choose a different language. \
You can download the language by running the following command in a python console: \
nltk.corpus.stopwords.download('lang') \
For a list of supported languages, please visit <https://www.nltk.org/book/ch02.html>.") from e
self.stopwords = self.stopwords.union(self.cust_sw)
def fit(self, X, y=None):
"""
Calls transform method.
Parameters
----------
X : pandas.DataFrame
The input data.
y : pandas.Series
The target data.
Returns
-------
pandas.Series
the clean text column.
"""
return self.transform(X)
def transform(self, X):
"""
Cleans the textual features in X as specified in the config file and returns the combined clean text column.
Parameters
----------
X : pandas.DataFrame
The input data.
Returns
-------
pandas.Series
the clean text column.
"""
for col in self.config['textual_features']:
X[f"{self.config['col_prefix']}_{col}"] = self._text_cleaner.transform(X[col].values)
if self.config['combined_text_col'] in X.columns:
X[self.config['combined_text_col']] = X[self.config['combined_text_col']] + '\n' + X[f"{self.config['col_prefix']}_{col}"]
else:
X[self.config['combined_text_col']] = X[f"{self.config['col_prefix']}_{col}"]
return X[self.config['combined_text_col']]
def fit_transform(self, X, y=None):
"""
Calls fit method.
Parameters
----------
X : pandas.DataFrame
The input data.
y : pandas.Series
The target data.
Returns
-------
pandas.Series
the clean text column.
"""
return self.fit(X, y)
if __name__ == '__main__':
newsgroups_train = fetch_20newsgroups(subset='train', categories=['<http://sci.space|sci.space>', 'alt.atheism',])
newsgroups_test = fetch_20newsgroups(subset='test', categories=['<http://sci.space|sci.space>', 'alt.atheism',])
X = pd.DataFrame({'text': newsgroups_train.data, 'label': newsgroups_train.target})
X['subject'] = X['text'].apply(lambda x: x.split('Subject: ')[1].split('Lines: ')[0])
X['body'] = X['text'].apply(lambda x: x.split('Subject: ')[0])
X_train, X_test, y_train, y_test = train_test_split(X[['subject', 'body']], X['label'], test_size=0.1, random_state=42)
label_encoder = LabelEncoder()
y_train_encoded = label_encoder.fit_transform(y_train)
tcw_conf = {'stopwords_languages': ['hebrew', 'english'],
'custom_stopwords': ['hellp', 'world', 'hi'],
'col_prefix': 'clear', 'textual_features': ['subject', 'body'], 'combined_text_col': 'clear_text'}
pipe_params ={'verbose': True}
text_cleaner = TextCleanerWrapper(**tcw_conf)
tfidf_vectorizer = TfidfVectorizer(ngram_range=(1, 2), max_features=200)
xgb_classifier = XGBClassifier()
pipeline = Pipeline(steps=[('text_cleaner', text_cleaner),
('vectorizer', tfidf_vectorizer),
('classifier', xgb_classifier)], **pipe_params)
pipeline.fit(X_train, y_train_encoded)
y_pred = pipeline.predict(X_test)
y_pred_encoded = label_encoder.inverse_transform(y_pred)
print(classification_report(y_test, y_pred_encoded))
local_bento_model = bentoml.picklable_model.save_model("sk_pipeline", pipeline,
signatures={'transform': {"batchable": True, "batch_dim":0},
'predict': {"batchable": True, "batch_dim":0},
'predict_proba': {"batchable": True, "batch_dim":0}},
custom_objects={'label_encoder': label_encoder},
)Jian Shen Yap
09/04/2023, 3:21 PMPicklingError: Can't pickle : attribute lookup TextCleaner on __main__ failedJian Shen Yap
09/04/2023, 10:01 PMwith Pool(cpu_count() -1) as pool:
return pool.starmap(self.clean_text, [(text,) for text in texts])
multiprocessing doesn't really play well with pickling, you may want to consider not using multiprocessing here?Elior Cohen
09/05/2023, 5:21 AMElior Cohen
09/05/2023, 11:22 AMbentoml.pickable_model without the multiprocessing it indeed works.
Sorry about the trouble 🙏 seems like this is not a bento issueJian Shen Yap
09/05/2023, 2:09 PMcloudpickle and sklearn integration uses joblib
Joblib serializes external modules by reference so for sklearn integration will not even work because it needs the source code to be present.
i raised a PR to address this https://github.com/bentoml/BentoML/pull/4168#pullrequestreview-1610063168
as for multiprocessing, it apparently just doesn't work will with pickling. There might be ways to make it work but i didn't dive in and try all the possibilities. maybe you could give it a tryJian Shen Yap
09/05/2023, 2:10 PMElior Cohen
09/06/2023, 7:25 AMre.sub takes a lot of time... In production indeed it is not necessary, but we wanted to avoid having to maintain two separate logic flowsJian Shen Yap
09/07/2023, 3:36 AMElior Cohen
09/07/2023, 5:51 AM