This message was deleted.
# ask-for-help
s
This message was deleted.
g
bentoml.sklearn.save_model('stroke_prediction', model, custom_objects={ 'dicVectorizer': dv })
service.py: import bentoml from bentoml.io import JSON model_ref = bentoml.sklearn.get('stroke_prediction:latest') dv = model_ref.custom_objects['dicVectorizer'] model_runner = model_ref.to_runner() svc = bentoml.Service('stroke_prediction', runners=[model_runner]) @svc.api(input= JSON() , output= JSON()) def classify(application_data): vector = dv.transform(application_data) prediciton = model_runner.predict.run(vector) return prediction
Could the issue be that I am inputting JSON and not a python dictionary?
m
Hard to say without seeing how you created the model.
g
Here is what I did: Data: df_train, df_test = train_test_split(df, test_size=0.2, random_state=1) y_train = df_train.stroke.values del df_train['stroke'] Train: train_dict = df_train.to_dict(orient='records') dv = DictVectorizer(sparse=False) x_train = dv.fit_transform(train_dict) smote (data imbalanced, this adds more rows of the minority data): x_train, y_train = training_smote(x_train, y_train) model = LogisticRegression(max_iter=1000, random_state=1,C=100,penalty='l2',solver='lbfgs') model.fit(x_train, y_train)
After that, I began the bento process above.
m
🤔
I ended up removing tons records to balance my dataset. When I tried to generate dummy ones I noticed some correlation changes. The smote part is the only one that stands out to me. Examine what it returned. I bet it doesn’t match up with your target now.
g
I am not sure if I understand what you mean but I had also checked with: len(x_train), len(y_train) output (7798, 7798)
👍 1
Also running this in the notebook: y_pred = model.predict(x_test) seen to run fine. A check and i got len(x_test), len(y_pred) output (1022, 1022)
m
I had similar errors but during model eval and it was always because I was trying to predict with my training set and scoring with my test set. Once I had those marched up right, it worked.
Sorry not much help.
gratitude thank you 2
g
I appreciate the effort
m
Keep at it. You’ll get it. It helps if you walk away for a bit and think of something else for a while. Come back with fresh eyes.
👍 1
g
I should add inputting into the local host:3000 a patient right from test, { "gender": "Male", "age": 69.0, "hypertension": 0, "heart_disease": 1, "ever_married": 1, "work_type": "self_employed", "residence_type": "Urban", "avg_glucose_level": 195.23, "bmi": 28.3, "smoking_status": "smokes", "obese": 0, "clearly_diabetes": 1 }
Eyeball and brain break will update if I have a stroke of genius 🤣
m
That’s what 12 features?
Take a look at dv.feature_names_out() and count what you have there.
g
14, but DictVectorizer makes it 21. I checked by pulling a row from post DictVectorizer .transform
x_test[21] output: array([ 69. , 195.23, 28.3 , 1. , 1. , 0. , 1. , 1. , 0. , 0. , 0. , 1. , 0. , 0. , 0. , 1. , 0. , 0. , 0. , 1. , 0. ])
m
Getting 12:
Copy code
d = {
    "gender": "Male",
    "age": 69.0,
    "hypertension": 0,
    "heart_disease": 1,
    "ever_married": 1,
    "work_type": "self_employed",
    "residence_type": "Urban",
    "avg_glucose_level": 195.23,
    "bmi": 28.3,
    "smoking_status": "smokes",
    "obese": 0,
    "clearly_diabetes": 1,
}
k = d.keys()
print(len(k))
g
You are correct, not sure why I thought it was 14.
👍 1
@Martin Uribe OMG soooo... it turns out my stroke_service.py was not updating from jupyter notebook cell with "%%writefile stroke_service.py" because the magic code was not at the top of the cell. I had a Google Colab title function above it. It now runs as it should. Thank you again for your time on this.
m
Wow, good to know. I try not to use those magic functions myself. Good job in finding the issue.