This message was deleted.
# ask-for-help
s
This message was deleted.
j
Hey @Andrej Zachar what does warm up means in your context?
according to your post above, the initial slow request is probably due to our adaptive batching. if you have a strict latency requirement, you can lower down the
max_latency_in_ms
parameter
a
The StackingClassifier from sklearn appears to benefit when I run it with a random input. This might assist in memory allocation or initializing the base classifiers. While this approach addresses the issue of the initial slow request, if there’s a 5-minute break, the first request afterward is again slow. Even though I executed this using a custom runner, the slowdown persists after several minutes. Will also experiment with max_latency_in_ms, besides do you have any other suggestions I should dive into?