The StackingClassifier from sklearn appears to benefit when I run it with a random input. This might assist in memory allocation or initializing the base classifiers. While this approach addresses the issue of the initial slow request, if there’s a 5-minute break, the first request afterward is again slow. Even though I executed this using a custom runner, the slowdown persists after several minutes. Will also experiment with max_latency_in_ms, besides do you have any other suggestions I should dive into?