upon further investigation, it looks like this does happen on the internal endpoint too - just a few seconds of "connection refused" then resumes working normally, which i suppose explains the 503 from the ingress, but i am still confused as to why it would happen in a deployment with multiple replicas. or even a deployment with a single replica, as i had assumed it would wait for the new pod to start and become ready before killing the old. this is pushing the boundaries of my knowledge of kubernetes, but it seems like the old pods are being killed before traffic stops being routed to them or something?