Anyone else is getting the sense that AWS ALBs are...
# general
g
Anyone else is getting the sense that AWS ALBs are a bit less stable than NLBs? Like, every once in a while the ALB simply stops responding (no traffic, no metrics) for 1-2 minutes before coming back like nothing happens. I never saw that with the NLB.
c
Are you using grpc traffic?
I had a horrible experience with ALB's and GRPC, so we switched to NLB.
g
we use http2, but not grpc
c
That could be the explanation. HTTP2 uses long-lived connections. Generally there's a
keepAlivePing
(well that is the GRPC name of it but it's an HTTP2 concept) that sends a packet just to check the health of the connection.
What happened to us was two separate things: 1. The ALB would respond to that keepalive ping and it wouldn't get transferred to the backend server, so the server got unhappy and thought the connection closed
2. The ALB didn't count the keepalive ping as "activity" and it would close connections after 10 minutes of inactivity
Eventually our clients would open the connection again.
g
Interesting 🧐 the combination of the two sounds like two engineers didn’t talk to each other or the PM. Each decision kinda sorta makes sense, but not both together…
c
To me, neither of them make sense (especially not the first one) because if the client feels the need to tell the server that "hey I'm still here!", then the ALB should tell the backend server that the client is still there. The ALB provides some nice abstractions but I think the implementation is poor enough that it gets in the way.
g
Yeah, this is exactly what I’m starting to suspect…
We use it for the built in api routing, which is really nice (and cost effective).
c
As a disclaimer, this was 18 months ago, so it could have changed by now.
👍 1
(on K8s I assume?) Yes, the routing directly to the pod is great, and I think there are some features to route to the
Service
endpoint that is in the same AZ as well to reduce cross-AZ costs. And there's one less network hop since the load balancer itself is the ingress controller. Probably uses some fancy
NodePort
magic under the hood.
We don't use it because: 1. It actually doesn't work well for LH, and 2. It's too AWS-specific for us so we would have to re-implement ingress at Google. We're using Gateway API now and we have to make no change to our actual CRD's to get it to work on Google. Pretty wild.
g
Yeah, the control over cross AZ traffic is great too
c
Here's some "documentation" https://stackoverflow.com/questions/66818645/http2-ping-frames-over-aws-alb-grpc-keepalive-ping Gotta go back to writing docs now 🙂 have a good day!
😂 1