Just found a nice resolution to a problem that was bugging us for the last few days... (sorry this will be slightly long)
TLDR;
Pay attention to what is where.... π
We are testing a new Database product to replace our current self-managed MySQL installation. we had several requirements, and so we started testing AWS Aurora (provisioned, not serverless).
Requirements: Latency, Latency & performance
Our main issue with MySQL was the lag (latency) of the replication instances, it was just too long to be able to route most read operations to a replica. Also, we wanted higher write throughput, and more than everything - avoid having to manage this...
So we started testing, all the standard things you'd expect - some generated load testing, sample interesting queries, failover behavior and speed, resizing (this is fun in contrast to the self-managed).
But, we wanted more. more confidence that when we migrate everything will be OK, and we won't have to rollback due to some unexpected, you find it only in production, problems... So we figured we needed another test method.
I came up with the genius idea (genius because it was immediately assigned to me - beware of having good ideas) - that we should find a high traffic (in regards to DB work) service and duplicate it incoming workload to a shadow copy of the service. The original instances will stay connected to the existing MySQL server and work as usual, supplying the clients with the data. While the shadow will be connected to a brand new Aurora cluster but do nothing with the query results. This will allow testing a real production load, checking Aurora "in-action" and we can see how the failover looks on a "production" env. So, go with it, they said... Work. that's why we pay you π
So, had to figure out how to do this duplication. 2 tracks - 1. Make sure the service will "like" this - make sure no unknown side effects (writing to places) will hurt production, so code changes to make sure that a differently named app (with the same code) will not overwrite metrics, logs, redis caches, etc. 2. Figure a way to duplicate the traffic reliably without adding a significant overhead to an already overloaded production environment ...
My first thought, a project I've been wanting to work with for years, and never got the chance
Goreplay - this is a network packet level app that can take packets and just shoot them someplace else, but with lots of filtering & control (we wanted filtering by host and port - it gives that and much more) - however, a network card level app like this is a major project to install in a prod env, with all the mess going on, the last thing we wanted is an app that might destroy every packet in the data-center...
Second thought - NGINX - it has a
mirroring module (I may have mentioned before?), this is entirely application level, and NGINX is already part of our stack - so it sounded great, configured, tested, tested again it worked great locally, deployed to multiple testing envs, no problemo, traffic is duplicated, looks great! Then we deployed to production - and within 5 minutes, our NGINX (which along with duplicating handled traffic between our backend services crashed. This means all our API clients started getting 502s... Not good, reverted (15-20 minutes long incident). BTW - this was "vetted" by our Dev/Sec Ops including reviewing the code and configuration (my main conclusion - never volunteer to configure production tools you are not an expert at). We wrote an incident report, I took a few days to relax... Just as an aside, as this is getting much longer than I thought - this was due to the fact that the module mirrors by using "subrequests" and the nginx actually waits for both requests to finish (the original and the dupe) - since we had only one shadow server on 7 main - the latency of the shadow slowed down processing, and created a huge request queue killing the nginx. But we only found this out after the fact obviously - if we had limited the rate, and put up enough shadow copies (which was the plan, but we wanted to start "small")
Third option (BTW all the options were in my original design doc, the order is following all those annoying meetings in which this was discussed)... - Add a little piece of code to the service to intercept and re-send all the incoming traffic - we originally didn't want this, as by design it would hurt the performance of the service, but I had enough (no one was offering, and we still wanted to do this test), from now on I would only be willing to test code I had control of. So I wrote an Interceptor (
Node.js + Nest), that had rate controls (you could send anywhere between 0 - 100% of traffic intercepted), and more useful configuration to control the load. We then deployed this very gradually (as we should have with NGINX), starting at 1 % traffic and growing as we saw it was stable.
This survived, and so we entered the data gathering phase, started monitoring the metrics and logs to see what we missed, and we did!
We found important cluster configurations we overlooked, Failover definitions that needed help, our isolation level was wrong (we only found this because we saw deadlocks) and more.
The test is still running, and a few days ago we noticed the client metrics showed higher latency on Aurora than MySQL - which puzzled us as we certainly didn't expect this, so we started looking, added some more dashboards, a ticket to AWS (which asked for Select samples), and such. We suspected the multi-AZ disk writes, we suspected network issues, so we started thinking of how to find this. I started writing a little client that we could deploy as lambda and test connectivity and latency from multiple points and varying queries. but we thought, lets check the network angle first (VPC Peering and such), we started pinging around, no luck - network was great, less than 0.5ms on average for a ping round trip. We looked more and were about to quit (and start that mysql client thing) and someone noticed that our main instance was in a different AZ than everything else (same region), due to one of our failover tests... So we failed over again. and what do you know? The latency dropped as soon as the instances started connecting to the new writer... That was a fun day!
You can see for yourself π (the spike up is the failover, the light blue line is the shadow latency with Aurora, the dark is the MySQL)
Important note about credit and ownership and teams: While I'm obviously the hero π there is a whole team with me, and it is a great collaborative fun, I wasn't the one who noticed the latency, but a colleague who decided to look at the client metrics, I didn't find the issue with the AZ, but one of our SREs who wouldn't give up even when the network report looked OK, I wasn't the one to offer the test - that was our CTO, I just offered the technical means to execute it. It takes a village as they say. Always share your thoughts, always support a good idea, and never say no - go for it!
hope this wasn't too much, always check your config... and where you put your stuff - matters!
thanks for reading, would love your questions and comments - I can't provide internal info of course, but whatever is ok, I'll be glad to share