Just found a nice resolution to a problem that was...
# general
m
Just found a nice resolution to a problem that was bugging us for the last few days... (sorry this will be slightly long) TLDR; Pay attention to what is where.... πŸ™‚ We are testing a new Database product to replace our current self-managed MySQL installation. we had several requirements, and so we started testing AWS Aurora (provisioned, not serverless). Requirements: Latency, Latency & performance Our main issue with MySQL was the lag (latency) of the replication instances, it was just too long to be able to route most read operations to a replica. Also, we wanted higher write throughput, and more than everything - avoid having to manage this... So we started testing, all the standard things you'd expect - some generated load testing, sample interesting queries, failover behavior and speed, resizing (this is fun in contrast to the self-managed). But, we wanted more. more confidence that when we migrate everything will be OK, and we won't have to rollback due to some unexpected, you find it only in production, problems... So we figured we needed another test method. I came up with the genius idea (genius because it was immediately assigned to me - beware of having good ideas) - that we should find a high traffic (in regards to DB work) service and duplicate it incoming workload to a shadow copy of the service. The original instances will stay connected to the existing MySQL server and work as usual, supplying the clients with the data. While the shadow will be connected to a brand new Aurora cluster but do nothing with the query results. This will allow testing a real production load, checking Aurora "in-action" and we can see how the failover looks on a "production" env. So, go with it, they said... Work. that's why we pay you πŸ™‚ So, had to figure out how to do this duplication. 2 tracks - 1. Make sure the service will "like" this - make sure no unknown side effects (writing to places) will hurt production, so code changes to make sure that a differently named app (with the same code) will not overwrite metrics, logs, redis caches, etc. 2. Figure a way to duplicate the traffic reliably without adding a significant overhead to an already overloaded production environment ... My first thought, a project I've been wanting to work with for years, and never got the chance Goreplay - this is a network packet level app that can take packets and just shoot them someplace else, but with lots of filtering & control (we wanted filtering by host and port - it gives that and much more) - however, a network card level app like this is a major project to install in a prod env, with all the mess going on, the last thing we wanted is an app that might destroy every packet in the data-center... Second thought - NGINX - it has a mirroring module (I may have mentioned before?), this is entirely application level, and NGINX is already part of our stack - so it sounded great, configured, tested, tested again it worked great locally, deployed to multiple testing envs, no problemo, traffic is duplicated, looks great! Then we deployed to production - and within 5 minutes, our NGINX (which along with duplicating handled traffic between our backend services crashed. This means all our API clients started getting 502s... Not good, reverted (15-20 minutes long incident). BTW - this was "vetted" by our Dev/Sec Ops including reviewing the code and configuration (my main conclusion - never volunteer to configure production tools you are not an expert at). We wrote an incident report, I took a few days to relax... Just as an aside, as this is getting much longer than I thought - this was due to the fact that the module mirrors by using "subrequests" and the nginx actually waits for both requests to finish (the original and the dupe) - since we had only one shadow server on 7 main - the latency of the shadow slowed down processing, and created a huge request queue killing the nginx. But we only found this out after the fact obviously - if we had limited the rate, and put up enough shadow copies (which was the plan, but we wanted to start "small") Third option (BTW all the options were in my original design doc, the order is following all those annoying meetings in which this was discussed)... - Add a little piece of code to the service to intercept and re-send all the incoming traffic - we originally didn't want this, as by design it would hurt the performance of the service, but I had enough (no one was offering, and we still wanted to do this test), from now on I would only be willing to test code I had control of. So I wrote an Interceptor (Node.js + Nest), that had rate controls (you could send anywhere between 0 - 100% of traffic intercepted), and more useful configuration to control the load. We then deployed this very gradually (as we should have with NGINX), starting at 1 % traffic and growing as we saw it was stable. This survived, and so we entered the data gathering phase, started monitoring the metrics and logs to see what we missed, and we did! We found important cluster configurations we overlooked, Failover definitions that needed help, our isolation level was wrong (we only found this because we saw deadlocks) and more. The test is still running, and a few days ago we noticed the client metrics showed higher latency on Aurora than MySQL - which puzzled us as we certainly didn't expect this, so we started looking, added some more dashboards, a ticket to AWS (which asked for Select samples), and such. We suspected the multi-AZ disk writes, we suspected network issues, so we started thinking of how to find this. I started writing a little client that we could deploy as lambda and test connectivity and latency from multiple points and varying queries. but we thought, lets check the network angle first (VPC Peering and such), we started pinging around, no luck - network was great, less than 0.5ms on average for a ping round trip. We looked more and were about to quit (and start that mysql client thing) and someone noticed that our main instance was in a different AZ than everything else (same region), due to one of our failover tests... So we failed over again. and what do you know? The latency dropped as soon as the instances started connecting to the new writer... That was a fun day! You can see for yourself πŸ™‚ (the spike up is the failover, the light blue line is the shadow latency with Aurora, the dark is the MySQL) Important note about credit and ownership and teams: While I'm obviously the hero πŸ˜‰ there is a whole team with me, and it is a great collaborative fun, I wasn't the one who noticed the latency, but a colleague who decided to look at the client metrics, I didn't find the issue with the AZ, but one of our SREs who wouldn't give up even when the network report looked OK, I wasn't the one to offer the test - that was our CTO, I just offered the technical means to execute it. It takes a village as they say. Always share your thoughts, always support a good idea, and never say no - go for it! hope this wasn't too much, always check your config... and where you put your stuff - matters! thanks for reading, would love your questions and comments - I can't provide internal info of course, but whatever is ok, I'll be glad to share
d
This sounds like a great piece of collaborative work - what tools or frameworks did you use to coordinate things at each stage? I'm thinking like GDocs / JIRA / Basecamp / etc. ?
m
Designs/ docs - we use mostly Confluence, we have a draw.net account for charting which is an important tool as well, Tasks / Project planning - JIRA - I hate it, and try to avoid this part, luckily my managers and others handle this, I'm the worst Direct Coordination, meetings & (in)direct communication - Slack/Huddle/Email/Zoom/Meet/In person - really anything goes as long as you get the message across well
πŸ‘€ 1
g
@Moshe Eshel you need a blog πŸ™‚
βž• 1
Did you at least post this on LinkedIn?
m
No, you think it should be?
g
It’s a great write up. Definitely worth sharing outside this group
m
High praise coming from you 🀩 To me it reads like a rant, can't see me reading it...
My default mode would be a review by a colleague, comments, cleanups and then, maybe I publish
g
Do it :)
πŸ‘ 1
I'm debating with myself whether to publish on Medium or LinkedIn I have a very old unmaintainable blog that I'm not sure is worth resurrection However both above platforms feel abused enough already. I think I'll go with LinkedIn, between the two I've always gotten better content there, at least that's my feel
πŸ‘ 1
Also, the story isn't over yet, so I think I'll wait for the conclusion. The solution turned out to be a false hope, the improvement we were seeing was the clients hitting the reader (DNS wasn't refreshed) and getting that read mode error on writes which skewed the πŸ“‰ (think heavy insert finishing way faster than it should have, no insert and no error)
g
ouch
Those ups-downs on the way to a solution are so relatable
m
That's an engineer's life, isn't it? If our machines would function without issues there'd be nothing new to accomplish
g
Any updates? I'm holding my breath on how you end up solving it πŸ™‚ Also, do you mind if I copy paste this (including followup, and with obvious attribution) to Hacking SaaS? It is so good, and I'm afraid you'll not publish it πŸ™‚
❀️ 1
m
Updates are we still have the issues, our main db guy was on vacation. We had a meeting with AWS, they didn't have a lot to contribute. I'm writing a little tool that will allow us various tests across DBs and accounts with lots of metrics. Hope to deploy it next week and get some usable stats. In addition we are hunting for bad queries. The main takeaway from the meeting is that we would be better off (higher confidence) if we ran the same test with aurora-mysql 5.74 Because the 8 engine is known to be significantly different (in addition to migration from MySQL to Aurora we are also jumping from v5.74 to v8.0.13 or so, mostly compatible, but many differences. So we will probably also do the exact same thing with a 5.74 cluster - if that shows parity we'll know to blame Aurora or V8 - or not πŸ₯΄ Regarding publication, I'll post it as is, unless you have feedback - to linkedin by tomorrow, then you can feel free to republish, quote or link to (or all three, whatever you like...) I'll just add that the story is far from over, and I will be sure to post updates and conclusions with lessons learned. That will be long
g
Great plan! πŸ™‚
m
first draft,a little forward and some typo fixes.
i want to add more graphics
make it proper
g
❀️ isn't it passed midnight at your $local_time?
m
yes, 00:51 to be exact, I took some serious stuff, and I'm sort of turning on and off - I'm assuming I have minutes before I'm out for the night
g
Take care, my friend.
I'm still swooning over this line: "Always share your thoughts, always support a good idea, and never say no - go for it!"
m
BTW, started editing, and also trying to find out within my company if there is a policy I need to follow, approve the posting/review etc. What a mess