We’re collecting best bug stories and would love t...
# general
y
We’re collecting best bug stories and would love to hear yours.
It worked yesterday
Today it has other plans
Welcome to Monday
What’s the most memorable bug you’ve ever found (or accidentally introduced)? Reply in the thread with: • Your favourite bug you’ve ever found (or accidentally created!) • Why it was memorable • What you learned No proprietary code please. Keep examples anonymous and avoid sharing confidential information. Entries close on Friday, 31 July. We’ll select our three favourite bug stories and announce the winners the following week. The top three entries will receive a prize and be featured in our community showcase.
👀 1
m
I’ll start. I was awoken at 3am (and then later at 6am due to a fire alarm in the building) due to an ongoing incident with our production system struggling with what looked like a DoS attack, targeting our login page. Investigation showed the attack coming from IPs all over the world, so we weren’t able to apply the usual network blocks. After investigation, I noticed that there was an errant script on the page that resulted in a loop hitting an API endpoint over and over and over. A minor emergency JS fix, and we all got back to bed. (except for me, of course, thanks to some burnt toast on level 5)
🙌🏻 1
y
(Disclaimer: I am one of the judges) I remember moving from a world of on-prem systems, to that of one based in the clouds. Effectively someone else’s computer. We thought we were pretty nifty with lambdas, and serverless helping connecting multiple services together, until something went wrong and we were trying to sift through the logs. There were a-lot of them, and it was like looking for a needle in a haystack, with welding goggles on. We very quickly learned about and implemented distributed tracing. That wasn’t the end of our troubles. We had an MI system, that would get populated by way of a csv uploaded to an S3 bucket. This process started via our Postgres database setup with logical replication. Wal2json provided json changeset messages which were sent on a kinesis stream, enriched and transformed, and uploaded to the S3 bucket. We got a call from management, that the payout system has been relatively quiet for the last few days. It transpired that we had run a large database update, pushing our changeset message over 1mb, which was the limit for kinesis messages. Our code was unaware of this error condition, so blindly would try and re-attempt. In the end, we had to fix forward by rescuing from the error and then splitting our message into chunks. The morale of this story, is bugs are everywhere, and you can only fix what you can see, so try to capture as much as you can from your running systems, so that you can lift off those welding goggles, blow away the haystack and be left with a treasure map to the root cause.
😬 1