
Just like Dyn is a very popular DNS service provider, AWS is very popular for many website to host their backend. The problem is, AWS is one of the largest CSPs and a lot of popular sites are hosted there, including Imgur, Medium, Slack, Adobe and Salesforce.com. It’s a no wonder for many internet trawlers, it feels like half the internet’s gone down, when one of the AWS S3 availability zone failed .
Amazon explained that they were experiencing some lag in their billing systems, which led to a debugging session accidentally removing more servers offline than expected.
In a note to customers, AWS wrote that “an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. The servers that were inadvertently removed supported two other S3 subsystems. One of these subsystems, the index subsystem, manages the metadata and location information of all S3 objects in the region.”
Or, in simpler terms, someone made a typo, and took down more servers than only the billing system that needed debugging.
Amazon has since put in some additional safety measures to prevent similar events from happening again. “We have modified this tool to remove capacity more slowly and added safeguards to prevent capacity from being removed when it will take any subsystem below its minimum required capacity level. This will prevent an incorrect input from triggering a similar event in the future. We are also auditing our other operational tools to ensure we have similar safety checks.”



