Two years ago a design review approved our document storage on one condition: every object is replicated to a second bucket in another region. In a tabletop exercise in July somebody asked which region, and I opened the console to show them. The replica bucket was in eu-west-1. So was the original.
Replication is set up by our storage module, which creates the primary bucket, the replica bucket, the role and the rule. The replica bucket uses an aliased provider, and the caller is meant to pass one pointed at the recovery region. The first repository to use the module did exactly that. The module's README, though, had an example that mapped the replica provider to the default one, so that it could be applied in a single sandbox account, and the three repositories written after it copied the example. Terraform was happy. An aliased provider pointed at the same region is a perfectly valid configuration.
Everything we monitored was green. Objects arrived in the replica within seconds, the replication metrics showed nothing pending, and the object counts matched to the last file. Each of those checks confirmed that a copy existed. None of them looked at where it was, and where it was happened to be the entire purpose of it. Three of our four document stores, about nine terabytes, would have gone down with the region they were supposed to outlive.
Fixing the data took four days of batch replication into new buckets in the recovery region, and a week of checking before we removed the old replicas.
Fixing the pattern took less. The module now has a precondition that reads the region of both providers and fails the plan if they are equal, with a message that names the variable to change. The README example uses two regions, even in the sandbox. A nightly job lists every replication rule in every account, asks the API for the actual region of each destination bucket, and alerts if it matches the source. And the exercise in July has become a standing question in design reviews: for anything described as a copy, show me where it lives, from the provider's API rather than from our code.
A disaster recovery copy has one property that matters more than all the others, and it is the one property that replication metrics never report.
– Sergey Shinder







