Warm standby beat multi-cloud, and it wasn't close
A regulated customer asked for multi-cloud resilience. What they actually needed was a second region and a rehearsed failover.
When a financial customer asks “what happens if your cloud provider goes down?”, the reflexive answer in a sales call is multi-cloud. It sounds like the most resilient option. In practice, it often isn’t.
What the question really means
Under DORA, the regulator doesn’t ask you to run on two clouds. It asks you to know your recovery objectives, to show you can meet them, and to prove you’ve tested it. Those are three very different things from “be on AWS and GCP at the same time”.
What we did instead
- A warm standby in a second region with the same provider.
- Data replicated continuously, infrastructure defined in the same Terraform modules.
- A failover runbook that we actually rehearse, with timings written down.
Resilience you haven’t rehearsed is a hypothesis, not a capability.
Why not multi-cloud
Every abstraction you add to stay portable is another thing that can fail at 3 a.m., and another skill set your team has to keep sharp. For most regulated SaaS, the provider-wide outage is far less likely than a bad deploy, a misconfigured IAM policy or a regional incident. Spend the complexity budget where the risk is.