The Backup in the Same Building

A fire in one of four buildings on OVHcloud's Strasbourg site took the other three down with it, because the power to the whole campus was cut while it was fought. This is an inventory of the failure domains a backup can quietly share with the thing it protects, and what a restore really costs once you price egress and rehearsal. Same-site copies still earn their place, and they cannot cover the failure people buy them for.

Four Buildings, One Switch

Overnight on the ninth and tenth of March a fire started in SBG2, one of four buildings on OVHcloud's Strasbourg site. The building was a total loss. SBG1 next to it was damaged, and then, as a precaution while the fire was being fought, the electricity was cut across the whole site, which took SBG3 and SBG4 down as well. Two of those three buildings were not on fire and had no other problem. They were simply on the same power feed as one that was.

The company's founder told customers to activate their business recovery plan, which is the correct thing to say and lands differently depending on whether the customer has one. Lichess went dark. Documentation and marketing sites for a spread of well-known projects went dark. Facepunch reported that its European Rust servers had lost their data and that it could not be restored, which is the sentence in the whole event that I would pin to a wall.

Nobody involved was careless. Data centers have fire suppression, compartmentation, alarms and drills, and they have them because this is a known risk that has been engineered against for decades. Buildings still burn. The interesting question is never whether a facility can fail; it is what else stops working when it does, and that is a question about the topology of your copies rather than about anyone's fire doors.

Which is where the uncomfortable part is. A great many of the customers offline that morning had backups. They had bought a backup option, seen it in the invoice, ticked the line in an audit questionnaire, and believed the question was closed. The copy and the original were in the same failure domain, and a backup inside the blast radius of the thing it protects is not a backup. It is a second copy.

What Somewhere Else Has to Mean

The building is where people stop. The obvious one, and the one everybody names. A copy in the same room burns with the room, floods with the room and loses power with the room. This is also the only one on the list that is intuitive to a non-technical audience, which is why it is the one that gets solved first and the one people believe they have finished after solving.

The site is several buildings. Strasbourg is the demonstration. Three buildings that did not burn went dark because the site electricity was cut while the fire was fought, and that decision was correct. Campuses share power feeds, network entry points, cooling plant, access roads and the emergency response that shuts all of it down at once. Separate buildings inside one perimeter are one failure domain wearing four addresses.

Then there is the provider, and the account inside it. A copy in another region of the same provider survives a fire and does not survive a billing dispute, a compromised administrative credential, an accidental account closure or a mistaken deletion propagating through the control plane. The physical separation is real; the administrative separation is zero. It is the domain that gets missed most often because the copies really are far apart.

The credential and the control plane form a domain of their own. If restoring requires logging into a console that is itself down, or fetching a key from a vault that lives in the affected region, the copy exists and cannot be reached. This is the domain that fails quietly, because it is invisible while everything works and it is only exercised in the exact circumstances where everything does not.

Two media used to mean two mechanisms. Three copies, two media, one off site comes out of Peter Krogh's work on photo archives in the mid-2000s, and it survives because it is a rule about independence you can hold in your head. In 2005 two media meant a disk and a tape, whose failure mechanisms were unrelated. Today it usually means object storage at one provider and object storage at the same provider, which is the same words around a very different property.

Last is the format, and the software that reads it. A copy that only one product can open shares a domain with that product's licensing, its version compatibility and its continued existence. It is the slowest of these to bite and the hardest to unwind. The test is whether you could read the archive with tools you could obtain in an afternoon from someone other than the vendor who wrote it.

A True Sentence Is Not a Property

"Backups are taken nightly and retained for thirty days" is a true statement about most systems I have looked at. It is also compatible with every failure above. It says nothing about where the copies are, who can delete them, how long a restore takes, whether anyone has done one, or whether the thing restored would actually run. It is a statement about a job that completes, and a job completing is the cheapest possible evidence.

That sentence passes audits, and I do not think auditors are the problem. The question on the form is answerable, the answer is verifiable, and both parties leave satisfied. The trouble is that the artifact produced by the exercise is a tick, and a tick has no gradient: it looks identical whether the copy is in another country under separate credentials or in the rack next to the source.

What you are buying is not storage. It is a tested restore, with an explicit blast radius and a number attached to how long it takes. Everything else is a file sitting somewhere.

And the reason people buy the other thing is that storage is a line item and recovery is a project. One of them can be approved in a procurement cycle by a person who never has to think about it again. The other needs someone to run a drill, find out that the drill takes eleven hours, and then go and argue for the budget to make it four. Nothing in how organizations allocate money favors the second.

Costing the Restore

Egress is priced to discourage you. Storing bytes with a cloud provider is cheap and getting them all back out is not, because outbound transfer is billed per gigabyte at rates that have barely moved in years. A full restore is the one operation that reads everything at once, so the exit price is a real number that appears exactly once, in the worst week you will have. Work it out now, while it is arithmetic rather than an invoice.

Retrieval tiers mean cheap storage buys you a delay. Archive tiers cost a fraction of standard object storage and impose a retrieval time measured in hours, plus a retrieval fee per gigabyte. That is an excellent trade for records you must keep and will never read, and a terrible one for the copy your business restarts from. The mistake is not choosing the cheap tier. It is putting both kinds of data in it because they were both called "backup".

Throughput sets your recovery time whatever the plan says. Recovery time is bytes divided by the narrowest link between the copy and the running system, and that link is usually not the one in the diagram. It is a single-threaded restore process, or a database that has to replay after the files land, or one network interface on one machine. An estimate that has never been measured is not an estimate.

Rehearsal hours are the part that actually works. A restore drill costs a few engineer-days a year and is the only line here that reliably prevents the outcome. It is also the only one that never survives a busy quarter, because it produces no artifact anyone outside the team can see. If a budget conversation forces a choice between more retention and one rehearsed restore, the rehearsal wins every time and it is not close.

A restore has to be of something real, to somewhere new. Not a checksum, not a listing, not a restore into the environment that already works. Bring a service up from the copy on infrastructure that did not exist that morning, have somebody use it, and record how long it took. Everything short of that tests the backup software, which is the part least likely to be broken. The duration is the output, and it moves as the data grows.

You need a break-glass path to your own copy. Getting to the backup during an incident needs a credential, a network path and a person, and all three have to work when the primary environment does not. Write down who can reach the copy from a laptop on a home connection with the corporate identity provider unavailable. If the answer is nobody, the copy's location does not matter.

Why the Local Copy Earns Its Place

Same-site backup is fast, cheap, and the right answer to the failures that actually happen. The overwhelming majority of restores are for a deleted file, a bad migration, a corrupted table or a bad deploy, and for all of those a local copy is restored in minutes over a link that costs nothing. A remote copy with a four-hour retrieval time would be strictly worse for every one of those events, and every one of them is a hundred times more likely than a fire.

The mistake is not buying it. It is buying it and believing the question is answered, because the two copies serve different failures and the cheap one cannot cover the expensive one. A local copy is an undo button and a remote copy is an insurance policy, and an organization that has confused the two has usually bought the undo button, because it is the one whose value is visible every month.

The question is not how many copies you have; it is how many independent events could take all of them at once, and the count is usually one and nobody has written it down. Copies are easy and cheap and they multiply. Independence is expensive, and it is the only property in the exercise that does anything.

I should concede that the honest version of that is expensive and I have made it sound like a weekend. A second provider means a second contract, a second set of credentials, a second thing to monitor and a transfer bill that recurs whether or not anyone ever reads the copy. For a small organization that is a real fraction of the infrastructure budget spent on an event that may never occur, and I have watched that argument lose, reasonably.

What I would not concede is the sequencing. The cheap half of this is knowing where the copies are and having restored one, and that costs a few days of somebody's year. If the drill then says the exposure is acceptable, that is a decision. What Strasbourg produced was a lot of organizations discovering the answer on Wednesday morning, which is the same information arriving at the worst possible price.

I want to be careful with the lesson, because the fire is doing rhetorical work here that it has not earned on its own. One building burned; the base rate for that is very low, and an organization that ran everything in one facility for a decade and got away with it was not being stupid, it was accepting a small risk that mostly does not land. The failure I am actually describing is not the fire. It is that almost nobody could say, that morning, what their copies would have survived.

The line I keep returning to is that the power was cut to all four buildings, and three of them were fine. Whatever the diagram said about separation, the site had one switch, and somebody had to throw it in order to fight a fire in a building that was not yours. Every backup arrangement has a switch like that somewhere, and the only question worth answering before Wednesday is where it is and who else is standing next to it.