Observation No. 23 · The backup that shares a fault
The help desk was in the building that lost power.
Overnight on August 13, storms hit Phoenix hard enough to damage the cooling at the PhoenixNAP data center. The room started to heat up. PhoenixNAP’s own notice described “elevated white space temperatures.” Servers do not tolerate heat, so the operators did the only safe thing and shut them down on purpose. More than 5,000 servers went offline.
Those servers ran Namecheap. Hosting, EasyWP, DNS, and private email all went dark, some services unreachable, others crawling. The blast radius reached past Namecheap’s own customers, too. Sites hosted somewhere else still broke if their nameservers pointed at Namecheap, because the thing that tells the internet where a website lives was in the hot room with everything else.
Then the part worth writing down. A Namecheap customer watching their site go down would naturally do one thing: open a support ticket to ask what was happening. They couldn’t. The support help desk lives in the same building. The channel you use to get help when the system fails had failed with the system, for the same reason, at the same moment. CEO Hillan Klein posted updates from outside it, two of four chillers back, a third expected within about three hours, service returning by mid-afternoon Eastern. The help desk itself had nothing to answer with.
One environmental failure, a storm and a cooling unit, took out both the service and the only built-in way to ask about the service. They shared a physical dependency nobody had thought of as shared, because on a normal day the walls of a data center are just walls.
Your business has a version of this, and it hides the same way. The backup that lives on a drive in the same office as the computer it backs up, so the flood or the fire or the break-in takes both. The emergency contact sheet saved as a file on the server that’s down. The spare key to the shop kept inside the shop. The generator that needs the same fuel delivery the outage just interrupted. The one employee who knows the alarm code and the safe combination and the vendor passwords, who is also the person you’d call to find out any of them.
The test is a single question asked twice. First, what backs this up? Then, does the backup depend on anything the original depends on? A generator behind the same locked gate as the power meter is not a backup during a lockout. A second phone line billed through the same provider is not a fallback when the provider goes down. A cross-trained second person who rides in the same truck to the same job is not redundancy on the morning that truck won’t start.
The failure is quiet because the shared dependency does its job every day. The drive in the office backs up fine, right up to the day the office is the problem. The help desk answers every ticket, right up to the storm that takes the building. You never see the coupling until the one event that hits the thing both sides sit on, and by then you are Namecheap’s customer, watching the site go down with no way to ask why.
Fixing it is usually cheap and slightly annoying, which is why it doesn’t get done. Move one copy of the backup off-site, or into a service that isn’t yours. Keep the emergency numbers on paper, somewhere that doesn’t need power to read. Put the spare key with a neighbor. Write the codes down and give the sheet to a second person who works a different shift. Each of these separates a backup from the thing it backs up, and each one does nothing at all until the day it is the only reason you’re still running.
The move for this week is small: pick your most important fallback, and trace what it sits on. If it sits on the same thing as the system it protects, you don’t have a fallback. You have a second copy of the same risk.
Does your fallback, your backup, or your support channel quietly share a dependency with the thing it is supposed to back up?
