Your backup plan is only as good as the last time you have actually done a restore.
I spent eight years as Director of Engineering for a healthcare company in Boston. Once a year, we held “Disaster Day,” an exercise that started after 9/11. Our office was near the USS Constitution, as well as a Liquid Natural Gas terminal. We started to think of the disaster we were planning for was terrorists detonating a LNG Tanker which would have destroyed our office and most of Boston’s waterfront
On a random day, my boss would walk in and announce: “It’s Disaster Day.” Then he and I would drive to a Panera in New Hampshire. I would be handed a laptop and told to prove that I could bring the site back from our offsite backups.
This was before AWS and GitHub had become the default answers to infrastructure and source control. Our code lived in SourceForge, while the replicated database and snapshots were on a Level 3 colocation machine in Denver. The point was not to review a disaster-recovery document and declare ourselves prepared. The point was to start from a different location, with limited equipment, and actually restore the system.
Then, one year, my boss came in and said: “It’s Disaster Day, and sadly Dan was at work when the building was destroyed. Jorge, come with me.”
Jorge was my lieutenant. And, as it turned out, only I knew the SSH password to the colo machine.
That was a successful test, even though it was an embarrassing one. It found a real single point of failure before a real disaster found it for us. We fixed it.
Backups are important, but “we have backups” is not the same thing as “we can recover.” Can you find them? Are they complete? Are the credentials still valid? Can someone besides the person who set everything up perform the restore? How long does it take? What data will be missing? Does the restored application actually work?
Do the actual restore. Document every step, including the things that went wrong and the assumptions that turned out not to be true.
Today, AI can help create recovery runbooks, automate parts of the process, identify missing dependencies, and turn the lessons from each test into better documentation. That makes monthly testing much more practical for many teams. But AI does not change the central rule: recovery is not real until you have restored the system and verified that it works.
Test your backups. Then test the people, credentials, documentation, and decisions that make those backups useful.

Leave a Reply