An untested disaster recovery plan is a hope, not a plan. Until you’ve actually restored your systems and timed it, you don’t know if your backups work, how long recovery takes, or whether the data you assume is protected even exists. Disaster recovery testing is the only thing that turns a document into a capability.
Most businesses we meet discover this at the worst possible moment — mid-incident, watching a restore fail or a backup turn out to be empty. The plan looked fine on paper. Nobody had ever proven it. This post covers why backups quietly fail, the four kinds of DR testing, how to actually run a test, and how it all ties back to your recovery targets.
Why your backup is probably lying to you
A backup job that reports “completed successfully” every night feels reassuring. The green tick tells you the job ran. It does not tell you the data is complete, readable, or recoverable. Those are different questions, and the gap between them is where businesses get destroyed.
Here are the failure modes we see most often, and they’re rarely dramatic — they’re quiet:
- Incomplete coverage. The backup protects the file server everyone remembers, but not the SQL database behind the line-of-business app, the configuration on a critical appliance, or the new cloud VM someone spun up in March. Coverage drifts as the environment changes, and nobody re-checks it.
- Corrupt or unreadable backups. The job runs, the file lands, but the backup itself is corrupt — a bad block, a half-finished snapshot, an encryption key nobody can find. You only learn this when you try to restore.
- Microsoft 365 assumed safe but not backed up. This is the big one, and it deserves its own section below.
- No recent restore test. The most common failure of all. The backups exist. Nobody has actually pulled data back out of them in months — or ever. A backup you’ve never restored from is an untested assumption, not a safety net.
A manufacturing business in Dandenong we work with came to us after a ransomware scare at a competitor prompted a nervous board to ask a simple question: “If this happened to us, could we get back up?” Their previous provider had been running nightly backups for years. When we ran a test restore, two of the four critical systems wouldn’t come back — one backup set was corrupt, the other had silently stopped including the database six months earlier. Every nightly report had said “success”.
The Microsoft 365 and SaaS backup gap
This catches out more Melbourne SMEs than anything else, so it’s worth being blunt. Microsoft 365 does not back up your data in the way most people assume. Microsoft replicates your data across its own infrastructure for availability — so the service stays up if a data centre has a problem. That is not the same as a backup that protects you from your own mistakes.
Under Microsoft’s shared responsibility model, the data in Exchange Online, SharePoint, OneDrive and Teams is your responsibility to protect. If an employee deletes a mailbox folder, a departing staff member wipes a SharePoint site, or ransomware encrypts files synced through OneDrive, Microsoft’s retention windows are limited and unforgiving. A deleted user’s mailbox is gone after 30 days by default. Versioning and recycle bins help with small accidents, but they are not designed to recover from a deliberate or large-scale data loss.
The same logic applies to other SaaS platforms — Xero, your CRM, your practice management system. Many businesses run their entire operation on cloud apps and have never asked who is responsible for backing the data up. The honest answer is usually: nobody. A proper backup strategy treats Microsoft 365 and key SaaS data as first-class assets to be backed up independently, with their own restore tests. We cover the full backup picture in our guide to backup and disaster recovery for Melbourne businesses.
The four types of DR testing
“Testing your DR” isn’t a single activity. There’s a ladder of tests, from cheap and frequent to involved and occasional. A mature business uses all four for different purposes.
| Test type | What it proves | Effort | How often |
|---|---|---|---|
| Restore verification | The backup is readable and a real file or record comes back intact | Low | Monthly |
| Tabletop exercise | Your people know the plan, roles and decisions under pressure | Low to medium | Twice a year |
| Partial failover | One system or workload actually recovers to a working state | Medium | Annually |
| Full failover | The whole environment recovers and the business can operate | High | Annually, where the risk warrants it |
Restore verification
The foundation. You pick data from a backup and actually restore it, confirming it opens and is intact. Done monthly, this catches corrupt backups and coverage gaps long before a real incident. It’s low effort and there is no excuse for skipping it — yet it’s the test most businesses never run.
Tabletop exercise
A tabletop is a structured walkthrough, no live systems touched. You gather the people who’d respond to a disaster, present a realistic scenario — “ransomware has encrypted the file server and the backup NAS at 9am on a Monday” — and talk through exactly who does what, in what order, using what. Tabletops surface the human gaps: nobody knows where the recovery runbook lives, the only person with the backup credentials is on leave, no one’s sure who decides to declare a disaster. Cheap to run, brutally revealing.
Partial and full failover
Failover testing is the real thing — you actually recover systems. A partial failover brings one workload back, often into an isolated environment so it doesn’t disrupt production. A full failover recovers the entire environment and confirms the business could genuinely operate from the recovered state. Full failover is the most demanding test and the most honest one. It’s where you discover that recovery takes eleven hours, not the three you’d promised the board.
How to actually run a test
Testing fails when it’s vague. “We should test our backups sometime” never happens. Here’s the approach that works:
- Schedule it. Put recurring dates in the calendar — monthly restore verification, a tabletop each half, an annual failover. If it isn’t booked, it won’t happen.
- Pick a real recovery scenario. Don’t test the easy case. Choose something that would actually hurt: the primary file server lost, the finance database corrupted, the email tenant compromised. Test what you’re afraid of, not what’s convenient.
- Time it against your targets. Start a clock. How long until the data is back and usable? Measure it against your Recovery Time Objective (RTO) and check the recovered data is recent enough to satisfy your Recovery Point Objective (RPO). A restore that works but takes three days when your RTO is four hours is a failed test.
- Document the gaps. Write down everything that went wrong or slowly — missing credentials, an out-of-date runbook, a system nobody knew wasn’t covered. The output of a test is a list of fixes, not a pass mark.
- Fix and re-test. Close the gaps, then test again to confirm. A test that finds problems you never fix is theatre.
The discipline here is the same one that separates a real plan from a hopeful one. If you can’t put a stopwatch on your recovery and read the number out loud to your directors, you don’t have a tested plan.
Tying results back to RTO and RPO
This is the point of the whole exercise. Your RTO is how long the business can tolerate being down before the damage is serious. Your RPO is how much data you can afford to lose, measured in time. These aren’t IT numbers — they’re business decisions about survival, and we walk through setting them in our piece on RTO versus RPO and how long your business can survive without IT.
A DR test exists to prove your recovery meets those targets. You set an RTO of four hours; the test reveals actual recovery takes nine. That’s not a failure of the test — it’s the test doing exactly its job, exposing a dangerous gap before a real incident does. Now you have a choice: invest in faster recovery, or honestly revise the target and tell the business what to expect. Either way you’re working from reality instead of a comforting assumption.
Untested, RTO and RPO are just numbers in a spreadsheet. Tested, they become commitments you can stand behind. That’s the difference between business continuity that exists on paper and business continuity that actually holds when the building floods or the ransomware note appears.
How managed backup with monthly restore verification works
The reason most businesses don’t test is honest: it’s tedious, easy to defer, and the day job always wins. That’s precisely why it belongs with a provider who does it as routine rather than something you’ll get to “next quarter”.
Under our managed backup and disaster recovery service, restore verification is scheduled, not optional. Each month we pull real data back from backups — including Microsoft 365 — and confirm it’s intact, not just that the job reported success. Coverage is reviewed as your environment changes, so the new server or cloud app doesn’t quietly fall outside protection. When something fails, we find out during a test on a quiet Tuesday, not during an incident at 9am on a Monday with the whole business watching.
TechAssist is a Melbourne-based MSP founded in 2014, with 13 Australian-employed engineers and a 24/7 NOC in Tecoma. We run backup as an operated service — monitored, tested and tied to your RTO and RPO — alongside managed IT, because a backup nobody verifies is the most expensive false comfort in IT.
Frequently asked questions
How often should we test our disaster recovery?
Restore verification monthly, a tabletop exercise twice a year, and a failover test annually where the risk warrants it. The cheap, frequent tests catch the most failures, so don’t skip restore verification just because it feels routine — that routine is exactly what protects you.
Isn’t my Microsoft 365 data already backed up by Microsoft?
No. Microsoft keeps your data available across its infrastructure, but protecting it from deletion, ransomware and human error is your responsibility under their shared responsibility model. Default retention is limited and won’t save you from a large or deliberate data loss. Microsoft 365 needs its own independent backup with its own restore testing.
What’s the difference between a tabletop and a real failover test?
A tabletop is a discussion — you walk through a scenario to check your people, roles and decisions, without touching live systems. A failover test actually recovers systems and proves the technology works. You need both: the tabletop finds the human gaps, the failover finds the technical ones.
What does a failed DR test mean?
It means the test worked. Finding a corrupt backup, a missing system or a recovery that blows past your RTO during a controlled test is the entire point — far better there than during a real disaster. A test that surfaces problems has just saved you. The failure is never testing at all.
Stop hoping, start proving
A plan you’ve never tested is a guess about the worst day of your business year. The fix isn’t more documentation — it’s putting a stopwatch on a real restore and reading the number honestly. Verify your backups, close the Microsoft 365 gap, run the tabletop, and prove your recovery actually meets the targets you’ve set.
If you’re not certain your backups would come back — or you’ve never actually tried — get in touch. We’ll run a real restore test, time it against your RTO and RPO, and tell you plainly where you stand. Better to find out now than to find out mid-incident.