What a real disaster recovery test looks like
A backup job reporting success proves the job ran. It proves nothing about whether you can recover, how long it takes, or whether the data is usable.
Ask an organisation when it last tested its disaster recovery and the answer is usually a version of “the backups are monitored and they are all green”.
That is a statement about a job completing. Recovery is a different event, and the gap between them is where continuity plans fail — not dramatically, but slowly, on the day it matters, while somebody discovers that the restore takes eleven hours and the recovery objective said four.
What “green” actually proves
A successful backup job proves that data was read from a source and written to a target without an error being raised.
It does not prove:
- that the data can be read back
- that the restore completes inside the time you committed to
- that the restored system starts, authenticates and talks to the systems around it
- that the backup covered everything, including the things added since the job was configured
- that the person on call knows how to run it
Every one of those has failed in an environment we have been called into, and in every case the backup reporting was green throughout.
The three levels of test
A restore test. Take one file, one mailbox, one virtual machine from a real backup and restore it. Time it. This is the minimum, it takes an hour, and it should happen quarterly. It catches the most common failure: backup data that cannot actually be read.
A recovery test. Restore a full system to a usable state in an isolated environment, then confirm it starts, authenticates against identity, connects to its database and serves a request. Annually, per critical system. This catches dependency failures — the system that restores perfectly and cannot work because it needs a certificate, a licence server or a service account that was never in scope.
A scenario exercise. Give a group of people a scenario, no warning, and the actual documentation, and have them work through it. Half a day, annually. This catches the failures that are not technical at all: nobody knows who declares an incident, the runbook is in the system that is down, the contact list is out of date, and the only person who knows the environment is on leave.
Organisations that do the first one are ahead of most. Organisations that do all three are rare and are the ones whose plans work.
The questions a test should answer
Write these down before the test, then answer them from what actually happened:
- How long did it take? Wall clock, from decision to service restored. Compare that number to your stated recovery time objective. If they differ, one of them is wrong and it is usually the objective.
- How much did we lose? The gap between the last usable backup and the incident. Compare to the recovery point objective.
- What was not in the backup? There is always something. Configuration held outside the system, a certificate, a scheduled task, a firewall rule, an integration key.
- Who could have done this? If the answer is one person, that is a finding worth more than the test.
- What did the documentation get wrong? It will have got something wrong. Fix it while the test is fresh.
Microsoft 365 is the common blind spot
The single most frequent gap we find is that Microsoft 365 is not backed up at all, because retention policies are being mistaken for backup. Microsoft protects the platform; your data in it is your responsibility, and retention does not help against a deletion discovered six months later, a compromised account, or a departed employee’s mailbox that was removed with their licence.
It is an inexpensive gap to close and an expensive one to discover. We have written about what Microsoft 365 does and does not protect in more detail.
What to schedule
- Quarterly: a timed restore of one real item per critical system.
- Annually: a full recovery test of each critical system, in isolation, to a working state.
- Annually: a scenario exercise with the people who would actually be involved.
- After every significant change: a check that the new thing is in scope of the backup. This is where coverage quietly erodes.
And keep the evidence. For regulated organisations the test record is the artefact an assessor asks for, and reconstructing it afterwards is neither convincing nor possible.
More on backup and business continuity.