End-to-End DR Walkthrough
This is a single, sequential runbook for exercising the entire disaster-recovery cycle by hand, so you can confirm your two-site setup works before you depend on it. Each step says what to do, what you should see, and how to verify it.
Use a disposable test workload for this walkthrough — never a production VM.
Before you start
Complete the one-time setup first:
- Both sites paired at the Cockpit and DRM layers — see Pairing Two Sites.
- Disaster Recovery licensed on both Cockpits; both DRM instances licensed or in trial — see Getting Started.
- A small throwaway test VM running at the primary site.
Terminology below: Site A = primary (where the test VM runs), Site B = recovery.
Step 1 — Protect the test workload
Do: In DRM's Workloads view, select Protect on the test VM. Choose the recovery site (Site B), a target datastore, an RPO (e.g. 15 minutes), and compression. Confirm.
Expect: The workload moves to Protected, and a baseline (full) replication starts.
Verify: When the baseline finishes, the workload shows at least one recovery point and a Last Successful Sync time. See Protecting Workloads.
Don't continue until the baseline is complete
A recovery point only appears once the replica has physically landed at Site B. If replication reports a failure instead, resolve it before moving on — see Monitoring & Troubleshooting.
Step 2 — Confirm replication is healthy
Do: Trigger a Sync Now on the workload, then write a small marker file inside the guest and Sync Now again.
Expect: Each sync adds a new recovery point; the incremental transfers far less data than the full disk.
Verify: The recovery-point count increases and the newest point's age is within your RPO. See Managing Replication.
Step 3 — Group it into a recovery plan
Do: In DRM's Recovery Plans view, create a plan, add the test workload as a step (action Failover, a priority group, an optional boot delay), and save.
Expect: The plan is created at Site A and replicates to Site B as a read-only copy within a reconcile cycle.
Verify: On Site B's DRM, the plan appears and is read-only (editing it there is refused). See Recovery Plans.
Step 4 — Rehearse with a test failover (non-disruptive)
Do: Run a Test of the plan (or Test Failover on the workload).
Expect: A copy of the workload boots at Site B in an isolated bubble network. The production VM at Site A keeps running, untouched.
Verify: The test copy is running at Site B with all its disks attached; Site A's VM is unaffected.
Then clean up the rehearsal: choose Cleanup test (or Stop Test Failover). The isolated copy and its temporary network are removed. See Failover & Recovery.
This is the safe checkpoint
If the test failover boots and cleans up correctly, your replica is genuinely recoverable. Only proceed to a real failover once this passes.
Step 5 — Simulate a disaster and recover
Now rehearse a real recovery. There are two ways, depending on what you're testing:
- Planned migration (Site A still healthy): choose Planned Migration on the workload. Site A gracefully powers the VM down, a final delta syncs, and it boots at Site B — no data loss, no split-brain.
- Emergency failover (Site A lost): make Site A unavailable (for a lab, power off the primary), then from Site B's DRM run the replicated plan.
Expect: The workload boots at Site B from the latest recovery point, keeping its identity (UUID preserved).
Verify: The recovered VM is running at Site B with all disks attached and your Step-2 marker file present. See Failover & Recovery.
Emergency failover moves production
A real failover relocates the workload. On the lab test VM this is fine; never rehearse it on production. DRM refuses an emergency failover if it detects the source still running (split-brain guard) unless you explicitly override.
Step 6 — Reprotect and fail back
Do: With the workload now running at Site B, choose Reprotect — Site B begins replicating it back toward Site A. Once Site A is healthy and a recovery point exists there, perform a planned failover back to Site A (a failback).
Expect: Replication direction reverses (B → A), then the workload returns to Site A.
Verify: The VM is running at Site A again with data intact; optionally Reprotect once more to restore the original A → B direction. See Reprotect & Failback.
Step 7 — Clean up
Do: Remove the test artifacts so nothing lingers:
- Disable replication on the test workload (DRM Workloads → Disable, or Cockpit → Disable Replication).
- Delete the test recovery plan.
- Delete the throwaway test VM and any leftover recovery copies (
<name>-dr) at the recovery site.
Verify: The workload is no longer protected, the plan is gone, DRM shows the protected count back to its starting value, and no -dr recovery domains remain.
What you've validated
Completing this walkthrough confirms the full cycle works end to end:
| Capability | Proven by |
|---|---|
| Replication to the recovery site | Steps 1–2 |
| Grouped, replicated recovery plans | Step 3 |
| Non-disruptive rehearsal | Step 4 |
| Real recovery (planned or emergency) | Step 5 |
| Return to the original site | Step 6 |
| Clean teardown | Step 7 |
Make it a habit
Re-run at least Steps 1–4 on a regular schedule and after any significant change (a version upgrade, a network change, a new workload). A test failover is non-disruptive and is the only reliable proof your DR is ready.