Skip to content

End-to-End DR Walkthrough ​

This is a single, sequential runbook for exercising the entire disaster-recovery cycle by hand, so you can confirm your two-site setup works before you depend on it. Each step says what to do, what you should see, and how to verify it.

Use a disposable test workload for this walkthrough — never a production VM.

Before you start

Complete the one-time setup first:

  • Both sites paired at the Cockpit and DRM layers — see Pairing Two Sites.
  • Disaster Recovery licensed on both Cockpits; both DRM instances licensed or in trial — see Getting Started.
  • A small throwaway test VM running at the primary site.

Terminology below: Site A = primary (where the test VM runs), Site B = recovery.


Step 1 — Protect the test workload ​

Do: In DRM's Workloads view, select Protect on the test VM. Choose the recovery site (Site B), a target datastore, an RPO (e.g. 15 minutes), and compression. Confirm.

Expect: The workload moves to Protected, and a baseline (full) replication starts.

Verify: When the baseline finishes, the workload shows at least one recovery point and a Last Successful Sync time. See Protecting Workloads.

Don't continue until the baseline is complete

A recovery point only appears once the replica has physically landed at Site B. If replication reports a failure instead, resolve it before moving on — see Monitoring & Troubleshooting.

Step 2 — Confirm replication is healthy ​

Do: Trigger a Sync Now on the workload, then write a small marker file inside the guest and Sync Now again.

Expect: Each sync adds a new recovery point; the incremental transfers far less data than the full disk.

Verify: The recovery-point count increases and the newest point's age is within your RPO. See Managing Replication.

Step 3 — Group it into a recovery plan ​

Do: In DRM's Recovery Plans view, create a plan, add the test workload as a step (action Failover, a priority group, an optional boot delay), and save.

Expect: The plan is created at Site A and replicates to Site B as a read-only copy within a reconcile cycle.

Verify: On Site B's DRM, the plan appears and is read-only (editing it there is refused). See Recovery Plans.

Step 4 — Rehearse with a test failover (non-disruptive) ​

Do: Run a Test of the plan (or Test Failover on the workload).

Expect: A copy of the workload boots at Site B in an isolated bubble network. The production VM at Site A keeps running, untouched.

Verify: The test copy is running at Site B with all its disks attached; Site A's VM is unaffected.

Then clean up the rehearsal: choose Cleanup test (or Stop Test Failover). The isolated copy and its temporary network are removed. See Failover & Recovery.

This is the safe checkpoint

If the test failover boots and cleans up correctly, your replica is genuinely recoverable. Only proceed to a real failover once this passes.

Step 5 — Simulate a disaster and recover ​

Now rehearse a real recovery. There are two ways, depending on what you're testing:

  • Planned migration (Site A still healthy): choose Planned Migration on the workload. Site A gracefully powers the VM down, a final delta syncs, and it boots at Site B — no data loss, no split-brain.
  • Emergency failover (Site A lost): make Site A unavailable (for a lab, power off the primary), then from Site B's DRM run the replicated plan.

Expect: The workload boots at Site B from the latest recovery point, keeping its identity (UUID preserved).

Verify: The recovered VM is running at Site B with all disks attached and your Step-2 marker file present. See Failover & Recovery.

Emergency failover moves production

A real failover relocates the workload. On the lab test VM this is fine; never rehearse it on production. DRM refuses an emergency failover if it detects the source still running (split-brain guard) unless you explicitly override.

Step 6 — Reprotect and fail back ​

Do: With the workload now running at Site B, choose Reprotect — Site B begins replicating it back toward Site A. Once Site A is healthy and a recovery point exists there, perform a planned failover back to Site A (a failback).

Expect: Replication direction reverses (B → A), then the workload returns to Site A.

Verify: The VM is running at Site A again with data intact; optionally Reprotect once more to restore the original A → B direction. See Reprotect & Failback.

Step 7 — Clean up ​

Do: Remove the test artifacts so nothing lingers:

  1. Disable replication on the test workload (DRM Workloads → Disable, or Cockpit → Disable Replication).
  2. Delete the test recovery plan.
  3. Delete the throwaway test VM and any leftover recovery copies (<name>-dr) at the recovery site.

Verify: The workload is no longer protected, the plan is gone, DRM shows the protected count back to its starting value, and no -dr recovery domains remain.


What you've validated ​

Completing this walkthrough confirms the full cycle works end to end:

CapabilityProven by
Replication to the recovery siteSteps 1–2
Grouped, replicated recovery plansStep 3
Non-disruptive rehearsalStep 4
Real recovery (planned or emergency)Step 5
Return to the original siteStep 6
Clean teardownStep 7

Make it a habit

Re-run at least Steps 1–4 on a regular schedule and after any significant change (a version upgrade, a network change, a new workload). A test failover is non-disruptive and is the only reliable proof your DR is ready.