2Labs Tech

What Should an MSP Disaster Recovery Plan Include?

A backup is a copy of data. An MSP disaster recovery plan explains how the business will use its people, systems, vendors and backups to operate again after something goes wrong.

That difference matters. A provider may be able to show that last night’s backup completed successfully without knowing whether the accounting system can be restored, how long recovery will take or which system should come back first.

A useful disaster recovery plan answers those questions before an outage, ransomware incident, fire, equipment failure or extended internet disruption.

Disaster recovery is one responsibility inside a broader managed IT plan. See what an MSP does and when a small business may need one for the larger context.

Start with the business, not the backup product

The first step is identifying which operations cannot wait.

Ask department leaders:

  • Which systems do you use every day?
  • What stops if that system is unavailable?
  • How long can the business reasonably operate without it?
  • Is there a manual workaround?
  • Which employees and vendors are needed to restore service?
  • Where is the information needed to contact them?

For a small business, the priority list may include email, shared files, accounting, point of sale, scheduling, phones, internet access and a specialized line-of-business application. A church may prioritize livestream, presentation and donation systems differently. A farm or rural operation may have seasonal systems that become critical during planting or harvest.

The MSP should understand those differences rather than restore equipment in whatever order is technically convenient.

What should be backed up?

The plan should identify important data wherever it lives:

  • Servers and shared folders
  • Workstations that store local business files
  • Microsoft 365 or Google Workspace data
  • Accounting and line-of-business applications
  • Configuration for firewalls, switches and other infrastructure
  • Website or cloud application data when the business is responsible for it
  • Documentation, license information and recovery credentials

Do not assume a cloud application includes the retention and restore options your business expects. Confirm what the software vendor protects, what the MSP protects and what remains the client’s responsibility.

Recovery time and recovery point, in plain English

Disaster recovery discussions often use two terms.

Recovery time objective (RTO) is the target for how long a system can remain unavailable before it causes unacceptable damage.

Recovery point objective (RPO) is the amount of recent data the business can tolerate losing. If a system is backed up once each night, a failure late in the day could mean losing that day’s changes.

These are business decisions, not merely technical settings. Faster recovery and more frequent protection usually require additional cost, so the owner should decide where that investment matters most.

A simple table can help:

System Maximum tolerable downtime Maximum tolerable data loss Temporary workaround
Email [business decision] [business decision] Personal phones or alternate contact method
Accounting [business decision] [business decision] Paper notes until controlled re-entry
Shared files [business decision] [business decision] Local copies only if approved
Point of sale [business decision] [business decision] Documented offline process

Complete the table with actual leaders before choosing technology.

Where should backup copies be stored?

A local copy can provide fast recovery from a failed computer or deleted file. It may not survive theft, fire, flooding or ransomware that reaches the same network.

A resilient approach usually separates copies so one incident cannot destroy all of them. Ask:

  • Is there an independent copy outside the primary location?
  • Can ordinary user or administrator credentials delete every copy?
  • Are backups encrypted?
  • How long are versions retained?
  • Are failed or incomplete jobs reported?
  • What happens if the backup appliance itself fails?

The right design depends on the systems and recovery targets. The important point is avoiding a single point of failure.

Restore testing is the proof

A backup dashboard can show successful jobs while hiding a problem that only appears during recovery. Testing verifies that data can be read, the process is documented and the people involved understand their roles.

Testing may range from restoring a sample file to conducting a full application or server recovery exercise. The plan should state what is tested, how often and how results are documented.

Ask your provider:

  • When was the last restore test?
  • What exactly was restored?
  • How long did it take?
  • Were any errors found?
  • What changed after the test?

“We have never needed to restore it” is not reassuring. It means the process has not been proven under pressure.

The plan needs people and communication

Technology is only part of recovery. The document should identify:

  • Who declares an incident
  • Who can authorize emergency purchases or system shutdowns
  • Who communicates with employees and customers
  • Who contacts the internet, software, insurance and equipment vendors
  • How leadership will communicate if normal email or phones are unavailable
  • Where printed or offline copies of key contacts and procedures are stored
  • Who documents decisions and actions during the incident

Cyber incidents may also involve legal counsel, cyber insurance and law enforcement. The MSP should know when to preserve evidence and avoid actions that could complicate an insurance claim or investigation.

Internet and power failures belong in the plan

Not every disaster is a cyberattack. Rural organizations may face carrier outages, damaged lines, storms and power interruptions.

Review:

  • Whether the firewall can use a secondary internet connection
  • Which systems need battery backup
  • How long critical network equipment can remain powered
  • Whether employees can work from another location
  • Which cloud services remain usable through a temporary connection
  • Who contacts the carrier and tracks the outage

A backup internet connection does not need to support every activity. It may only need to keep payment, phone or essential cloud services working until the primary connection returns.

What should happen during a ransomware incident?

Employees should report the warning immediately and avoid experimenting with the affected system. The provider’s first job is containment: isolate affected devices or accounts, determine the scope and protect clean recovery copies.

The plan should not assume that paying a demand will restore data or solve the incident. Recovery decisions may involve insurance, legal counsel and law enforcement. The organization needs a trusted communication chain rather than a technician and owner making major decisions alone in the first few minutes.

For common rural-business security gaps, see 5 Cybersecurity Mistakes Small Rural Businesses Make.

For a detailed ransomware response and recovery checklist, review CISA’s #StopRansomware Guide.

Questions to ask about an MSP disaster recovery plan

  1. Which systems and cloud services are protected?
  2. How often does each backup run?
  3. How much recent data could we lose?
  4. How long should recovery of each critical system take?
  5. Where are independent copies stored?
  6. Who responds to backup failures?
  7. When was the last restore test, and what were the results?
  8. Which system will you restore first, and why?
  9. What equipment, licenses or vendors could delay recovery?
  10. How will we communicate if email, phones or internet are unavailable?
  11. What responsibilities belong to us?
  12. How often will the plan be reviewed and tested?

If the provider cannot answer those questions, the business has backup software, not a recovery plan.

Keep the plan small enough to use

A plan does not need to begin as a hundred-page manual. Start with the critical systems, contacts, recovery order, backup locations, decision authority and communication method. Test one realistic scenario and improve the document from what you learn.

Related MSP guides

2Labs Tech provides monitoring and patching and cybersecurity support for rural Kansas organizations. A Practical Tech Checkup can identify critical systems and backup gaps before you decide what recovery targets are reasonable.