Business continuity and backup plan for SOC 2
Your cloud provider's resilience is not your continuity plan. What an auditor wants is your recovery objectives, your backup schedule, and proof you have restored something.
For a SOC 2 audit, a business continuity and backup plan has to state four things: recovery time and recovery point objectives per system tier, what is backed up and how often, how long backups are retained and where they live, and how you verify a restore actually works. The fourth is the one auditors ask for evidence of, and a configured backup with no tested restore is the most common gap on this control in a first audit.
If you run entirely on a public cloud, most of the classic continuity content does not apply to you. Delete the sections on alternate work sites, generator fuel and data centre failover. What replaces them is a plain statement of what you would do if a region became unavailable, if your database were corrupted, or if a key vendor went down for a day.
1 per year Restore tests most plans commit to, and the minimum an auditor expects
35 days Backup retention that comfortably clears a monthly restore cycle
What recovery objectives should we commit to?
Tier your systems before you name any numbers. A single set of objectives applied to everything either makes your internal wiki as critical as your production database or lets your database inherit an undemanding target. Three tiers is enough.
| Tier | What is in it | RTO | RPO | Backup frequency | Retention |
|---|---|---|---|---|---|
| Tier 1 | Production application and its primary database | 4 hours | 1 hour | Continuous plus daily snapshot | 35 days |
| Tier 2 | Supporting services, object storage, search indexes, queues | 24 hours | 24 hours | Daily | 35 days |
| Tier 3 | Internal tools, analytics warehouse, documentation | 72 hours | 7 days | Weekly | 90 days |
RPO is a promise about data loss, not about speed
A one hour recovery point objective says you accept losing up to an hour of data. If your database backs up nightly with no point-in-time recovery configured, your real RPO is 24 hours no matter what the document says. Check the configuration first, then write the number. Auditors compare the two.
The restore test, which is the control that matters
Backups are a configuration. A restore is a control, because it is the only one that proves the configuration works. Write into the plan that a restore is tested at least annually, name what gets restored, and keep a record with the date, who ran it, what was restored, where it was restored to, how long it took and whether the data was verified as correct.
The test does not have to be a full disaster simulation. Restoring the production database to a scratch environment, running a row count and a spot check of recent records, and writing down the elapsed time satisfies both the control and the point of the exercise. What it must not be is a screenshot of a backup job that succeeded, which is evidence of a backup, not of a restore.
Twice a year is better than once, since the first attempt in a new environment usually finds something: a missing permission, an unexpected duration, an encryption key nobody had documented. Doing it once at month eleven means finding that out the week before fieldwork. The timeline page covers where the point-in-time controls like this one sit relative to the observation window.
Comparing firms for this? Tell us what you need and it goes to the ones in the directory that do this work. No charge, and no phone number required.
What continuity means when you own no hardware
Write your plan around four scenarios that can happen to a cloud company, and say what you would do in each. Cloud provider region failure: whether you fail over, and if you do not, say so and say what the customer impact would be. Data corruption or accidental deletion: point-in-time restore, which is the scenario most likely to occur. Loss of a critical vendor for a day, such as an authentication provider or a payment processor. And loss of key people, which for a small company is a real continuity risk and is worth a paragraph about documentation and shared access.
Be careful with the shared responsibility line. Your provider is responsible for the resilience of its infrastructure. You are responsible for the architecture you built on top, for your backups, and for your configuration. A plan that says "our cloud provider maintains 99.99% availability" and stops there is not a plan, and its accuracy is not something you can evidence anyway.
Does your draft contain these clauses?
0 of 8 done ·
What gets a continuity plan turned into a finding?
"Backups are tested regularly." Cut it. Regularly is not a cadence, and the auditor will ask when the last test was. If the answer is that nobody has run one, the control fails on the spot. Write annually or semi-annually and then run it in the first quarter of the window.
A stated RPO your database cannot meet. The auditor pulls the backup configuration. Nightly snapshots against a one hour RPO is a written commitment contradicted by a screenshot, and it is embarrassing in a way a missing document is not.
Alternate site and pandemic sections inherited from the template. A plan describing a hot site and a battery room for a company with a dozen laptops tells an auditor the document was not written by anyone who works there. It does not fail a control by itself, but it changes how carefully everything else you hand over gets read.
Backups that live inside the blast radius. A backup stored in the same cloud account with the same credentials as production does not survive the scenario most likely to need it. Write down where backups live and how access to them differs, and check that a compromised production credential cannot delete them.
The policy templates page covers editing an inherited pack down to what you run, and the policy checklist will tell you whether you need continuity as a standalone document, which mostly depends on whether you have taken on the availability criterion alongside security.
Get a continuity plan that is tested, not filed
The plan is the easy half. The tabletop and the record of it are what the auditor samples. Tell us your scope and compare firms that do both.
Get matchedCommon questions
Do we need a business continuity plan if we only take the security criterion?
Yes, in practice. Backup and recovery sit under the common criteria on system operations regardless of whether you added availability, and every auditor request list asks for the plan and the restore evidence. Taking the availability criterion raises the bar on the objectives and on monitoring against them, but it does not create the requirement.
Is a disaster recovery plan the same thing?
Not quite. Disaster recovery is the technical restoration of systems. Business continuity covers how the company keeps operating, including people and communication. Small companies routinely combine them into one document and no auditor objects, as long as the technical objectives are in there.
What does a restore test record need to contain?
The date, the person who ran it, what was restored, the target environment, the elapsed time, how the restored data was verified and the result. A paragraph and a screenshot of the verification query is enough, and it beats a polished report with no evidence of the actual check.
Do we have to fail over to a second region?
No. Multi-region failover is an architectural choice with real cost, and most companies at this size do not have it. Say so plainly and write recovery objectives that reflect a single-region deployment. Claiming failover you have not built is the failure mode, not the absence of it.
Where does data residency come into this?
In your backup location clause. If you have told Canadian customers their data stays in Canada, backups and any restore environment have to honour that too, and a restore into a US scratch environment quietly breaks the promise. Write the permitted regions into the plan.
How long should we keep backups?
Long enough to survive a problem you did not notice immediately, which is usually 30 to 35 days for production data, and no longer than your retention policy allows for personal information. Those two constraints pull in opposite directions, so the continuity plan and the data retention policy have to name the same numbers.