Shipping during a SOC 2 observation window
You do not have to slow down. You have to make the record of each change a by-product of shipping it, because every merge in the window is in the population.
Nothing about a SOC 2 observation window requires you to deploy less often. Teams shipping thirty times a week pass Type 2 audits routinely, and teams shipping once a month fail them. What changes is that every production change between day one and the last day of the window enters a population the auditor samples from, typically 25 to 45 items for a three to twelve month window, and each sampled change has to show the same four facts. If those four facts are a by-product of your pipeline, volume is irrelevant. If they are something a person assembles afterwards, volume is fatal.
4 Facts each sampled change must show
25 to 45 Changes typically sampled from a full window
0 Deploys you need to cancel for the window
What does an auditor need from each change?
- It was reviewed by somebody other than the author
- Evidenced by the approval record on the merge, with a different identity from the committer. A convention that engineers review each other is not this. Branch protection enforcing it is.
- It passed whatever gate you told the auditor about
- If your description says tests gate deployment, the sampled run has to show tests passing. If it says nothing gates deployment, that is fine and testable. The exception is the gap between the two.
- It was authorized, in the sense that somebody wanted it
- A ticket, an issue, or a linked description. This is the weakest of the four in practice and the one auditors are most flexible about, but a merge titled "fix" with no linked anything is the one they circle. The change ticket evidence page covers what the record has to carry.
- It reached production through the path you described
- Deployed by the pipeline, not by somebody with a shell. Every manual push to production during the window is an item you will be explaining.
How many of our changes will actually be looked at?
Sampling is proportionate but strongly sublinear, which is the fact that should calm anyone worried about deploy frequency. Doubling your deploy rate does not double the sample. The bands below are the ranges Canadian CPA firms work to for change management populations, and the same figures are in the table under the chart.
| Merged changes in the window | Items sampled | Share of population |
|---|---|---|
| Under 25 | All of them | 100% |
| 25 to 99 | 10 to 16 | Around 20% |
| 100 to 499 | 20 to 27 | Around 8% |
| 500 to 1999 | 25 to 37 | Around 3% |
| 2000 or more | 30 to 45 | Under 2% |
The first row is the one that surprises people. A team that deliberately slowed to twenty deploys across a six month window did not reduce its audit burden, it guaranteed that every single change gets inspected. High volume is protective. Sloppy high volume is not.
Comparing firms for this? Tell us what you need and it goes to the ones in the directory that do this work. No charge, and no phone number required.
What creates an exception, and what does not?
| Situation | Exception? | What to do instead |
|---|---|---|
| Founder merges own pull request at 2am to fix an outage | Yes, if there is no emergency path | Define an emergency change process with after-the-fact approval inside one business day, and use it |
| Automated dependency bumps merging without human review | Usually not, if described | Name bot-authored dependency updates in the system description as a separate change class with its own control |
| Feature flags flipped in production without a deploy | Yes, if flags change behaviour and nothing records it | Log flag changes with actor and timestamp, and say in the description that this is the control |
| Terraform applied from an engineer's laptop | Yes | Move infrastructure changes into the same pipeline before the window opens, not during it |
| Database migration run by hand during a release | Yes, if it is not in the change record | Ticket it like any change and attach the executed statement |
| Revert of a bad deploy | No | Nothing. Reverts are changes, they follow the same path, and auditors read them as the control working |
| Contractor with commit access and no confidentiality agreement on file | Yes, in the people section rather than change | Fix the agreement before the window opens |
| A single deploy on the last day of the window with no review | Yes, and it is fully in scope | The window ends when it ends. Nothing gets waved through in the last week |
What do we set up before the window opens?
Everything below is a configuration change, not a process change, and each one is the difference between a record that already exists and a record somebody has to assemble from memory in month four.
0 of 0 done ·
Pipeline retention is the one that ends programs. Log and run-history retention cannot be applied backwards, so a default that expires at ninety days quietly destroys the first half of a twelve month window. Check it before day one, alongside the rest of the evidence routine.
The case for slowing down anyway
There is one, and it is not about the audit. If your change history is undisciplined, unreviewed merges, direct pushes, infrastructure by hand, then opening a window on top of it converts every week of shipping into another week of findings you will be writing management responses to. In that situation the right move is to delay the window start by four to six weeks and fix the pipeline, not to keep shipping into a population you already know is dirty. Where a customer date makes that delay look impossible, SOC 2 by a date that is not possible covers what to tell them instead. Delaying a window start is cheap. Explaining eleven exceptions in a report a customer will read is not.
The test is simple: export last month's merges and check the four facts on ten of them at random. If eight or more pass, open the window and keep shipping. If fewer, you have found your remediation work, and the timeline page shows what the delay does to your report date.
Do we have to freeze deploys during a SOC 2 audit?
No. There is no freeze requirement anywhere in SOC 2, and a freeze during fieldwork creates a suspicious gap in a population the auditor is already looking at. Keep shipping normally through fieldwork.
Does deploying more often make the audit harder?
No, and it usually makes it easier. Sample sizes grow far more slowly than populations do, and a team deploying daily almost always has the automation that produces the evidence, whereas a team deploying monthly often does things by hand.
What happens if we find an unreviewed merge mid-window?
Record it, note why it happened and what you changed, and carry on. One documented instance with a response is an exception the auditor can describe as isolated. The version that hurts is an instance the auditor finds and you had not noticed, because it makes the monitoring look absent as well.
Can we exclude a repository from scope?
Only if it genuinely does not affect the system in your description, and you say so in the description. Excluding a repository that deploys to production is a scope misstatement rather than a scoping choice. Where the boundary sits is the subject of the scope page.
Do dependency bot pull requests count as changes?
Yes, they are in the population. Most auditors accept them as a distinct change class with automated review, provided your system description says so before the window rather than after the sample comes back.
Have somebody check the pipeline before you start
A pipeline review before the window opens is a few days of work and it is the cheapest point to find the unreviewed path.
Get matched