Most companies still test their blue team the same way they ran fire drills in the 90s. There is an exercise scheduled for the second Thursday of the quarter. The SOC knows it is coming. The IR team has pre-allocated their afternoon. Detection engineers have tuned their content for the playbook in advance. The exercise runs, everyone gets a green check, and the report says the controls are working.
That report describes a tabletop. It does not describe how your defenses behave on a regular Tuesday at 3 a.m. when an actual adversary is not on the calendar.
This is the gap continuous adversary emulation is built to close.
What unannounced means in practice
The pipeline runs MITRE ATT&CK aligned playbooks against your live environment on a deliberately randomized cadence. Some weeks there is one run. Some weeks there are three. The timing inside the week is also random. A run might land at 10:14 on a Monday morning when the SOC is fully staffed, or at 03:45 on a Saturday when the on-call analyst is the only person paged.
This is not adversarial against the blue team. It is the only way to find out what the blue team's actual response time looks like when the alert is not expected, when the analyst is mid-coffee, when the senior on-call is asleep. Real adversaries do not send a calendar invite. The emulation should not either.
The playbook itself is chosen from a library of TTP profiles, also randomized within the scope the customer approves. APT29 one month, Volt Typhoon the next, FIN7 the month after. The phases run end to end, with the same techniques the underlying group is observed to use in the wild, on the kill chain the group is known for.
A run, from the inside
A representative run is a six-phase walk down a real group's playbook. The APT29 (G0016) profile, for example, looks like this when it lands in the pipeline:
- phase 1 · T1566.001 spearphishing attachment → landed
- phase 2 · T1059.001 PowerShell scriptblock execution → detected
- phase 3 · T1053.005 scheduled task persistence → landed
- phase 4 · T1003.001 LSASS memory dump → blocked
- phase 5 · T1021.002 lateral movement via SMB admin shares → landed
- phase 6 · T1041 exfiltration over C2 channel → landed
Three outcomes per phase, with no editorializing. Landed means the technique executed and nothing stopped it. Detected means something fired in the SIEM or EDR but the chain continued. Blocked means a control on the endpoint or the network actively prevented the step from completing.
A run is not graded on landed-versus-blocked alone. A blocked LSASS dump is a strong signal, and a detected PowerShell scriptblock is meaningful, but the real question is whether the blue team noticed the sequence in time to interrupt it. The pipeline measures that separately, by correlating the technique timestamps with the incident timestamps on the defender side. If phase 4 fired at 12:30 and the SOC acknowledged at 12:34, that is one kind of outcome. If they acknowledged at 16:00, that is a different one.
Why this is purple team, not pure red
Continuous adversary emulation is purple team work, not a covert red team engagement. The customer's blue team is not the target of a long, undisclosed campaign trying to evade them for six months. The runs are open at the executive level, the scope is agreed, and the artifacts are shared with the defenders after each cycle.
What is hidden from the SOC is the schedule. The fact that emulation is part of the program is known. The fact that the next run is happening right now is not.
This is the configuration that produces useful signal. If the SOC knows nothing, you are running a red team. If the SOC knows when, you are running a tabletop. If the SOC knows that it could be any time, you are running adversary emulation the way it was meant to work.
What you find that announced exercises do not
A few patterns show up almost universally in the first months of continuous runs, regardless of how mature the program looks on paper:
- Detections that pass in a controlled exercise miss in production because the data source they depend on has been quietly broken for weeks. The agent forwarding PowerShell logs from one fleet of workstations stopped sending three deploys ago. Nobody noticed because nobody was looking. The next emulation run is the thing that notices.
- The shift handover loses tickets. A technique fires at 17:55, the day analyst marks it as low priority, the night analyst inherits the queue without context, the chain completes by 21:00. This is invisible in a Thursday-afternoon exercise. It is the most common real outcome.
- Detection content that was tuned against last year's APT29 profile silently degraded when an EDR update changed the field name on a relevant event. The rule still runs. It just no longer matches anything.
None of these are exotic. All of them are the kind of thing that only an unannounced run, against the live stack, on a random day, will expose.
MITRE ATT&CK as evidence, not as aspiration
Every run produces a structured artifact that maps each phase to its ATT&CK technique, the outcome, the timestamp, and the corresponding defender activity. Over months, this accumulates into a heatmap of the matrix that is grounded in your environment, not in a vendor's marketing slide.
The point of the heatmap is not to claim coverage. The point is to claim measured coverage. T1003.001 is green because you have blocked it on three separate runs across three different playbooks over the last quarter, not because somebody ticked the box in a spreadsheet. T1021.002 is yellow because it lands every time and your SOC catches it inconsistently. T1059.001 is detected reliably and your team's mean time to acknowledge has dropped from twelve minutes to four over six months.
That is the report a board wants to read. Not a percentage of ATT&CK techniques theoretically covered, but a record of what actually fired, what actually blocked, and what the blue team actually did when the alert was real.
What the blue team gets out of it
The framing for defenders is the part that most programs get wrong. The runs are not designed to make the SOC look bad. They are designed to give the SOC the one thing it normally cannot get: hundreds of realistic incidents per year, with ground truth attached, against the controls they are actually responsible for.
Every junior analyst who joins the team will see APT29 land in their queue inside their first six months. They will see the technique, they will see the alert, they will work the ticket, and after it closes they will see exactly what was on the other side. There is no faster way to build operator instinct, and there is no other source of training data that is also production-real.
The senior engineers get a different kind of value. Every run is a regression test on their detection content. When a rule that used to catch a technique stops catching it, the next emulation cycle says so. When a new control is rolled out, the next cycle measures it. The detection backlog stops being a guess and starts being a list with evidence behind every item on it.
If you want to see what an unannounced APT29 run produces against your environment, we can scope a single playbook against a segmented subset of your estate and walk you through the artifact afterwards. The first run is usually the one that tells the team the most.