Few terms in cybersecurity are as misused as wargame. It gets sold as a pentest, confused with Red Team, and is frequently purchased without the client knowing exactly what is being measured. The confusion is not trivial: it changes who is being evaluated, what counts as success, and what the board receives in the final report.
This article separates the concepts and explains why, in a well-designed wargame, the trophy does not belong to the attacker — it belongs to the defender.
The spectrum of offensive exercises
Instead of treating each service as an isolated box, it is more useful to see them on a spectrum — from most focused on "finding flaws" to most focused on "testing the defenders":
| Type | Question it answers | Who consumes the result |
|---|---|---|
| Pentest | What vulnerabilities exist and how far can I get by exploiting them? | Infrastructure and application teams |
| Red Team | Could a real adversary reach a critical objective without being detected? | Security leadership |
| Purple Team | Red and Blue together: what did we detect, what didn't we, and how do we fix it now? | Detection engineering |
| Wargame / Adversary Emulation | Faced with a realistic adversarial scenario, how well does the organization detect, decide, and respond? | The entire security operation and management |
A pentest answers a question about vulnerability. A wargame answers a question about defensive capability. These are different goals, and treating them as equivalent is the root of nearly every frustration with this kind of engagement.
In a pentest, the scoreboard belongs to the attacker
In a pentest, success is defined by the tester's progress: escalating privilege, chaining vulnerabilities, reaching the objective. The deeper they go — and often the more quietly — the more valuable the result. The focus is on proving that the path exists.
This is essential, but it answers a narrow question: "is it possible?". Almost always, the answer is yes. Every sufficiently complex environment has exploitable paths. Knowing that is rarely the bottleneck for a mature operation.
In a wargame, the scoreboard belongs to the defender
In a wargame, the Red Team also operates for real: it escalates privileges, moves laterally, collects and exfiltrates data. The attack chain is genuine. The difference lies in three design decisions that completely reposition the exercise.
1. The primary metric is the reaction, not the advance
The Red Team's advance stops being the objective and becomes the stimulus. What is under evaluation is the response to that stimulus: did the SOC notice? How fast? Did it classify correctly? Did it escalate to the right people? Did it contain? That is why a wargame's deliverables revolve around defensive metrics:
- MTTD (Mean Time to Detect)
- MTTA (Mean Time to Acknowledge)
- MTTC (Mean Time to Contain)
- MTTR (Mean Time to Respond)
- MITRE ATT&CK coverage — which techniques were detected, partially detected, or went unnoticed
2. Assumed breach: the attacker is presumed to be already in
Instead of spending weeks trying to break down the front door (reconnaissance and initial access), many wargames assume the breach: they start with the adversary already inside the environment — an insider, a compromised workstation, a leaked credential. The rationale is statistical and mature: over a long enough horizon, the attacker will get in. The testing budget goes further by measuring what happens after the door is broken — because that is precisely where most organizations are blind.
3. Representative, observable TTPs — not maximum stealth
A classic Red Team aims to be as quiet as possible to prove it can slip through undetected. In a wargame, that would be counterproductive: if the attacker is invisible and nothing is detected, you learn nothing about the maturity of your controls — you only learn that the Red Team is good at hiding.
In a wargame, each technique is executed cleanly and observably enough to give the defensive team the chance to detect it. Excessive stealth masks the very reality you set out to measure.
What stays secret from the Blue Team are the exact timings, the vectors, and the specific techniques — the so-called semi-blind model. The team knows something is coming; it does not know what or when. If it knew the schedule, it would measure theater, not real capability.
Why "the Red Team exfiltrated" is not a failure
This is the point that most unsettles executives — and the most important one to communicate before the exercise. In a wargame, the attacker's technical success is not the final grade: it is an input. The question that matters is not "did they make it?", but:
- How long did it take to detect the activity?
- Was the alert correctly prioritized, or lost in the noise?
- Was there clear escalation and communication between teams?
- Did the existing playbooks work under pressure?
- Where are the blind spots in telemetry and coverage?
An exercise where the Red Team advances and the Blue Team quickly detects, classifies, and contains is an excellent result. An exercise where nothing is detected is the most valuable warning an organization can receive — before a real adversary delivers it.
An established standard in the financial sector
This model is not a vendor invention. Regulators of critical infrastructure have formalized entire frameworks with exactly this DNA:
- TIBER-EU — Threat Intelligence-Based Ethical Red Teaming (European Central Bank)
- CBEST — Bank of England
- iCAST / AASE — Hong Kong and Singapore
All share the same pillars: a threat-intelligence-based scenario, legitimate offensive execution, a White Cell controlling the exercise, and a Blue Team kept semi-blind. Their central goal, in every case, is to validate detection and response capability — not to catalog vulnerabilities.
How to choose the right exercise
The practical rule is simple:
- If you need to know where the flaws are in a system or application: pentest.
- If you need to know whether a determined adversary would reach your most critical objective without being caught: Red Team.
- If you need to evolve your detection engineering collaboratively and immediately: Purple Team.
- If you need to measure and prove, with objective metrics, how well your operation detects, decides, and responds under pressure: wargame.
Security maturity is not proven by the absence of vulnerabilities — that is unattainable. It is proven by the speed and quality of the response when the inevitable happens.
How Antisec runs it
At Antisec, we design and execute exercises across this entire spectrum — from pentest and Red Team to MITRE ATT&CK-based adversary-emulation wargames, with White Cell governance, a semi-blind model, and objective detection-and-response metrics. Our Blue Team practice closes the loop by turning every finding into concrete detection-engineering improvements, and our DevSecOps and vCISO services ensure the lessons learned become a sustainable security posture.
If your organization wants to stop asking "can we be breached?" and start answering "how well do we react when we are attacked?", let's talk.