A lead has arrived, and an analyst has assessed it. This section covers what the regime does next, and doing it proportionately, which decides whether a verification regime survives its first ambiguous case. The module then closes as it opened: with the picture as a whole, written for somebody who has to act on it.
Comparison With Declarations
The first step is the cheapest: compare the observed activity with what was declared, and ask for clarification. Most anomalies resolve here, which is the purpose. A regime that escalates every anomaly has no capacity left for the ones that matter.
Escalation
When clarification does not resolve the anomaly, the steps are, in order:
- Task additional collection: collect again, with a different stream.
- Seek corroboration: find an independent source that sees the same thing.
- Trigger an audit or challenge inspection: the treaty's own tools, and the point at which identification passes to resolution (2.4.3).
Each step costs more, politically and diplomatically, than the one before. Thresholds are set against both error costs: a false alarm costs the regime credibility; a miss costs the treaty.
The sequence ends in a regime action by design. A military strike is not a step in it and not a verification finding: it can destroy a facility but cannot establish whether a violation occurred. Resolution stays with the regime's adjudication. The worked case is the Iran record, where an intelligence tip opened the file and IAEA complementary access resolved it.
The Record
Whatever action the regime chooses, the record carries the confidence attached to the assessment, any dissent among the analysts, and the unresolved blind spots: what nobody could see, as distinct from what was examined and found clean. An assessment that hides its own uncertainty is worth less than one that states it, because the next decision-maker cannot calibrate on it.
Summary
- Intelligence identifies; the regime resolves. A tip is not a finding.
- The layer works without the monitored actor's permission and without new machinery: Wasil's national technical means, Six Layers' Layer 6, Scher and Thiergart's national-intelligence building blocks.
- Eight collection disciplines: literal collection (OSINT, HUMINT, SIGINT, CYBER, FININT) and nonliteral collection (IMINT, GEOINT, MASINT).
- The technical signatures are supplemental in Six Layers' scheme: circumventable on their own, but circumvention triggers the other layers.
- Power is the most quantitative signature and the one that decays fastest: performance per watt improves about 1.6× a year, roughly 500 candidate sites exist, and covert generation is plausible at the 130 MW scale.
- Money is the stream the field rates highest; open sources are the one it has written least about.
- People are the most enforceable and least reliable mechanism in the course, and the channel built for them decides whether an insider reports.
- Corroboration across kinds of evidence is what justifies escalating an anomaly, and base rates decide whether an alarm means anything.
Written Output
Write an overview of what intelligence-based mechanisms can see this year (what works, what is missing, what could be built) for a reader who has not taken this module. Budget: two hours. There is no source packet; use the module's materials and the open web. The limitations table from 2.3.6 is a draft of the memo's middle section. Carry it over, then add what the drill could not supply: this year's artifacts and the bottom line.
Written output · 2.3
What We Can See Today
The question: if a state were covertly training a frontier model right now, what could outside observers actually see? Write an overview a decision-maker can read in ten minutes. Cover each mechanism that exists today — overhead and thermal imagery, power and grid analysis, procurement, customs and financial tracking, open sources, human sources — with one sentence on what it establishes and one on what defeats it. Research is part of the task: find at least three public artifacts from the last two years yourself (a commercial satellite product, an enforcement action, a public tracker or filing) and cite each with a link. Open with the bottom line: the two mechanisms you would rely on most, and why. Close with the blind spot that concerns you most and the other verification layer that covers it. Do not draft treaty language: the question is what can be seen, not what should be signed.
- Budget
- about 900 words
- Reader
- A policymaker trying to understand what we can actually track, and what we cannot.
Judged on
- Accuracy about what each mechanism can establish
- Currency: public artifacts from the last two years, found and cited with links
- Limits stated as the papers state them: the layer reduces cheating and does not defeat it
- A prioritization with reasons, not a list of mechanisms
- Named blind spots, each with the mechanism elsewhere in the course that covers it
Before you start, read an example of the genre:
Optional: Nuclear Arms Control: U.S. May Face Challenges in Verifying Future Treaty Goals
An audit office asks, of the treaty goals that follow New START, what the United States could verify, with which methods, and where the gaps are. It is the memo's question at full scale. Read the highlights page first: what the office says can be verified, what cannot, and what it recommends.
U.S. Government Accountability Office (2023) | 20 min
Five controversies on which serious people disagree. The aim is to be able to state both sides.
The course is self-paced, so a debate here is a short two-sided memo rather than a discussion. Pick one controversy and read the others. State the strongest case for each side in turn, a short paragraph each, written to satisfy that side's strongest advocate. Then commit to a position in a sentence or two and name the evidence that would change it. The aim is to state both sides, not to win. For a live opponent, argue one side against an AI tutor arguing the other, and have it test your weakest point before you write the commitment.
1. How much verification is enough: "adequate" or "effective"? Carter-era doctrine required adequate verification (detect militarily significant cheating in time to respond); the Reagan administration raised the bar to effective verification, which Paul Nitze operationalized as "high confidence, not perfect confidence" (CRS R41201). The deeper point (Krass): there is no objective technical answer to how high a detection probability must be. The threshold is a political judgment, and no treaty is ever fully verifiable. At the far end of the argument: sometimes the verification burden is not worth it (Kimball, "Trust, but Don't Verify"). For an AI pause, a defensible confidence target has to be chosen, together with its false-positive and false-negative trade-off.
2. NTM sufficiency versus the necessity of intrusive access. National technical means (satellites, SIGINT) were sufficient to count large fixed objects, silos and launchers, and treaty law protected them with noninterference clauses. The IAEA record shows the limit: comprehensive safeguards plus NTM could verify declared material yet were structurally blind to undeclared activity (Iraq, North Korea), which is why the intrusive Additional Protocol had to be created. Verifying the correctness of a declaration is easy; verifying its completeness, that nothing undeclared exists, is the hard problem, and it is the AI-pause problem.
3. Commercial satellite imagery: gain or liability? The gain: commercial Earth observation has grown rapidly, many firms now sell frequent sub-5-metre optical and radar imagery, the IAEA runs a Satellite Imagery Analysis Unit, and NGOs do independent verification. The liability: imagery has hard physical limits (uranium mines look like copper mines, in-situ leaching leaves almost no footprint, thermal readings throw false positives); "open" imagery is not neutral, since firms withhold under government request (shutter control); and AI-generated satellite imagery threatens the evidentiary value of the pictures themselves.
4. Sources-and-methods versus sharing, and the credibility problem. Sharing a tip with a rival or an international body risks burning the collection that produced it, so arms-control treaties impose no duty to share; sharing is voluntary and inconsistent (Baker). Shared evidence "might be fake": the standing warning is the Niger documents of 2002, papers purporting to show that Iraq had sought uranium from Niger, passed between intelligence services and cited publicly before the IAEA found them to be forgeries (the WMD Commission's report covers the episode). For a US–China pause, the design question is how to make a tip credible to Beijing without exposing how it was obtained. That question leads to the confidentiality module's cryptographic tools (2.0), which this section does not cover.
5. Is the nuclear analogy the right frame? Arms control is the richest verification precedent available, and this module relies on it. A serious critique holds that the nuclear frame is misleading for AI: weights are not fissile material, there is no bulky unavoidable input to meter (chips are the closest), diffusion is faster, and dual-use is total. A systematic review of how the field uses the analogy catalogs these objections while arguing analogies can still inform policy despite the domain differences. The module's position: analogies can inform reasoning; they cannot prove a conclusion.
Write your two-sided memo here.
Review Questions
Seven questions on the module's objectives. Six of seven is the bar. A wrong answer names the section to revisit; revisit, then try again.
Which signature family is the cheapest to collect and the easiest to sanitize?
In Baker’s record of the four states investigated for undeclared facilities, how did most investigations begin?
The sequence, in order:
Commercial imagery shows a new substation beside a campus declared as logistics. Procurement and financial records show nothing unusual. Which step?
A program moves onto behind-the-meter gas generation to stay invisible to grid data. Which collection is most likely to catch the workaround itself?
“Our rule fires on only 1% of sites per year, so an alarm almost certainly means a covert program.” The error?
Splitting a training run across many smaller sites exploits which stated limit of this layer?

