A signature is not a finding, and a tip is not non-compliance. This section covers the step between them: how an analyst turns an anomaly into an assessment that can be defended, and how much confidence that assessment can carry.
Observation, Inference, Judgment
Three terms that ordinary speech runs together must stay apart here.
- An observation is what the sensor, the filing, or the source actually recorded.
- An inference is what that observation implies given some assumption.
- A judgment is what an analyst concludes given the inference, the alternatives, and their own confidence.
Most disputes that appear to be about evidence are disputes about which assumption is being relied on.
Sources of Error
- Alternative explanations. Asking for them is the default step, not the last resort. A power draw has more than one cause.
- Dual-use ambiguity. The same building, the same chips, and the same people serve permitted and prohibited activity. Ambiguity is the normal case, not the exception.
- Base rates. How many facilities of this kind exist, and how many are violations? Around 500 datacenters worldwide exceed 10 MW (the count given in 2.3.3). A signature that fires on all of them establishes nothing.
Evaluating a Source
Four questions apply to every input:
- Reliability. Has this source been right before, and about this kind of question?
- Timeliness. Does it describe the present, or the facility as it stood two years ago?
- Corroboration. Does an independent stream see the same thing? Streams that share an origin are not independent.
- Confidence. What would have to be true for this to be wrong, and how likely is that?
From Anomaly to Suspected Non-Compliance
The sequence is anomaly → verification lead → suspected non-compliance. Each step is a decision that someone has to justify, and each has a different evidentiary bar. Escalating too fast costs credibility; escalating too slowly loses the window in which the activity is observable.
Corroboration across kinds of evidence — physical, procurement, financial, digital, organizational — justifies a step up. They fail in different ways, so their agreement is informative.
Good analysis states its uncertainty rather than manufacturing certainty. The cautionary case is the 2002–03 Iraq WMD estimate, "dead wrong" and held with high confidence, which the post-mortem drill below takes apart. The phrase is the Silberman–Robb WMD Commission's own, and its report is the public post-mortem.
First, a question to answer before the drill: among roughly 500 candidate sites, how often does a detection rule that seems cautious fire on an innocent one? Commit a guess in the drill before doing the arithmetic.
Drill Bench
Three drills. The first turns the base rate into arithmetic; the second is a post-mortem of an assessment that was confident and wrong; the third records, for each signature family, its standard failure case and the mechanism elsewhere in the course that covers it. The third table is a draft of the memo's middle section: 2.3.7 asks for the same table, mechanism by mechanism, with the artifacts you find.
Limits of the Layer
Four papers agree. NTM is "a valuable starting point" but limited against disguise, distribution, and software-level violations (Wasil et al.); detecting a secret datacenter "is unlikely to be the load-bearing part of a verification regime" (Scher and Thiergart); imagery, OSINT, and financial audits are "merely supplemental", although an attempt to circumvent them may trigger other, more reliable mechanisms (Six Layers); and military facilities are a named blind spot (MIRI).
These limits are why the layer is supplemental rather than load-bearing, and they do not make it useless. Its value is deterrence, lead generation, and raising the cost and complexity of cheating: each limit is also a cost imposed on the cheater, and another mechanism of the course covers each named gap, as recorded in the third drill. The limits widen over time, because efficiency gains shrink the large-facility signature each year and distributed development spreads compute across smaller sites. The correct statement of the layer's value is therefore that it reduces the effectiveness of cheating; it does not defeat it.
Optional: Does Distributed Training Undermine Compute Governance?
A quantitative study of the distributed-development limit. A simulator calibrated on published low-communication training runs finds that every existing or proposed compute threshold can be crossed on nodes below every proposed monitoring threshold, over consumer-grade internet. Nodes of that size are invisible to thermal and electrical means, so enforcement falls to network monitoring and in-person observation. Appendix G grades the countermeasures. The ones that hold are this module's: whistleblowers (a distributed operation needs a workforce that grows with its node count), financial and procurement intelligence, chip registries that count memory as well as compute, and challenge inspections coordinated across the whole suspected network. Bandwidth caps and traffic monitoring do not hold. Read the whole paper, appendices included.
Robi Rahman | Machine Intelligence Research Institute (2026) | 35 min
For each card from Scher and Thiergart's Locating Compute Building Blocks table, record your own feasibility rating and timeline with a short reason before the card shows the authors' rating, then compare and argue any disagreement. The intelligence deck covers this module's mechanisms; the full table is optional further practice and preparation for 4.1, which unseals your Module 2 ranking.
Reading
Optional: Psychology of Intelligence Analysis
The primary source for the biases named in the post-mortem drill, free in full from the CIA's own press. Part III covers the biases; chapter 8 covers analysis of competing hypotheses.
Richards J. Heuer, Jr. | CIA Center for the Study of Intelligence (1999) | 30 min
Optional: A Tradecraft Primer: Structured Analytic Techniques for Improving Intelligence Analysis
Methods applied directly in the post-mortem drill: Quality of Information Check (pp. 10–11), Indicators or Signposts of Change (pp. 12–13), and Analysis of Competing Hypotheses (pp. 14–16).
US Central Intelligence Agency (2009) | 15 min

