Every other method in 2.3 assumes some access: to a chip, a cloud provider, a facility, a person. Open-source intelligence (OSINT) does not.
Open-source intelligence is the application of the intelligence process (define the question, collect, evaluate, corroborate, state a confidence, deliver to a specific reader) to information that is legally and publicly available. A permit is open-source data. A news article about the permit is open-source information. An answer to the question "does site X contradict the Prover's declaration, and how sure are we" is open-source intelligence. Only the third is a verification product; the first two are inputs.
The MIRI draft agreement recommends open-source intelligence among the measures available before any agreement exists:
Invest in verification expertise and know-how, for instance by funding pilot verification efforts using open-source intelligence or satellite data.
Scher, Abecassis, Barnett, and Abeyta | MIRI (2025)
But what does it mean?
What Open Sources Reveal
| Category | What to look at | What it establishes | Where it is taught |
|---|---|---|---|
| Energy and grid | Interconnection requests (ERCOT, PJM, FERC, state PUCs), power purchase agreements, grid-operator load data, new substations and transformers, on-site generation permits | Upper bound on cluster power; commissioning timeline | 2.3.3 |
| Imagery | Optical (Planet, Maxar), thermal (SatVu), construction progress, cooling equipment, substations, parking lots, night lights | Physical existence and scale of a facility | 2.3.2 |
| Chip supply chain | Customs records (HS codes, Panjiva, ImportGenius), TSMC, Nvidia, SK Hynix disclosures, export licences, intermediaries | Where accelerators went and in what quantity | 2.3.4 |
| Financial disclosures | 10-K and 10-Q, capex on earnings calls, project finance and bond issues for data centres, cloud contracts and backlog, state tax incentives | Whether stated scale matches money spent | 2.3.4 |
| Permits and local records | Building, environmental, water and zoning permits, council hearings, local news, lawsuits, lobbying filings | Early signal of intent; often earlier and more precise than press releases | 2.3.1, 2.3.2 |
| Agreements | Cloud deals (Microsoft and OpenAI, Oracle and OpenAI, CoreWeave), government contracts, voluntary commitments, industry-body membership | Who is using whose compute; what was promised | 2.3.1, 2.4.4 |
| Technical traces | Model cards, papers, benchmarks, compute estimates (Epoch AI), API behaviour, public repositories | Back-calculated FLOP for models already trained | 2.3.1 |
| Network infrastructure | BGP and ASN announcements, IP ranges, dark-fibre leases, cable capacity, inter-site latency | Whether clusters are linked across sites | 2.3.5 |
| People | Hiring, job postings, career moves, conference talks, leaks, litigation between firms | Indirect signals; the verifier's route to interviewees | 2.4.1 |
Most rows are disciplines in their own right, and the sections named in the last column teach their methods.
The fullest treatment of open sources in the core literature is one entry in Six Layers:
Open-source intelligence (OSINT): Open-source information, such as social media posts and news reports, could reveal information such as hidden data center constructions, especially in combination with inspections of suspect sites (“Inspections for undeclared AI clusters” above). Open-source information could also find signs of non-compliant AI deployment, e.g., through its economic, scientific, or military impacts. However, these impacts might not be evident before the deployment causes harm or yields an unfair advantage.
Baker, Kulp, Marks, Brundage, and Heim | RAND (2025)
Read the whole of §4.4, the fifteen supplementary verification mechanisms, of which the OSINT entry above is one. The section explains why the paper calls them supplemental (limited in scope, or circumventable even when implemented well) and why an attempt to circumvent them tends to trigger the other layers.
Baker, Kulp, Marks, Brundage, and Heim | RAND (2025) | 5 min
Open Source Contributions to Verifiability
The NATO handbook on open sources, written for coalition staffs without their own classified collection, names the contributions open sources make to the rest of the intelligence effort. All of them transfer to verification, and together they say more precisely what "supplementary" means.
- Tip-off. Open data is usually first to show an anomaly: a grid filing before a press release, a permit before a building.
- Targeting. Open data tells the verifier where to spend the expensive instruments: which site to inspect, which cluster to attest, which cloud records to request.
- Context and validation. Open data lets the verifier read the result of the expensive instrument. Imagery shows a building; the permit says what it was built for and when.
- Plausible cover. If the verifier learned something through a closed channel, an open source that shows the same thing lets the verifier put the finding to the Prover without revealing the method. The AI verification literature does not discuss this role, and for a US-China agreement it matters: how do you allege a violation without disclosing how you know.
Optional: NATO Open Source Intelligence Handbook
If you are interested you can find the handbook here but it was written for staffs who use open sources so as not to burden a national intelligence service. A verifier is in the opposite position: the closed sources may not exist, or may belong to the other side. The roles still hold, but the reason to invest in open sources is not economy. It is that for much of the problem, open sources are all there is.
NATO | Supreme Allied Commander Atlantic (2001) | 85 min
What Open Sources Cannot Reach
Open sources have a few key limitations, notably:
- Open sources estimate capacity, not activity. Nothing visible from outside distinguishes training from inference from idle.
- All of it is evadable by a Prover who is trying: distributed training across sub-threshold sites, rented compute in a third country, hardware bought through intermediaries.
- Open sources therefore mostly produce discrepancy that justifies the more intrusive layers.
Rules for Handling Open-Source Evidence
Rules from the intelligence and investigative-journalism literature that apply unchanged to a verification file.
- Collect sources, not information. Keep an evaluated register of sources (what it is, who runs it, how reliable it has proved, how often it updates), not an archive of facts. The register is what lets the next analyst start in an hour rather than a week.
- Provenance is half the product. Every claim carries where it came from, when it was published, and when it was collected. Treat an undated or unsourced item as suspect. Archive every page, because pages change.
- Separate information from fact. The reader must be able to see, line by line, what is known, what is inferred, and what is guessed. State a confidence level using a shared scale (2.3.6 covers the scales).
- Corroborate across independent sources. Two sources that copy from the same press release are one source. Independence means a different collector with a different reason to be right.
There are public trackers that collect and publish information relevant for training runs. For example, the power explainer in 2.3.3 runs on Epoch’s data-center dataset. It is becoming increasingly easy to create automated trackers based purely on open-source information, since many things are uploaded to the Internet, and when they’re not, you can often find traces of activities anyway.
The Pentagon Pizza Index
You probably have heard of one of the most classic OSINT examples: the Pentagon Pizza Index. The observation dates to the Cold War: on the nights before major US military operations, pizza deliveries to the Pentagon and the CIA spiked, because staff were working late. Journalists noticed it before the 1989 Panama invasion and the 1991 Gulf War. Now Google Maps publishes live “popular times” for every pizzeria including those near the Pentagon; an account called Pentagon Pizza Report on X.com posts screenshots when several of them run unusually busy late at night, and it flagged activity hours before the June 2025 US strikes on Iran.
Tracking Hyperscale AI Data Center Growth with Satellite Imagery
Read Case Study 2 (feature identification at the Colossus site) and the Key Takeaways & Analysis that follow. The author, an imagery analyst, identifies the Memphis site's features one by one and shows which open sources confirm each: the utility's annotated Google Earth image, which located two substations and stated their capacities; the permit that limited the turbines to fifteen; the civil-society flight that counted thirty-five; the October 2025 satellite image that counted twelve. The takeaways state the method's limits: better for construction than for ongoing activity, useful as a check on announcements, blind to the inside of a finished building. The conclusion this section rests on: imagery "should be analyzed in conjunction with other publicly-available information sources."
Christina Krawec | Federation of American Scientists (2026) | 13 min
Exercise
Optional: Introduction to Investigative Journalism: Fact-Checking
How investigative journalism, which works from open sources, checks its own work before publication: record the source of every statement of fact, prefer primary documents to anything that does not cite its source, archive every link because pages change, and test each weak claim with the questions a fact-checking desk asks: can we say this with confidence, and can we show a skeptic the evidence. The same discipline applies to a verification file.
Mariam Elba | Global Investigative Journalism Network (2024) | 10 min
Going Further
If you want to spend more time on this, Research Clinic is a library of internet-research links that Paul Myers, an internet-research specialist, has maintained since 2003 for his training courses: search engines and academic sources, social-media research, people and company finders, domain and IP lookups, government and charity registers, geolocation and mapping, verification tools. The site is free. It is a directory, not a course: go to the category your question belongs to.
The Berkeley Protocol is the current standard for open-source investigation: collection, preservation, verification and chain of custody, written for investigators who will have to defend their findings to a hostile reader. It is long; the chapters on verification and documentation are the ones that transfer.

