# Scrutica — Complete Reference for AI Agents *Last updated: 2026-07-23. License: CC BY-SA 4.0 (data) except where source-attributed otherwise. Canonical URL: https://scrutica.com* This file is the long-form agent-consumption format described at [llmstxt.org](https://llmstxt.org/). For the navigational index version see https://scrutica.com/llms.txt. The contents below mirror the substantive content of the platform's methodology, entity model, and data-quality discipline in a single flat document suitable for full-context ingestion by Claude, GPT, Gemini, Perplexity, or any other agent fetching pages by URL. --- ## What Scrutica is Scrutica is the AI compute infrastructure intelligence platform built for governance research. It tracks where the compute lives, who controls it, and who can't get it — the cross-substrate join over facilities, corporate control, supply chains, and export-control exposure. Facility coverage centers on Epoch AI's frontier-DC catalog (frontier-scale US, UAE, and China sites), the IM3 OpenStreetMap US data-center bulk corpus, the New York ISO load-interconnection queue (and analogues across the seven US ISOs/RTOs), and sovereign-program-linked international sites; tier-2 sovereign-facility ingest (HUMAIN, Yotta, Tencent, EuroHPC) is in active sprint. The platform connects layers historically siloed across separate research outputs: facility-level compute infrastructure (Epoch Frontier Data Centers, Epoch GPU Clusters, OpenStreetMap-derived facility coverage, PeeringDB facility registry), bilateral supply-chain relationships (a licensed supply-chain database, SEC EDGAR Exhibit 21 ownership filings, CSET ETO aggregations), export controls (BIS Entity List from the Federal Register, OFAC SDN list, EU dual-use list, Japanese end-use restrictions), sovereign AI procurement (USAspending, EU Tenders Electronic Daily, UK Find a Tender, federal solicitations across multiple jurisdictions), capability evaluations (Epoch's notable-models database, MLPerf systems benchmarks, METR autonomy time-horizon evaluations), and financial substrate (a licensed corporate-ownership database for fundamentals and industry classification, a licensed fund/LP database for LP→fund chains, and licensed entity-hierarchy data held under subscription). The core dataset (counts as of 2026-07-23): - 4,550 compute-related facilities across 110 countries - 21,037 bilateral supply-chain edges (15,049 canonical-deduplicated) - 129,061 organizations in the canonical universe; 24,873 compute-relevant; 5,093 appearing as supplier or customer in the supply-chain graph - 3,422 BIS Entity List (and related export-control) designations cross-referenced onto compute-universe entities - 36 sovereign AI programs with announced-versus-deployed reconciliation - 33,627 investment / capital-flow records - 4,317 discrete compute-deployment records (chip-count × facility × vendor disclosures) These counts stream into the homepage and the Dataset JSON-LD on every refresh; they are not estimates, they are direct row counts from the production substrate. The audience is AI safety and governance researchers at frontier labs, universities, and independent research institutes; policy analysts working export controls, compute policy, and national-security technology assessment; journalists covering AI infrastructure and the semiconductor supply chain; academics; congressional staff; and increasingly the AI-mediated discovery layer itself — Claude, ChatGPT, Perplexity, Gemini, and the MCP-using agents that wrap them. Scrutica's distinctive asset is the data-shape rather than data-volume: a single integrated cross-layer query surface, per-record provenance with full source-URL auditability, a methodology page where the user can adjust assumption variables and re-derive in browser, and the same substrate exposed through an MCP server for agentic discovery. Primary-source provenance, transparent uncertainty, and named methodology with versioning are the design commitments every record carries. ## How to cite Scrutica Standard citation: ``` Scrutica. "[Page Title]." [Methodology Version], [Date]. https://scrutica.com/[path] ``` BibTeX: ```bibtex @misc{scrutica2026, author = {Scrutica}, title = {AI Compute Infrastructure Intelligence Platform}, year = {2026}, url = {https://scrutica.com}, } ``` Per-entity citation: every facility, company, sovereign program, BIS designation, supply-chain edge, scenario, and methodology page exposes a "Cite This" panel that emits APA + BibTeX + Chicago + JSON-LD with the entity's canonical Scrutica ID and a methodology version stamp. Citing a specific estimate should include both the methodology version (`v2.1`, etc.) and the estimation path (e.g., "Hardware Path, v2.1") so readers can reproduce the calculation against the same parameter set. The corrections log at https://scrutica.com/corrections records every methodology version change with the rationale and what previous calculations would now yield differently — this lets citing work remain accurate even as Scrutica's methodology evolves. For agent integration: the Model Context Protocol server at `https://scrutica.com/api/mcp` exposes the same substrate as ten tools (search, get entity, supply-chain edges, export-control lookup, Entity List change log, FLOP estimation, facility comparison, scenario lookup, sovereign-program data, methodology), each returning structured JSON with provenance fields preserved. --- ## Site map (by category) Scrutica's information architecture organises routes into six analytical categories plus a cross-category tools rail. The hub page (`/`) renders all six categories inline; the sidebar exposes the same hierarchy on every interior page. The MCP server (`/api/mcp`) exposes a subset of these surfaces as agent-callable tools. ### Infrastructure (4 routes) Facility coordinates, GPU deployments, FLOP capacity estimates. The substrate layer of the platform — where the physical compute actually lives, who operates it, what chips it runs, how much power it draws. - `/map` — Interactive deck.gl + Mapbox map of every tracked compute facility worldwide. Supply-chain arcs animate the bilateral relationships between supplier and customer organizations; heatmap overlays surface concentration by region. - `/supply-chain` — Approximately 21,037 bilateral supplier-customer relationships from a licensed supply-chain database + SEC Exhibit 21 + CSET ETO. Filter by relationship type (supplier, customer, subsidiary, parent, competitor, JV partner, co-investor), tier classification (1=direct, 2=2-hop, 3=3-hop+), supply share where disclosed, 3-month price correlation. - `/entities` — The Entities instrument: facility, company, and country registers plus the Ownership Investigator as sibling views; the facility table anchors the root. Filter by operator, country, hardware type, power tier, freshness; sort by capacity, GPU count, last update. - `/flop-engine` — Facility-level AI training capacity via three independent estimation paths (hardware, power, capital). Every parameter adjustable and auditable in browser; sparsity factor, utilization fraction, peak chip FLOPS all assumption-adjustable. ### Chokepoint Analysis (9 routes) Cascade propagation, scenario simulation, concentration scoring. The "what breaks when access gets cut off" layer. - `/cascade` — Dependency propagation across thousands of supply-chain edges, weighted by supply share (where disclosed), 3-month price correlation (the markets' real-time dependency-strength signal), and editorial criticality (expert-assessed substitutability factor for the input class). - `/chokepoint-anatomy` — Where AI compute supply narrows to single suppliers or single jurisdictions. HHI (Herfindahl-Hirschman Index) computed per supply-chain layer; node-level downstream-dependency tracing. - `/compute-leverage` — Veto-count instrument over the frontier-AI compute stack: per leverage point, who holds it, the veto verdict at 18 months and 5 years (leverage = switching-cost × redeployment-time, not market share), and the accountability tag that checks the holder. Per-run verdicts re-resolve via accelerator/site overrides. - `/scenario-analysis` — Geopolitical stress tests across five named scenarios: Taiwan Strait blockade, IRGC missile range, Saudi grid (Abqaiq 2.0), Tokyo earthquake, submarine cable severing. - `/stress-index` — Composite 0-100 score across four dimensions (demand pressure, supply concentration, capacity utilization, power availability), refreshed quarterly with explicit methodology version history. - `/compute-supply-model` — Forward 2-5 year compute-supply projection — which constraint (power, advanced-node chip production, HBM packaging) binds when. Built on substrate trajectories and announced-vs-delivered tracking. - `/concentration-monitor` — Ranked AI-compute concentration with HHI, substitute availability, demand growth trajectory per chokepoint. The reference surface for "which single-supplier dependencies are tightening fastest." - `/export-controls/diversion` — Diversion view of the Export Controls instrument: a diversion-thesis audit at the primary-source layer, rendering the Culper 2026-05-13 NVDA short-seller thesis as a worked example, with licensed supply-chain edges and entity-resolved counterparties anchoring the analysis. - `/chokepoints/early-warning` — Quarterly anomaly + consolidation detection across seven chokepoint layers (advanced-node logic, HBM, EUV lithography, CoWoS packaging, EDA tools, advanced photoresist materials, deposition equipment). ### Governance (9 routes) Export controls, sovereign programs, allied-coordination gaps. The legal-access layer. - `/export-controls` — BIS Entity List designations cross-referenced onto compute infrastructure; multi-regime view across US (EAR 3A090 / 4A090 — the advanced-computing items classification), Netherlands / EU semiconductor equipment, Japanese end-use, UK strategic-export. - `/export-controls/changes` — Entity List change log: Federal Register actions grouped one event per notice, derived from the designation substrate (rows group by the first FR citation they carry; unresolvable citations are listed unparsed rather than guessed at). Per-entity records with removal flags and company cross-references; RSS 2.0 twin at `/export-controls/changes/feed.xml`, one item per notice over the last 24 months, plus a per-organization feed at `/export-controls/changes/entity//feed.xml`. - `/export-controls/brief` — Quarterly Brief view: a one-screen policy artifact — what changed on the Entity List over the trailing 90 days, a per-destination chart, and which compute-infrastructure organizations a designation newly reaches. Quiet quarters render the honest "no actions" artifact with the last action stated. - `/sovereign-ai` — 36 national programs with announced-vs-deployed reconciliation, government-only vs headline figures, reality ratios, interdependence scores, inflation decomposition. - `/threshold-atlas` — Which facilities can train models across the EU AI Act 10²⁵ FLOPs trigger, China CAC ~10²⁴ FLOPs (inferred from registration pattern), and historical US EO 14110 10²⁶ FLOPs threshold (since rescinded but retained as a calibration reference). Reverse calculator plus the SemiAnalysis 100K-H100 cross-validation. - `/geopolitics` — Per-country compute capacity across seven supply-chain layers (chip design, foundry, HBM, packaging, equipment, materials, data-center hosting) with the export-control constraint overlay rendered as visual gates. - `/compute-visibility` — Two distinct lenses on attribution: ownership transparency (Tier 1-4 capacity-weighted index) and hosted-vs-controlled attribution by country (where compute physically sits vs which jurisdiction's entity has operational control). - `/coordination-gaps` — Facilities under coordination arbitrage — sitting in an allied export-control regime but reaching a non-coordinated jurisdiction through the ownership chain. The structural-arbitrage map. - `/export-controls/impact` — Impact view of the Export Controls instrument: simulate the structural reach of a hypothetical BIS Entity List addition. What downstream entities would be affected and through what bilateral edges? ### Economics (4 routes) $/PFLOP-day, training-cost curves, trade-impact modelling. The price-signal layer. - `/cost-index` — $/petaFLOP-day across AWS p5 / p5e, Azure ND H100 v5 / NDmi A100 v4, GCP A3 mega / A2 ultra, CoreWeave HGX H100, Lambda 8xH100 — hourly published rates normalized via Epoch hardware specs. - `/export-controls/trade` — Trade view of the Export Controls instrument: did US export controls reduce China's compute access or redirect chip flows through Singapore / UAE / Malaysia? 15 years bilateral trade data pre- and post- each major BIS rule. - `/training-cost` — Frontier training cost rises 10,000× while replication cost falls 2-18× per Epoch's published cost trajectory. The divergent curves shape which capability-access controls bind where. - `/semiconductor-cycle` — 40 years of WSTS billings + Epoch GPU shipments + BIS event markers as a single timeline (1986-2026), with cyclicality overlay. ### Capabilities (3 routes) Scaling laws, training evidence, autonomy benchmarks. The model-side layer that compute substrate enables. - `/capabilities/scaling-explorer` — Benchmark scores mapped to training compute across 31 evaluations. Per-benchmark log-linear fits with R²; hardware provenance and export-control status on every data point. - `/capabilities/training-evidence` — Cross-reference vendor training claims against facility power, chip availability, financial data, export-control status. Evidence aggregation, explicitly not verification. - `/capabilities/autonomy-monitor` — METR time-horizon scores against training compute + regulatory FLOP thresholds. ### Reference (7 routes) Methodology, derivation chains, corrections, glossary. The auditability layer. - `/methodology` — Complete derivation chains for every estimation path. Formulas in full; parameters adjustable in browser; authority-tier definitions; cross-validation methodology; version history with append-only changelog. - `/related-work` — Prior art and adjacent research on AI compute infrastructure — Epoch AI, CSET, RAND, Stanford HAI, EPRI, GovAI — with an explicit account of what Scrutica covers and deliberately does not. - `/corrections` — Public log of data corrections and methodology fixes. Each entry: what was shown, the correct value, why it broke, what prevents recurrence. - `/methodology/glossary` — Technical definitions and governance relevance for 77 terms spanning semiconductors, AI training hardware, export controls, compute policy. The Glossary view of the Methodology instrument. - `/compute-in-context` — Scale from API call to global compute buildout — citation-grade explainers for policy briefings. - `/changelog` — Recent updates across tracked facilities and organizations. Hourly feed. - `/about` — Platform scope, data limitations disclosure, CC-BY-SA 4.0 license, methodology anchoring. ### Tools (3 cross-category instruments) - `/compare` — Multi-entity comparison across facilities, organizations, countries, sovereign programs — radar overlays, resilience scoring, metric tables, export to CSV / JSON / BibTeX. - `/ownership` — Walk any compute entity to ultimate beneficial owner with multi-source provenance per hop (SEC filings / Companies House / licensed corporate-ownership database / analyst-database citation chain). - `/query` — Natural-language interface to the full substrate via Claude Sonnet 5 (failing over to Claude Opus 4.8 then GPT-5.5) routed through a documented tool surface (55 named tools across 14 layers). Every numerical claim wraps in a citation tag with authority-tier provenance; a server-side verifier flags any cited record_id absent from the underlying tool-result history. ### Agent-native interfaces - `/api/mcp` — Model Context Protocol server (Streamable HTTP transport per spec 2025-11-25) exposing ten tools. - `/api/mcp/llms.txt` — LLM-optimized API reference for the MCP server. - `/.well-known/mcp/server-card.json` — MCP discovery manifest (SEP-1649 draft). - `/.well-known/mcp` — MCP server-discovery directory (SEP-1960 draft). --- ## Data quality principles Every record on Scrutica carries `data_source` and `source_url` columns. Every derived value carries `is_estimated: true`. The platform's data discipline is enforced by 16 architectural rules summarized below; the full text lives at https://scrutica.com/methodology and in the project's source-control-tracked CLAUDE.md. Each principle is paired with the concrete failure mode that motivated it — these are scars from past mistakes encoded as architectural rules, not abstract best-practice posturing. **1. No fabricated data.** Every entry traces to a real facility, transaction, or filing. If a value is unknown, it is stored as `null` rather than estimated through plausible-sounding guesswork. No synthetic rows backfill missing observations. Failure mode this avoids: the "looks reasonable" estimate that future analysts treat as primary data — a single fabricated megawatt rating, propagated through a derivation chain, can shift an entire country's compute-capacity bucket by a percentage point. **2. Transparent estimates.** Records with derived or estimated values flag `is_estimated: true` and document the derivation methodology in the row's notes or in the linked methodology page. Tables where every row is primary-source by construction (Federal-Register-anchored BIS designations, EDGAR-anchored SEC filings) or carries a dedicated quality-metric column (`confidence_score` on `bis_crossref_matches`, `authority_tier` on export_control_designations) supersede the boolean — but justify any deviation explicitly in the data-access module's header comment so the suppressed boolean does not become a hidden footgun. **3. Source documentation.** Every record has a `data_source` field naming the corpus (e.g., `epoch-ai`, `licensed-supply-chain`, `bis-entity-list`, `usaspending`, `peeringdb`) and a `source_url` pointing to the underlying record where applicable. Provenance is non-negotiable. The Methodology Explorer renders the chain via PROV-O `wasDerivedFrom` JSON-LD so agents can verify any number by walking the chain back to the primary measurement. **4. Honest uncertainty.** Confidence intervals on every estimate. The FLOP Estimation Engine emits point estimates with explicit upper and lower bounds derived from the underlying assumption ranges. The methodology page lets users adjust assumption variables (utilization fraction, sparsity factor, peak chip FLOPS, power efficiency) and re-derive the estimate in browser. Where the cross-path divergence exceeds 2× (the Scrutica-editorial contradiction threshold), the facility intelligence card surfaces the discrepancy explicitly rather than picking a winner. **5. No false precision.** Where a value cannot be determined with confidence, the platform stores `null` rather than guessing. The Methodology page distinguishes Tier 1 (primary measurement / filing) from Tier 2 (research database) from Tier 3 (press / analyst secondary) from Tier 4 (inferred / estimated); confidence cascades through derived metrics so a Tier-4 input cannot produce a Tier-1 output. Statistical rigor — error bars, confidence intervals — applied to compute-governance data where the conventions are still being established. **6. Temporal provenance.** Every entity has a `data_vintage` field — when the underlying observation was produced, not when it was accessed. Stale data (>12 months for fast-moving markets such as GPU shipments or HBM allocation) is flagged explicitly. The append-only `compute_capacity_snapshots` table tracks history so the platform can show "previous value" alongside the current value where they differ. Time is a first-class column, not a metadata afterthought. **7. Contradiction surfacing.** When primary sources disagree — Hua Hong's announced capacity vs SMIC's disclosed capacity vs Federal Register entity-list designations, or NVIDIA earnings-call commentary vs licensed-supply-chain-derived supply-share numbers vs CSET ETO aggregations — the platform shows both with an explicit cross-source-divergence flag and the editorial disposition (which value the platform's UI prefers and why). The `organization_data_quality_flags` table records the contradiction, the resolution disposition, and the dispositioning rationale. Silent winner-picking is a violation; visible disagreement is the discipline. **8. Verify, never hedge.** Words "likely," "probably," "presumably," "assume," "both could be correct" are banned when discussing facts. They signal unverified claims masquerading as assertions. The correct response is to check the primary source, not to hedge linguistically. This applies to architectural decisions, record counts, API schemas, column names, organization aliases. If a value cannot be verified, the platform stores `null` and surfaces the gap rather than emitting a hedge that future analysts read as a soft commitment. **9. Sense-check agent outputs.** Agent outputs from research subagents, MCP-tool calls, or LLM-mediated synthesis are not verified findings. Before any agent finding lands in the substrate or in a user-visible surface: (a) is the source primary or secondary? (b) is data vintage being presented as current? (c) does the magnitude pass a smell test against known benchmarks? (d) what coverage limitations would change the conclusion? (e) is the metric measuring what it claims to measure? Estimates from secondary sources are labeled as such, never presented as verified data points. Agent outputs go through an explicit dispositioning pass before they can advance into substrate; the pass leaves a record at https://scrutica.com/corrections so future analysts can trace why a particular value was accepted, rejected, or escalated. **10. Lookup data in the database, not in Python dicts.** Natural-key → canonical-id mappings (LP-name → org-id, fund-manager-name → org-id, alias variants for organizations whose name transliterates differently across English / Chinese / Japanese / Korean sources) live in SQL tables with UNIQUE constraints — not hand-curated Python dicts. Adding a mapping is a SQL insert, never a code edit. The `organization_aliases` table contains alias rows resolving variant spellings, transliterations, and acquisition-related name changes onto canonical organization IDs. Code that needs the mapping loads it once at run start as a read-only cache; the canonical pattern is documented at `scripts/lib/sovereign_lp_resolver.py` and `scripts/lib/fund_manager_resolver.py`. The reason: hand-curated dicts become orphaned and re-introduce stale mappings on every refresh; SQL-backed mappings are diff-visible, gate-checked, and survivable across pipeline rewrites. **11. Write-time resolution over post-hoc backfill.** FK columns whose nullability would let ingests punt (e.g., `supply_chain_links.source_org_id` accepting NULL while a free-text `counterparty_name` column carries the actual value) resolve at write time via a database-backed alias table. Backfill scripts not wired into the canonical refresh path are dead code — they get re-introduced on every refresh as the underlying source data overwrites the manual fix. Any new substrate write path emitting a NULLable FK either resolves via the established alias-index resolver (case- and whitespace-insensitive lookup against `organization_aliases ∪ organizations.name`) or queues the unresolved value into an `unresolved_*` table where it surfaces for editorial dispositioning. **12. Free-text resolver tokenization audit.** When a resolver reads from free-text source data (vendor `investor`, `name`, `description` fields, licensed-supply-chain `counterparty_name` strings, USAspending recipient-name fields), the platform samples the tokenization shape — multi-value strings (e.g., "KKR; Blackstone"), embedded delimiters (e.g., "Goldman Sachs / Apollo"), parenthetical aliases (e.g., "Samsung (Korea Investment Corp subsidiary)"), encoding variants (Chinese transliterations, Japanese kana, Korean Hangul) — before declaring atomic-name lookup sufficient. Composite-string sources fail silently under naive lookup; the canonical pattern (`_split_composite_name()`) strips relational tokens, splits on a multi-delimiter set (`/`, `;`, `,`, ` and `, ` & `), expands parenthetical aliases, deduplicates, and only then runs the alias-index lookup. **13. Column-population gates need paired closure.** Parity-check gates that flag NULL drift (e.g., "investments.lead_investor_org_id has 4,210 NULL values") surface a problem but do not close it. Every gate has paired closure: a curation pipeline that resolves a slice of the NULL rows automatically, an auto-stub generator that creates organization stubs from free-text names, a one-shot bulk seed that lands editorial dispositions, or an explicit accept-as-known-residue documented in the verification log. A gate-without-closure produces honest accounting that masquerades as completeness. The closure pattern: gate-as-alarm + closure-as-pipeline + accept-as-known-residue + editorial-dispositioning ladder. **14. Removal-vs-relocation discipline.** When a route is removed alongside others being relocated within the same commit, the absence of a redirect is deliberate — the removal commit's message documents intent. Audits flagging "missing redirect" on a deleted route check the originating commit before treating absence as a defect. Generalizes: when a multi-route cleanup commit combines relocate-with-redirect and remove-without-redirect decisions, treat the asymmetry as deliberate until the commit message says otherwise. The convention preserves the editorial signal that a removed surface was intentionally removed (not "lost"). **15. Analytical-artefact removal discipline.** Killing a peer-comparison surface requires same-session paired pipeline-audit + remediation tasks + consumption-site flags + premise-test. Removing an indicator from a dashboard without surfacing the removal in every downstream consumption site that referenced it is a partial closure — the dashboard says "this indicator is gone" but the facility-detail page still asserts the indicator's narrative. Scrutica-specific consumption sites that need flag-at-consumption when a peer-comparison surface is killed: `/facilities/[id]` (the facility intelligence card), `/map` tooltips, `/cost-index` $/PFLOP-day surfaces, `/sovereign-ai` exposure rollups, `/concentration` plus `/cascade` criticality weights. **16. Editorial corrections via SQL row, not one-shot scripts.** Name-typo corrections for primary sources (SEC filings, licensed-database records, BIS Federal Register entries with transliteration variants) MUST land as a new row in `editorial_name_corrections` — NEVER as a one-shot substrate UPDATE script. One-shot scripts are overwritten on the next ingest refresh and violate principle 11. The corrections layer preserves the verbatim primary-source form (so the audit trail back to the Federal Register / SEC filing remains exact) while surfacing the corrected display form at render time. The pattern: substrate stores verbatim source; display layer applies corrections; the user sees the corrected form with a hover-state showing the verbatim original. --- ## Methodology — FLOP Estimation Engine The FLOP Estimation Engine derives facility-level AI training capacity through three independent estimation paths, each with explicit confidence tiers and adjustable parameters. The full methodology, including the formulas in symbolic form and the assumption ranges per parameter, is at https://scrutica.com/methodology#flop-engine. **Path 1: Hardware Inventory.** Where chip-count disclosures exist (SEC filings, vendor press releases, licensed-database equipment-purchase data, BIS-related export-control records, EDGAR 10-K narrative sections, the Epoch GPU Clusters corpus, MLPerf systems benchmark records), the engine computes peak FLOPS as: ``` peak_FLOPS = hardware_peak_per_unit × unit_count × utilization × sparsity_factor ``` Peak FLOPS per unit comes from vendor-published specifications: - NVIDIA H100 SXM5: 989.5 BF16 TFLOPS dense (NVIDIA datasheet) - NVIDIA H200: same compute as H100, increased HBM3e capacity (141 GB) and bandwidth (4.8 TB/s) - NVIDIA B100: vendor-published TFLOPS (Blackwell architecture, 2024) - NVIDIA B200: ~2× B100 in dense BF16, with FP4 / FP6 lower-precision modes for inference workloads - NVIDIA GB200 NVL72 / GB300 NVL72: rack-scale systems integrating Grace CPU + Blackwell GPU + NVSwitch fabric - AMD MI300X: vendor-published TFLOPS (CDNA3 architecture, 304 compute units, 192 GB HBM3) - AMD MI325X / MI355X: MI300 refresh series with HBM3e capacity / bandwidth upgrades (1000 W TDP for MI325X) - Google TPU v4 / v5p / v6e / v7: per-chip BF16 TFLOPS from Google Cloud documentation - Intel Gaudi 3: vendor-published TFLOPS (2024) - Huawei Ascend 910B / 920: published specifications cross-validated against teardown analysis (Tier 2-3 signal — primary-source vendor disclosure is partial) - Cerebras WSE-3: wafer-scale design; vendor-published TFLOPS - Tenstorrent Wormhole / Blackhole: vendor-published TFLOPS Utilization defaults to 0.50 for sustained training workloads (representative of academically reported MLPerf training-system utilization for transformer-class workloads). The methodology page exposes utilization as an in-browser adjustable assumption with the default range 0.30 to 0.70; values outside that range require justification in the assumption block. Sparsity factor: dense (1.0×), 2:4 structured (1.0× — Scrutica defaults to dense BF16 because the 2:4 sparsity factor's 2.0× theoretical doubling rarely materializes in trained-model wall-clock throughput; the methodology page explains the contradiction between NVIDIA's published 2:4-with-sparsity numbers and observed MLPerf training-throughput records). Where a facility's workload is documented to use FP8 or INT8 precision in practice (inference clusters, some training systems), the engine renders those values as a separate annotation rather than collapsing them into the BF16-equivalent. **Path 2: Power Envelope.** Where power-capacity disclosures exist (Epoch's Frontier Data Centers dataset, OSM-derived facility coverage with associated permitting documents, county-level interconnection-queue filings for US facilities, NYISO load-interconnection queue, IM3 OpenStreetMap US data-center bulk corpus), the engine bounds compute capacity by power. The conversion: ``` peak_FLOPS_bounded = power_capacity_MW × utilization_power × power_efficiency_factor ``` The power-efficiency factor is the chip-family-specific FLOPS-per-watt curve from vendor specifications. Utilization here is a separate variable from Path 1 — the question being asked is "what fraction of nameplate power capacity is actually drawn by compute" rather than Path 1's "what fraction of chip-peak FLOPS is actually utilized." The methodology page exposes both as separately adjustable assumptions because they decouple in practice (a facility can be power-constrained-but-fully-utilized, or power-overcapacity-but-underutilized, or both constraints binding). Power envelope estimates carry intrinsically wider confidence intervals than hardware-inventory estimates because the chip mix is unknown — a 50 MW facility could host H100 generation chips (lower FLOPS per watt) or Blackwell generation chips (higher FLOPS per watt) or AMD MI300X (different FLOPS-per-watt curve again). The methodology page documents the assumption ranges that drive the resulting bounds. **Path 3: Capital Cost.** Where capital expenditure is disclosed (10-K filings, earnings transcripts citing specific facility CapEx, licensed-fund/LP-database valuations of GPU-cloud equity rounds, licensed-corporate-ownership-database capex disclosures, USAspending federal procurement awards with facility attribution), the engine decomposes facility cost into hardware components (chips + servers + storage + networking), construction components (shell + cooling + power-distribution), and operational components, then back-derives implied chip count from the hardware component using vendor-published unit pricing. Path 3 produces the widest confidence intervals of the three. The underlying capex-to-chip-count chain compounds uncertainty over chip-mix assumptions, per-chip pricing (which varies with allocation tier, contractual commitments, and large-customer negotiation), and the construction-vs-hardware split (which varies dramatically with cooling architecture — liquid cooling commits significantly more capex to facility plant than air cooling). The methodology page exposes each as separately adjustable. **Cross-path agreement and divergence.** The three paths typically agree within ~1.5× for well-disclosed facilities (the Scrutica-editorial cross-path-divergence threshold for "high agreement"). Where they diverge by >2×, the discrepancy surfaces in the facility intelligence card as a contradiction flag with the disposition — which path is preferred and why. The decision rule is rough: prefer Path 1 when chip-count disclosure is Tier-1 confident; prefer Path 2 when power capacity is Tier-1 confident and chip mix is dominantly inferred; prefer Path 3 when both upstream paths carry Tier-3 signal and capex is the strongest available primary source. **External cross-validation.** The methodology cross-validates against Epoch AI's Frontier Data Centers corpus (for the facilities both platforms cover) and against MLPerf systems benchmark records (where the named training system reports observed FLOPS at a documented hardware configuration). The cross-validation pass is published at https://scrutica.com/methodology#external-validation with the residual analysis broken out by tier. **Methodology versioning.** The complete version history is at https://scrutica.com/methodology#version-history. Per-version changes are append-only; older estimates carry the version stamp under which they were computed so an analyst citing a 2024 estimate sees the methodology that produced it, not the latest version. Citations should include the version stamp explicitly. **What the engine does not do.** The FLOP Estimation Engine produces capacity estimates, not training-run estimates. Capacity is the upper bound — the facility's theoretical training output if dedicated to a single run. Actual training runs are sub-allocations within the capacity envelope; the Threshold Atlas at `/threshold-atlas` reasons about training-run feasibility (model architecture × token volume × precision × duration) against the capacity envelope but is a separate analytical pass. --- ## Methodology — Compute Cost Index ($/petaFLOP-day) The Compute Cost Index normalizes published hourly pricing for AI-training-class GPU instances across major clouds (AWS p5, p5e; Azure NDmi A100 v4, ND H100 v5; GCP A2 ultra, A3 mega; CoreWeave HGX H100; Lambda 8xH100; Crusoe Cloud H100; Together AI H100 / H200) into a single unit: dollars per petaFLOP-day. The conversion uses Epoch AI hardware specifications for chip-peak FLOPS (BF16 dense by default — adjustable to FP16 / FP8 / INT8 in browser) and assumes 24-hour utilization at the documented utilization fraction. The Index is computed nightly via the `/api/cron/cloud-pricing` handler, which pulls published rate-card pages, parses GPU-instance pricing into chip-counts × hourly-rate, applies the Scrutica Compute Unit (SCU) conversion, and writes to the `compute_pricing` table. The full $/PFLOP-day formula and intermediate values are published per row, with the underlying assumption variables adjustable on the cost-index page. **Conversion formula.** For each cloud instance: ``` SCU_hourly = chip_count × per_chip_peak_BF16_TFLOPS × utilization $/PFLOP-day = (hourly_USD × 24) / (SCU_hourly / 1000) ``` The published table renders both the absolute $/PFLOP-day and the multi-cloud spread (range across providers for an equivalent capability class — the more useful policy metric, because it surfaces export-restricted-region price premiums and contractual-discount-driven variance). The platform deliberately separates point estimates per provider per hardware tier per region per procurement model from the cross-provider spread, because the two answer different governance questions. **Caveats acknowledged on the methodology page.** Published rate-card pricing is a ceiling, not the actual contracted rate — large customers negotiate down 30-60% in committed-use discounts, and the Index does not attempt to estimate the discount curve because it varies with customer profile, term length, geographic commitment, and platform-engineering details that are not public. Spot and reserved pricing differ from on-demand by another factor of 2-5×. Cloud-egress and storage costs are not included. Vendor utilization disclosures are not standardized — some quote peak, some sustained, some "typical workload" with no method definition; the Index applies a uniform 0.50 utilization assumption to make values comparable across providers. **Region adjustments.** Pricing varies by geography for the same hardware tier — US-East tends to be lower than US-West, EU regions vary with carbon-electricity pricing, APAC regions reflect different commercial models, and export-restricted regions show distinctive pricing patterns. The Index surfaces per-region pricing without attempting to "normalize" out the geographic variance, because the variance IS the data — a 30% premium for compute in a particular jurisdiction is itself a governance-relevant signal. **Hardware-tier classification.** The Index splits the universe into roughly five tiers: - Tier S: GB200 NVL72 / GB300 NVL72 rack-scale systems (where pricing is even published; many configurations are committed-use-only) - Tier A: H100 SXM5 and equivalents (NVIDIA flagship through 2024; AMD MI300X) - Tier A-restricted: H800, H20, B30 (export-compliance variants — lower delivered FLOPS at chip-name parity) - Tier B: A100 and equivalents (the prior-generation flagship; still the dominant training rental tier in 2024-2025) - Tier C: V100, A10, T4 (older or specialized configurations) **Export-restricted-pricing caveat.** Pricing for export-restricted destinations (mainland China access via specific cloud regions, hardware variants designed for export-compliance such as H800 / H20 / B30) appears at the same hourly rates as unrestricted variants in vendor published documentation, but the actual delivered BF16 TFLOPS is lower than the export-permitted ceiling. The Index surfaces this as an asterisk on the relevant tier, not a corrected price — the published price is what the customer pays, even if the delivered capability is lower than the chip-name suggests. Where Scrutica can derive the delivered-FLOPS adjustment from primary sources (vendor-published export-compliance variant specifications, BIS-related published guidance), the methodology page renders both the published price and the implied effective-rate-per-actual-FLOP. **Cost trajectory over time.** Frontier training cost rises roughly 10,000× over the 2018-2025 frontier-training-runs trajectory per Epoch's published cost dataset; replication cost (training a model at the previously-frontier capability level) falls 2-18× over the same period due to algorithmic improvement, chip efficiency gains, and price reductions. The divergent curves are rendered at https://scrutica.com/training-cost as side-by-side trajectories with the methodology decomposition. **Spot vs reserved vs on-demand.** The Index focuses on on-demand pricing because it is the most widely published and the most useful baseline. Spot pricing is highly volatile and not consistently published across providers; reserved pricing is commitment-dependent. The methodology page documents the rationale and notes where spot or reserved pricing would shift the relative ranking among providers. --- ## Methodology — Chokepoint Cascade Simulation The Cascade Simulation propagates supply-chain disruption through Scrutica's bilateral-edge graph (21,037 raw edges; 15,049 canonical-deduplicated). The simulator is a weighted BFS over directed edges; weights combine three signals. **Signal 1: Supply share.** The disclosed fraction of the customer's input that the supplier provides (where disclosed in SEC filings, vendor press releases, or analyst reports — Tier 1-2 signal where present). For most edges in the AI compute supply chain, supply-share disclosures are partial — companies disclose top-10 customers in 10-K filings, name top-3 suppliers in earnings calls, and the disclosure quality varies dramatically by jurisdiction. Where supply share is disclosed, it is the strongest single signal; where it is not, the engine defaults to NULL and lets the other two signals carry the weight. **Signal 2: 3-month price correlation.** The rolling correlation between supplier and customer equity returns over 90 trading days. This proxies the markets' real-time assessment of dependency strength and is the most reliable cross-edge weighting signal available outside disclosure-based estimates — a supplier whose stock moves with its customer's stock has a structural relationship that the market is pricing in continuously. The 3-month window is short enough to be responsive to disruption events (e.g., the 2022 H100 announcement, the 2023 BIS Oct 7 rule, the 2024 HBM3e shortage signals) but long enough to filter out daily noise. The correlation signal has known limitations. Both supplier and customer can co-move with broad market beta — a high correlation between two AI-supply-chain stocks may reflect common factor exposure (interest rates, AI-sector sentiment) rather than specific dependency. The methodology page documents the decomposition (correlation residualized against sector index returns) but the default cascade weight uses the raw correlation because it is the most reliable signal where supply-share disclosures are absent. For Tier-1 edges where supply-share is disclosed, correlation is a secondary check. **Signal 3: Editorial criticality.** A 0.0-1.0 expert-assessed substitutability factor for the input class. Editorial criticality applies to inputs where substitution is structurally constrained: CoWoS-class advanced packaging (single-supplier from TSMC), EUV lithography (single-supplier from ASML), specific HBM SKUs (three-supplier oligopoly with multi-quarter qualification lead times), advanced photoresist materials (Japanese suppliers concentrated), specific deposition chemistries (per-tool-vendor specific). The methodology page documents which edge classes carry editorial weight and the assessment provenance — a published industry-research source or, where the assessment is Scrutica-editorial, the analyst's reasoning chain. **Pre-built scenarios.** Five named scenarios anchor the cascade page: - **TSMC disruption family** — Four sub-scenarios: blockade (sea-lane interdiction, fab-operations-continuing-but-shipping-blocked), targeted strike (fab-physical-damage, 12-24 month recovery), grid failure (Taiwan grid disruption, fab-operations-halted-temporarily), export restriction (US adds TSMC-specific advanced-node restriction). - **ASML export-halt** — EUV + DUV shipment freeze to specific jurisdictions; effects cascade through equipment-customer chain. - **US-China full decouple** — Maximal bilateral trade restriction; effects propagate through both the export side (NVIDIA, AMD, Intel cannot ship to China) and the import side (US cannot import Chinese-jurisdiction memory, packaging, materials). - **CoWoS bottleneck** — TSMC advanced packaging capacity-constrained; effects propagate to NVIDIA, AMD, and any other CoWoS-dependent customer. - **HBM supply disruption** — One of the three HBM suppliers offline; effects propagate through every AI-accelerator customer. Each scenario is parameterized by seed nodes (the entities directly affected) and a propagation decay rate (how much disruption attenuates per hop). The methodology explorer lets users sweep the decay rate from 0.30 to 0.95 and see how the affected-node count and maximum propagation depth respond. Users can adjust per-edge weights, swap the seed set, and replay the simulation. **Output schema.** The simulator returns the affected-node set, the maximum propagation depth reached, the per-node disruption magnitude, the per-edge propagation trail, and the by-tier breakdown of the affected nodes. The output is consumable in the Compare and Query interfaces and exportable as JSON / CSV with full provenance preserved. **Limitations.** Cascade analysis is structurally stronger on topology than on edge weights. Graph structure derives from primary-filing sources (SEC EDGAR, a licensed supply-chain database, CSET ETO — Tier 1-2 signal); substitutability decay rates are expert-assessed (Tier 3 signal). The methodology is useful for identifying structurally critical nodes (where graph degree and weighted centrality concentrate); it is less reliable for precise severity prediction at multiple hops, especially under simultaneous-disruption scenarios where the linear-propagation assumption breaks down. The methodology page is explicit about what cascade analysis is and is not. It is a tool for surfacing which nodes are structurally exposed under named disruption regimes — the analytical lens is "where does the graph topology indicate concentration risk" rather than "what would the dollar-cost impact of the scenario be." For dollar-cost impact analysis the analyst should pair cascade output with licensed-database quantitative impact estimates and external market models; Scrutica's cascade output is the topology-layer input to that joint analysis. --- ## Methodology — Supply Chain Interdependence The supply-chain graph contains approximately 21,037 unique bilateral edges between firms in the AI compute supply chain (15,049 after canonical deduplication on the supplier-customer-product key), derived from five primary sources: **Source 1: Licensed supply-chain database (Tier 1).** Firm-level supplier-customer disclosures from a licensed supply-chain relationships dataset, accessed via Harvard institutional credentials and held under subscription (not redistributed). The source covers approximately 600,000 supplier-customer relationships globally; Scrutica filters to the AI-compute subset by joining against the canonical organization table and the chip-vendor / data-center-operator / equipment-supplier classifications. The most recent extraction (5,403 new relationships in the May 2026 pull) awaits the canonical-org-resolver workstream to land them into the canonical `supply_chain_links` table with resolved org IDs rather than free-text `counterparty_name` fields. Per CLAUDE.md principle 11, this resolution work runs at write time via the alias-index resolver; the data has not yet been merged into the canonical table because composite-string sources require the `_split_composite_name()` tokenization audit per principle 12. **Source 2: Licensed corporate-ownership database (Tier 1).** Competitor relationships and industry-network data from a licensed corporate-ownership database, accessed via institutional credentials and held under subscription (not redistributed). It provides industry-classification context (sector codes) that complements the supply-chain source's direct-disclosure signal. **Source 3: SEC EDGAR Exhibit 21 (Tier 1).** Subsidiary disclosures from SEC 10-K filings for SEC-registered companies. Parsed via the official SEC EDGAR API; provenance per row preserved. Exhibit 21 captures parent-subsidiary relationships for US-registered entities but not the broader supplier-customer graph (which Exhibit 21 is not designed to disclose); used as a cross-check for ownership-chain integrity. **Source 4: CSET ETO (Tier 2).** Research-database aggregations from Georgetown's Center for Security and Emerging Technology's Emerging Technology Observatory. Used as a cross-check on the licensed-database-derived graph for the entity classes CSET ETO covers; preserves the per-source authority distinction so Scrutica's display can show "this edge is Tier 1 via licensed-database primary disclosure" vs "this edge is Tier 2 via CSET ETO aggregation." **Source 5: Editorial primary-source extraction.** For entities where the structured-source coverage is thin (privately held Chinese fabs, government-linked Russian / Iranian compute infrastructure, certain sovereign-AI vehicles), Scrutica's editorial process extracts supply-chain edges directly from primary sources: vendor press releases, earnings transcripts, government procurement disclosures, regulatory filings. Each editorial edge carries explicit provenance (source URL + extraction date + analyst initials) so the audit trail is preserved. **Per-edge attributes.** Each edge stores: - `source_org_id` — Canonical organization ID for the supplier (FK to `organizations`) - `target_org_id` — Canonical organization ID for the customer (FK to `organizations`) - `relationship_type` — One of: `supplier`, `customer`, `subsidiary`, `parent`, `competitor`, `JV-partner`, `co-investor`, `distributor`, `licensee`, `licensor`, `reseller` - `product_service` — Free-text product/service classification (e.g., "HBM3e memory", "EUV lithography systems", "CoWoS advanced packaging", "5nm wafer manufacturing", "GPU server rack design") — used in deduplication to distinguish multi-product relationships between the same supplier-customer pair - `supply_share` — Numeric fraction (0.0-1.0) of the customer's input that the supplier provides (NULL where not disclosed) - `price_correlation_3m` — Numeric correlation coefficient (-1.0 to 1.0) between supplier and customer equity returns over 90 trading days (NULL where one or both entities are not publicly traded) - `tier` — Integer tier classification: 1 = direct supplier-customer, 2 = supplier-of-supplier (2-hop relationship), 3 = 3-or-more hops - `editorial_criticality` — Numeric 0.0-1.0 expert-assessed substitutability factor (NULL for edges where editorial assessment has not been applied) - `data_source` — Source corpus identifier (e.g., `licensed-supply-chain`, `licensed-corporate-ownership`, `sec-edgar-exhibit-21`, `cset-eto`, `editorial-extraction`) - `source_url` — URL to the underlying record - `is_estimated` — Boolean; defaults to false for primary-source edges; true for editorial-criticality-derived weights - `data_vintage` — Date the underlying observation was produced **Canonical deduplication.** Multiple sources frequently disclose the same supplier-customer relationship from different angles (the licensed supply-chain database captures the relationship; SEC Exhibit 21 captures the parent-subsidiary chain; CSET ETO captures the academic-research aggregation). The canonical-dedup pass merges these into a single edge while preserving the per-source authority chain so the display can show "this edge is corroborated by 3 sources." **Coverage bias disclosure.** Coverage is structurally biased toward entities with public disclosure obligations: SEC-registered companies, entities covered by the licensed databases (which require either public listing or substantial private-market visibility), and the supplier classes that academic-research databases prioritize. State-owned enterprises, privately held companies (especially in China, Russia, Iran, and other jurisdictions with limited reporting), and the entities most relevant to compute governance are underrepresented relative to the public-disclosure majority of the graph. The methodology page documents per-source coverage rates and the limits each corpus carries. **Cascade simulation feedback loop.** The supply-chain graph is the substrate for the cascade simulation. Edge weights derived from supply share + price correlation + editorial criticality feed directly into the cascade propagation; per-edge confidence-tier classification flows through to the cascade output as a per-affected-node confidence rollup. The cascade methodology page documents the relationship between graph-edge-confidence and cascade-output-confidence. **Exports and downstream consumption.** The graph is queryable through the supply-chain explorer at `/supply-chain`, the cascade simulator at `/cascade`, the Query Builder NL interface at `/query`, and the MCP `scrutica_get_supply_chain` tool. Export is available as CSV + Croissant JSON-LD (with full provenance preserved) and as named subsets via the Datasets layer at `/datasets`. The full graph is scheduled for Zenodo publication with a citable DOI; see the dataset-packaging session prompt for the deposition roadmap. --- ## Methodology — Sovereign AI Programs The Sovereign AI dashboard tracks 36 national AI compute programs across 30+ countries with two distinct figures per program: - **Announced**: the headline number from official announcements, often conflating government-only commitments with public-private blended figures (e.g., the US CHIPS Act $52.7B headline conflates Section 48D ITC plus direct funding awards plus loan guarantees plus state-level matching, all of which have different fiscal mechanisms and different timelines). - **Deployed**: actual government-disbursed compute capacity as verified through primary-source procurement disclosures (USAspending for US, EU TED for European programs, UK Find a Tender for the UK AI Compute Programme, federal solicitations for individual jurisdictions where directly searchable, and ownership-chain traversal where the program flows through fund-of-funds vehicles). **Reality ratio.** `deployed_capacity / announced_capacity` per program. Programs with ratios below 0.3 indicate substantial announce-vs-execute divergence; programs with ratios above 0.8 indicate strong follow-through. The dashboard surfaces both values inline with the methodology behind each, so an analyst can see "this country announced $X billion but has actually disbursed $Y" with full provenance for both figures. **Inflation decomposition.** Where multi-year announcements span periods of significant currency drift or inflation (US dollar against various reference currencies 2020-2026; EUR-USD parity changes; KRW / JPY weakness against USD), the dashboard surfaces the inflation-decomposed real-terms figure alongside the headline number, marked explicitly so the reader can choose which framing applies to their question. The decomposition pipeline runs at substrate-refresh time; the methodology is documented at https://scrutica.com/methodology#sovereign-inflation-decomposition. The inflation decomposition is particularly load-bearing for headline numbers that are constructed from disparate funding components — e.g., the UAE program's $518B headline is the sum of (a) the Stargate UAE 5 GW campus aspiration ($500B aspirational), (b) Microsoft's $15.2B cumulative AI infrastructure pledge through 2029 (subsuming prior $1.5B Apr 2024 G42 equity), and (c) Abu Dhabi's $3.54B Government Digital Strategy 2025-2027 budget. Treating these as a single $518B announcement obscures their distinct funding mechanisms, timelines, and execution risks. The decomposition row-by-row makes each component visible. **Interdependence score.** Per-program weighted dependency on foreign suppliers across four dimensions: - `nvidia_dependency` — Reliance on NVIDIA accelerators (low / medium / high) - `tsmc_dependency` — Reliance on TSMC advanced-node fabrication (low / medium / high) - `us_dependency` — Reliance on US-jurisdiction entities including hyperscaler cloud (low / medium / high) - `bis_jurisdiction_reach` — Whether the program's compute substrate is reachable by US BIS export-control authority (low / medium / high) Higher interdependence indicates less sovereign-AI independence even with high announced figures — a $50B sovereign-AI program whose substrate is entirely NVIDIA accelerators sitting in AWS / Azure / GCP regions has high BIS-jurisdiction reach and is not sovereign in the policy-meaningful sense. **Program inclusion threshold.** $100M government-only commitment OR named in (CRS public reports / NVIDIA earnings sovereign-AI commentary / IDC sovereign-AI tracking / CNAS Sovereign AI Index). The threshold is documented alongside the program count. Sub-threshold national initiatives are tracked in the methodology page narrative but not surfaced as separate program rows. **Governance reach.** A per-program score combining capacity-weighted control rights (where the sovereign program retains direct compute allocation authority — e.g., a national supercomputing facility vs cloud-credit subsidies to private firms), legal authority for export-control coordination with the program's chip suppliers (e.g., the UAE program's 2024 G42 divestment of Chinese tech holdings as a condition of Microsoft investment), and the program's institutional integration with cross-jurisdictional governance fora (Bletchley AI Safety Summit, GPAI, etc.). Computed per https://scrutica.com/methodology#sovereign-execution-classification. **Per-program data quality flags.** Many programs carry explicit data quality flags surfacing known accounting ambiguities. The flag categories include: - `CONFLATION WARNING` — The headline number combines disparate funding mechanisms (government direct + private capex + tax incentives + loan guarantees) in a way that the announcement does not disaggregate - `UNVERIFIED` — The figure is widely cited but Scrutica's editorial process could not corroborate against a primary source (typically a press claim attributed to "industry sources" that Scrutica could not trace to an actual document) - `VERIFIED` — Cross-checked against a primary source explicitly cited in the program's `sources` array - `AVOID-DOUBLE-COUNT` — A particular sub-figure is a subset of a larger announcement and should not be added separately - `GPU COUNT GAP` — Known absence of public GPU-count disclosure for a specific operator or sub-program The flags are rendered inline on each program detail page with the originating source citation. **Coverage limitations.** Sovereign AI program reconciliation depends on the quality of public procurement disclosures. Programs in jurisdictions with limited procurement transparency (or programs that route through privately-held vehicles, fund-of-funds, or holding companies) will undercount their actual disbursement. China's Big Fund III, Russia's national AI program, and Iran's compute infrastructure are particularly undercovered relative to their actual size because the underlying disclosure quality is weak. The methodology page documents per-program reconciliation source coverage so the reader can see "this program's reality ratio is computed against N primary sources covering X% of the announced commitment." --- ## Methodology — Threshold Atlas The Threshold Atlas surfaces which compute facilities have the FLOP capacity to train models above named regulatory triggers: **EU AI Act 10²⁵ FLOPs trigger.** Active as of the AI Act's general-purpose AI obligations. The July 2025 GPAI Guidelines provide facility-level interpretation guidance — what counts toward the threshold, how multi-facility training runs are aggregated, what the look-back window is for re-classification. **China CAC ~10²⁴ FLOPs.** The ~10²⁴ figure is analyst-reported from CAC internal guidance (authority tier 2), not codified in published regulatory text — the Generative AI Interim Measures state a qualitative scope trigger ("public opinion properties or capacity for social mobilization"), and China's regulator publishes registration outcomes rather than a bright-line FLOP threshold. The Threshold Atlas renders the figure with that tier-2 attribution attached. **Historical US EO 14110 10²⁶ FLOPs.** The dual-use foundation model reporting trigger introduced by President Biden's October 2023 Executive Order. Rescinded by President Trump's January 2025 Executive Order 14148 (NOTE: not EO 14179, which is a separate Trump-era order on a different subject — the correction was applied per CLAUDE.md principle 16 via the editorial_name_corrections table). Retained in the Atlas as a calibration reference point because the 10²⁶ figure remains the most widely-cited frontier-threshold anchor in academic and policy literature. **Per-facility threshold computation.** For each facility with sufficient hardware-inventory or power-envelope data, the Atlas computes the time-to-threshold under specified training-run assumptions: - Model architecture (default: dense transformer; alternative: sparse mixture-of-experts with effective-active-parameter fraction) - Sparsity (default: dense BF16 — per FLOP Engine methodology) - Precision (default: BF16; alternatives: FP8, FP16, INT8 for inference-focused workloads) - Training-token volume (default: Chinchilla-optimal scaling; user-adjustable) - Compute utilization (default: 0.50 sustained; range 0.30-0.70) The output is rendered as days-to-threshold per facility per training-run assumption set. A facility with 10 EFLOPS dense BF16 capacity can train a 10²⁵ FLOP model in 10⁵ seconds ÷ 86,400 ≈ 1.16 days of dedicated capacity; the same facility can train a 10²⁶ FLOP model in ~11.6 days. The Atlas surfaces these as a per-facility threshold-crossing matrix. **Reverse calculator.** Users can input a target FLOP budget and back-solve for required facility characteristics — minimum power capacity, minimum chip count, hardware tier required, expected training duration. The reverse-calculator output supports questions like "what facility characteristics would be required to train a 5 × 10²⁵ FLOP model in 30 days under H100-class hardware assumptions?" **SemiAnalysis 100K H100 calibration anchor.** The Atlas uses SemiAnalysis's published decomposition of the 100K-H100 cluster training capability (published 2024) as a cross-validation point. The methodology page renders the Scrutica-computed threshold-crossing figures for the equivalent hypothetical facility alongside the SemiAnalysis published estimates so the calibration residual is visible. **Coverage.** Approximately 110 countries with sufficient facility characterization to compute a threshold-crossing assessment. The Atlas surfaces 6-peer comparison panels per country (with the disclosed-comparable peer set defined per https://scrutica.com/methodology#peer-comparison) and a 110-country bar-chart overview ranked by facility-class threshold-crossing capacity. **Limitations.** The Atlas reasons about capacity, not about regulatory enforcement. A facility's threshold-crossing capacity does not mean a model trained at that facility would be subject to the named regulation — the regulation's scope depends on the entity training the model, the entity's jurisdiction, the model's deployment scope, and other factors outside the Atlas's analytical lens. The Atlas is most useful as a "where could this be trained" map rather than a "where will this be regulated" map. --- ## Methodology — Compute Visibility Index The Compute Visibility Index measures, per country, what fraction of in-country compute capacity is attributable to a public-filing-grade ultimate beneficial owner. The Index has four tiers: - **Tier 1**: ultimate beneficial owner identified through SEC, FCA, or equivalent regulatory filings (typically US, UK, EU listed entities under quarterly-reporting obligations) - **Tier 2**: ultimate beneficial owner identified through corporate-registry filings (Companies House for the UK, EU Business Registers for member states, US state-level Secretary of State filings for non-public US entities) - **Tier 3**: ownership documented through licensed analyst databases (held under subscription; not redistributed) — Tier-3 sources are reliable but secondary to the originating filing - **Tier 4**: ownership opaque — terminus chain at a state-owned enterprise of record, private trust, holding-company shell, or undocumented opaque parent The Index is capacity-weighted: a 100 MW facility under Tier 1 transparency counts more than a 10 MW facility under Tier 4 opacity. The metric is computed per-country and aggregated globally; the latest values are at https://scrutica.com/compute-visibility. **Methodology version.** The Visibility Index is at v1.2 as of the last methodology refresh; the version history documents the per-version changes (e.g., v1.1 narrowed the Tier-2 definition to exclude jurisdictions whose corporate-registry filings are filed-but-not-verified; v1.2 added the bilateral hosted-vs-controlled lens as a complementary metric). **Hosted vs Controlled.** A second lens on attribution: where compute physically sits (hosted) versus which jurisdiction's entity has operational control (controlled). The two diverge when a US-headquartered hyperscaler operates a facility in a foreign jurisdiction with a foreign operating subsidiary; the physical hosting attributes to the foreign country but the operational control attributes to the US parent. The lens is particularly informative for: - US hyperscalers operating EU / APAC regions (hosted in EU/APAC, controlled by US) - Sovereign AI programs hosting NVIDIA accelerators (hosted in sovereign jurisdiction, controlled by sovereign program but supply-chain-exposed to US BIS authority) - Chinese hyperscalers operating overseas regions (hosted overseas, controlled by Chinese entity subject to Chinese cybersecurity law) The hosted-vs-controlled rendering at /compute-visibility shows the bilateral matrix: rows are physical-hosting countries, columns are controlling-jurisdiction countries, cell values are megawatt-weighted facility counts. The diagonal (own-country hosted and controlled) is the sovereign substrate; off-diagonal entries are the international flows. **Coordination Gap detection.** The Index feeds the Coordination Gap Analyzer at /coordination-gaps. A facility under coordination arbitrage sits in an allied export-control regime but reaches a non-coordinated jurisdiction through the ownership chain — e.g., a Tier-1-transparent Singapore facility under a holding company whose ultimate beneficial owner is in a non-coordinated jurisdiction. The Analyzer surfaces these as a structurally-arbitrage-exposed surface. **Limitations.** The Index measures transparency, not actual control. A Tier-1-transparent facility under a holding company with non-coordinated jurisdiction reach is governance-arbitrage exposed even with maximum transparency; the Coordination Gap Analyzer is the complementary surface for surfacing these cases. The Index is also reactive to the underlying filing quality — improvements in transparency-disclosure regimes (e.g., the EU's Whistleblower Directive operationalization, the OECD beneficial-ownership-transparency initiatives) shift the Tier distribution over time. The methodology page documents the historical Tier distribution per country so the cross-time-period analysis is supported. --- ## Compute Leverage — who can stop a frontier AI run The instrument at https://scrutica.com/compute-leverage is a veto count over the frontier-AI compute stack: per leverage point, who holds it, whether the holder can veto a training run at an 18-month and a 5-year horizon, and the accountability that checks them. Leverage is defined as switching-cost × redeployment-time, NOT market share — a holder can halt a run when the lab cannot route around them within the horizon, regardless of what share of the market the holder books. Market share renders as context only, never as the verdict input. **Veto semantics.** A veto verdict (YES / NO / UNCERTAIN) is assigned per leverage point × horizon: YES when a dominant or sole actor — or a tightly-coupled institutional stack acting as a single gate — exists AND switching-cost × redeployment-time exceeds the horizon; NO when the lab can route around the holder in time; UNCERTAIN when the routing time straddles the horizon. The two horizons ask different questions: 18 months is the window to field a new frontier run; 5 years is structural independence. In aggregate the 18-month veto count is the number of gates to STAND UP the next run, not to stop one already training — halt-capable gates (power, the data center) clear the moment a run is energized. A closing window (HBM tri-sourcing maturing) reads YES→NO across the horizons; a permanent lock (EUV lithography) reads YES→YES. Verdicts re-resolve per named frontier run via per-run accelerator and site overrides — a vertically-integrated lab's run clears the accelerator gate that a merchant-GPU run does not — so no verdict is a universal constant; the published defaults are the cross-run committed-lab case. **Accountability tags.** Each veto-holder carries a categorical tag with a stated, auditable rule — never a blended score. The tag is the strongest governance hook over the layer under a fixed priority order: EXPORT_CONTROL (a BIS / Dutch / multilateral export-control regime governs the layer — directional: it governs diversion, e.g. to China, not a domestic lab's purchase) > ANTITRUST (a named competition authority with a live or available proceeding) > PUBLIC (the holder is itself a public body, democratically / FOIA accountable) > COMPETITION (no legal or public hook, but an alternative supplier exists in principle) > NONE (no governance hook of any kind). In-horizon reachability lives in the veto, not the tag: that a competitor exists but is unreachable in 18 months is a VETO=YES with a COMPETITION hook, not an absence of accountability. **The ledger** (11 leverage points, materials to capital; verdicts are the cross-run defaults): - **EUV lithography** — holder(s): ASML. Veto at 18 months: YES; at 5 years: YES. Accountability: EXPORT_CONTROL. - **Advanced logic foundry (≤5nm/3nm)** — holder(s): TSMC, Samsung Foundry, Intel Foundry. Veto at 18 months: YES; at 5 years: UNCERTAIN. Accountability: EXPORT_CONTROL. - **EDA tools (chip design)** — holder(s): Synopsys, Cadence, Siemens EDA. Veto at 18 months: YES; at 5 years: YES. Accountability: EXPORT_CONTROL. - **HBM memory (HBM3E / HBM4)** — holder(s): SK Hynix, Samsung, Micron. Veto at 18 months: YES; at 5 years: NO. Accountability: EXPORT_CONTROL. - **Advanced packaging / CoWoS** — holder(s): TSMC, OSATs (ASE, Amkor, SPIL). Veto at 18 months: YES; at 5 years: UNCERTAIN. Accountability: NONE. - **AI accelerator** — holder(s): NVIDIA, Google TPU, AWS Trainium, AMD. Veto at 18 months: YES; at 5 years: UNCERTAIN. Accountability: EXPORT_CONTROL. - **Scale-up interconnect (NVLink / NVSwitch)** — holder(s): NVIDIA. Veto at 18 months: YES; at 5 years: UNCERTAIN. Accountability: EXPORT_CONTROL. - **Scale-out interconnect (InfiniBand vs Ethernet)** — holder(s): NVIDIA/Mellanox InfiniBand, Broadcom/Arista Ethernet. Veto at 18 months: NO; at 5 years: NO. Accountability: COMPETITION. - **Cloud / datacenter capacity** — holder(s): AWS, Microsoft Azure, Google Cloud, CoreWeave, Oracle, neoclouds. Veto at 18 months: UNCERTAIN; at 5 years: NO. Accountability: COMPETITION. - **Power / grid interconnection** — holder(s): ISO/RTO (MISO, PJM, ERCOT, SPP), local utility, state PUC, FERC, EPA (institutional stack — counted as ONE gate, flagged as a different kind of concentration than a private firm). Veto at 18 months: YES; at 5 years: UNCERTAIN. Accountability: PUBLIC. - **Capital ($1B+/yr runs)** — holder(s): hyperscaler balance sheets, sovereign funds (QIA, PIF, Mubadala), mega-VC. Veto at 18 months: UNCERTAIN; at 5 years: NO. Accountability: COMPETITION. **The headline finding** is the intersection veto=YES ∩ accountability=NONE — concentrated control that answers to no one. At the 18-month horizon, 8 of the 11 leverage points read veto=YES; the intersection set is Advanced packaging / CoWoS — the sole layer that can halt a run and answers to no governance hook of any kind. **Provenance.** Every coefficient (switching cost, redeployment time, share context) is a sourced record carrying authority tier (1 primary … 4 estimated/inferred), vintage, and an is_estimated flag; the two genuinely unsourced Tier-4 coefficients carry a visible needs-verification marker, and sourced-but-estimated values carry a quiet "est." marker. The interactive page exposes the full methodology panel (the leverage definition and every coefficient with its source), a per-run selector, the concentrated-×-unaccountable matrix, and a CSV / JSON export of the ledger with authority tiers preserved. --- ## Methodology — deep-dive worked examples This section pairs the methodology surfaces above with worked examples grounded in the canonical-facility substrate. Each example is a complete derivation chain: the underlying primary-source disclosures, the assumption set, the intermediate values, and the final published estimate. The intent is to make every Scrutica-derived number reproducible by a careful analyst without having to walk the methodology pages line-by-line — the worked examples below are the calibration anchors. ### Worked example 1: Microsoft Fairwater Wisconsin FLOP-engine derivation (Hardware Path) The facility detail page at https://scrutica.com/facilities/fac-epoch-microsoft-fairwater-wisconsin carries the Epoch-derived disclosure: 555 MW power capacity, two-story building structure with GPU racks networked vertically, paired with a one-story CPU/storage building, ~558,827 H100-equivalent GPUs (Epoch estimate, derived from permit applications and rack-layout disclosure). The hardware-path FLOP-engine derivation applies the H100 SXM5 vendor specification (989.5 BF16 TFLOPS dense), the 0.50 sustained-utilization default per the methodology page, and the dense-BF16 sparsity factor (1.0× — Scrutica's editorial default per the cross-path-divergence narrative). The intermediate value: 558,827 × 989.5 TFLOPS × 0.50 = 276,481,159 TFLOPS sustained = 276,481 PFLOPS sustained ≈ 0.28 ZFLOPS. The capacity bound on a single training run: 276,481 PFLOPS × 86,400 seconds/day = 23,887,968,576 PFLOP-day per day of dedicated training, or roughly 2.4 × 10¹⁹ FLOPs/day. A 10²⁵-FLOP training run at this facility — under fully-dedicated-capacity assumption — completes in approximately 0.42 days. A 10²⁶-FLOP training run completes in approximately 4.2 days. The published Threshold Atlas value carries the same arithmetic with the assumption-adjustable parameters exposed in browser; an analyst who prefers 0.40 utilization rather than 0.50 sees the threshold-crossing duration scale by 1.25×. ### Worked example 2: Anthropic-Amazon New Carlisle (Project Rainier) Power-Path bound The facility detail page at https://scrutica.com/facilities/fac-epoch-anthropic-amazon-new-carlisle carries the Epoch-derived disclosure: 1,092 MW total power capacity (the largest single-site disclosed compute in the Scrutica substrate as of the May 2026 snapshot), ~685,914 H100-equivalent GPUs (Epoch estimate), direct-air cooling, $34.86B total capital cost (2025 USD). The Power-Path bound applies the chip-family-specific FLOPS-per-watt curve: at 60% IT-load utilization (a representative assumption for direct-air-cooled facilities under sustained training workloads), the 1,092 MW envelope yields approximately 655 MW of compute-attributable power. Under H100 SXM5 FLOPS-per-watt assumption (~1.4 TFLOPS/W dense BF16 sustained at the systems level, accounting for memory and network overhead), the Power-Path FLOP capacity is approximately 916 PFLOPS sustained. The Hardware-Path estimate for the same facility (Epoch's 685,914 H100-equivalent × 989.5 TFLOPS × 0.50 utilization) yields approximately 339 PFLOPS sustained. The cross-path divergence here is approximately 2.7× — above the 2.0× contradiction threshold — and surfaces on the facility intelligence card as a flag with the editorial disposition: the Hardware-Path estimate is anchored on the higher-confidence Epoch GPU-count disclosure (Tier 2) and is the preferred display, while the Power-Path bound is preserved as an upper-envelope cross-check. The divergence reflects that the 1,092 MW figure includes substantial non-compute power overhead (cooling, networking, storage, redundancy) that the Power-Path bound's simple FLOPS-per-watt application does not fully capture. The methodology page documents the per-facility power-overhead allocation pattern that would close the gap if applied consistently across the substrate. ### Worked example 3: OpenAI Stargate Abilene cross-validation against SemiAnalysis The facility detail page at https://scrutica.com/facilities/fac-epoch-openai-stargate-abilene carries the Epoch-derived disclosure: 590 MW power capacity, ~510,358 H100-equivalent GPUs, $16.14B total capital cost (2025 USD), satellite-imagery-validated air-cooled chiller layout, the named first site in the broader Stargate Project. The Hardware-Path estimate: 510,358 × 989.5 TFLOPS × 0.50 = 252,500 PFLOPS sustained. SemiAnalysis's published 100K-H100 cluster decomposition (the 2024 reference paper Scrutica cross-validates against) implies approximately 49,500 PFLOPS sustained per 100K H100 SXM5 cluster under the same utilization assumption — multiplied by Stargate Abilene's 5.1× scaling factor: 49,500 × 5.10 = 252,450 PFLOPS sustained. The two estimates agree to within 0.02% (residual is rounding-driven). The cross-validation pass at https://scrutica.com/methodology#external-validation publishes the full per-facility cross-comparison residual table; Stargate Abilene's near-zero residual is the methodology's strongest external calibration anchor and motivates the published default utilization of 0.50 (alternative utilization assumptions that would invalidate the cross-validation fall outside the published 0.30-0.70 range). ### Worked example 4: Cascade simulation seed-to-output trace (TSMC blockade sub-scenario) The TSMC blockade sub-scenario at https://scrutica.com/scenario-analysis/taiwan-strait seeds on TSMC's direct customers in the supply-chain graph. Seed set: NVIDIA (`org-nvidia`), AMD (`org-amd`), Apple (`org-apple`), Qualcomm (`org-qualcomm`), MediaTek (`org-mediatek`), Broadcom (`org-broadcom`), Marvell (`org-marvell`), Cerebras (`org-cerebras`), Tenstorrent (`org-tenstorrent`), and the top-N customers ranked by 3-month price correlation with TSMC over the past 90 trading days. The first-hop cascade weight for each seed: `(supply_share where disclosed) × (price_correlation if no supply_share) × (editorial_criticality)`. For NVIDIA the editorial-criticality multiplier is 1.0 (CoWoS-class advanced packaging is single-supplier from TSMC — full substitution-impossible weighting). For Apple the editorial criticality is 0.8 (Apple's M-series and A-series fabrication is TSMC-exclusive at advanced nodes; partial substitution available at mature nodes via secondary supplier). For Qualcomm the editorial criticality is 0.7 (Snapdragon flagship fabrication is TSMC; Samsung Foundry has historically taken portions of mid-tier Snapdragon SKUs). The cascade BFS propagates from these first-hop nodes outward through their own customer edges; the second-hop set includes NVIDIA's hyperscaler customers (Microsoft, Amazon, Alphabet, Meta, Oracle, CoreWeave, OpenAI, xAI, Mistral) and Apple's downstream-product distribution; the third-hop set includes the cloud regions' end-customer fan-out. The cascade output renders the affected-node count, the maximum propagation depth reached, the per-node disruption magnitude, and the per-edge propagation trail; the worked-example values are reproducible by a careful analyst against the published scenario seed set and per-edge weights. ### Worked example 5: Sovereign-AI reality-ratio derivation (UAE composite headline) The UAE program's $518.74B headline is the sum of three distinct funding components per the inflation decomposition: $500B Stargate UAE 5 GW campus aspirational (FDI character, multi-year construction window, NVIDIA / Oracle / OpenAI / Cisco joint structure), $15.2B Microsoft cumulative AI infrastructure pledge through 2029 (subsumes the prior $1.5B April 2024 G42 equity injection plus subsequent datacenter capex and local operating commitments), and $3.54B Abu Dhabi Government Digital Strategy 2025-2027 (governmental direct funding, including the Oracle Sovereign AI Supercluster with 4,000+ NVIDIA Blackwell GPUs in the OCI Abu Dhabi region). The headline-vs-government-only split reflects that the $518B figure is overwhelmingly private FDI character; the government-only commitment is $3.54B. The committed-vs-announced ratio: $18.74B committed / $518.74B announced = 0.036. Treating the $518B as a single reality-ratio numerator would substantially overstate UAE's government-execution velocity; treating only the $3.54B government-only figure relative to the same denominator understates UAE's compute substrate (which actually has 4,000+ Blackwell GPUs operational at OCI Abu Dhabi as of November 2025). The dual-figure rendering on the sovereign-AI dashboard at https://scrutica.com/sovereign-ai shows both the headline and the government-only with the CONFLATION WARNING data quality flag inline, surfacing the disaggregation rather than picking a "winner" framing. ### Worked example 6: Compute Cost Index $/PFLOP-day derivation (AWS p5.48xlarge vs CoreWeave HGX H100) The AWS p5.48xlarge instance (8× NVIDIA H100 SXM5) lists at approximately $98.32/hour on-demand in US-East-1 (rate-card pull, vintage May 2026). The CoreWeave HGX H100 instance (8× NVIDIA H100 SXM5) lists at approximately $32.00/hour reserved-1-year (CoreWeave published rate, vintage May 2026; on-demand is higher and varies by capacity-tier). The SCU conversion: 8 × 989.5 TFLOPS × 0.50 utilization = 3,958 sustained TFLOPS = 3.958 PFLOPS per instance. The $/PFLOP-day derivation: AWS on-demand: (98.32 × 24) / 3.958 = $596 per PFLOP-day. CoreWeave reserved: (32.00 × 24) / 3.958 = $194 per PFLOP-day. The cross-provider spread for equivalent H100 SXM5 capacity is approximately 3.1× — substantially larger than the 1.5× that would be expected if the market were efficiently arbitraged. The spread reflects committed-use commitment-length variance (CoreWeave's reserved rates require multi-year commitments; AWS on-demand has zero commitment), capacity-tier allocation differences (CoreWeave's neocloud business model concentrates inventory in named training-customer commitments while AWS preserves on-demand inventory for spot variability), and the underlying economic difference between hyperscaler vs neocloud capacity-economics. The Cost Index renders the spread inline with the methodology page documenting which procurement-model and commitment-length assumptions produce which point estimate. --- ## Entity taxonomy Scrutica's entity model uses TEXT primary keys with meaningful, multi-source-stable identifiers. The pattern is deliberate — opaque autogenerated keys would break the cross-source provenance chain that Scrutica's data discipline depends on. Where an entity has a canonical identifier in a primary source (e.g., `epoch-dc-001` from Epoch's Frontier Data Centers, `osm-w341892` from OpenStreetMap), Scrutica preserves that identifier in the primary-key column. Where multiple sources reference the same entity under different identifiers, the canonical alias-resolution table maps variants onto a single canonical ID. **Facility** (`facilities` table; 4,550 rows). Primary key examples: `epoch-dc-001` (Epoch corpus identifier), `osm-w341892` (OSM way ID for OSM-derived facilities), `fac-pdb-1234` (PeeringDB facility ID), `sov-humain-jeddah-1` (sovereign-program-linked facility identifier). Each facility carries: - `name` — Display name, with the canonical form preserved - `facility_type` — One of: `data_center`, `fabrication_plant`, `packaging_plant`, `lithography_facility`, `materials_plant`, `assembly_plant`, `research_facility` - `country` — ISO-2 country code - `region` — Sub-national administrative region (state, prefecture, province) - `city` — City name - `operator_org_id` — FK to `organizations`; the entity operationally running the facility - `owner_org_id` — FK to `organizations`; the entity legally owning the facility (frequently distinct from operator) - `power_capacity_mw` — Numeric; nameplate power capacity in megawatts (NULL where not disclosed) - `gpu_count` — Integer; disclosed GPU count (NULL where not disclosed) - `gpu_model_mix` — JSONB; per-model breakdown where disclosed - `location` — PostGIS geography(Point, 4326) with WGS84 CRS and GiST index; the spatial primary key for map rendering - `status` — One of: `operational`, `under_construction`, `announced`, `decommissioned`, `closed` - `commissioning_year` — Integer year the facility came online (NULL for non-operational facilities) - `data_source` — Identifier of the originating corpus - `source_url` — Direct URL to the underlying record - `is_estimated` — Boolean - `authority_tier` — Integer 1-4 - `data_vintage` — Date the underlying observation was produced - `updated_at` — Timestamp of the row's last refresh Detail page at https://scrutica.com/facilities/{id}. **Organization** (`organizations` table; 129,061 canonical rows, of which 24,873 are compute-relevant by joining against the facility / supply-chain / BIS / investment relationships). Primary key examples: `org-nvidia`, `org-tsmc`, `org-huawei`, `bis-entity-huawei` (BIS-only designations where the BIS entity ID is distinct from the canonical org), `licensed-comp-12141` (licensed-database variant identifier; resolves to canonical via `canonical_org_map`). Each organization carries: - `name` — Display name in the canonical form - `legal_name` — Full legal entity name (frequently longer than the display name) - `country_hq` — Country of headquarters (ISO-2 code) - `org_type` — One of: `public_company`, `private_company`, `state_owned_enterprise`, `holding_company`, `government_organization`, `research_organization`, `fund`, `fund_of_funds`, `special_purpose_vehicle`, `industry_consortium` - `parent_org_id` — FK to `organizations`; the immediate parent (NULL for ultimate-parent rows) - `ticker` — Public-market ticker (NULL for non-listed entities) - `lei` — Legal Entity Identifier (the ISO 17442 standard) - `same_as_uris` — JSONB; URIs to canonical external identifiers (Wikidata Q-IDs, ROR identifiers, Wikipedia URLs, official website URLs) — populated by the L2 jsonld-entity-schemas session - `founding_year` — Integer year of founding - `employees` — Integer headcount (NULL where not disclosed; vintage-tagged where disclosed) - `market_cap_usd` — Numeric USD market capitalization (NULL for non-public entities) - `is_estimated` — Boolean Aliases stored in `organization_aliases`: alias rows resolving variant spellings (e.g., "NVIDIA Corp" / "NVIDIA Corporation" / "Nvidia"), transliterations (e.g., "Hua Hong" / "华虹半导体" / "Huahong Semiconductor"), and acquisition-related name changes (e.g., "ARM Holdings" / "Arm Ltd"). Adding an alias is a SQL insert per principle 10. **Supply Chain Link** (`supply_chain_links` table; 21,037 rows). Foreign keys: `source_org_id` + `target_org_id` (both reference `organizations`). Per-edge attributes as documented in the supply-chain methodology section above. **Sovereign AI Program** (`sovereign_programs` table; 36 rows). Per-country program with the columns documented in the sovereign methodology section above. **Export Control Designation** (`export_control_designations` table; 3,422 rows). Each designation: - `entity_name` — Verbatim name from the Federal Register / OFAC SDN / EU dual-use / Japanese end-use designation - `designation_date` — Date of the originating publication - `jurisdiction` — One of: `us-bis-entity-list`, `us-bis-mil-end-user`, `us-ofac-sdn`, `us-treasury-ccp-mil-co`, `netherlands-export-license`, `eu-dual-use`, `japanese-end-use-restriction`, `uk-strategic-export` - `control_reason` — Free-text reason from the originating regulation - `source_url` — Federal Register URL (or equivalent for non-US authorities) - `control_type` — One of: `entity_list`, `sdn`, `ccl`, `mil_end_user`, `unverified_list` - `authority_tier` — Always 1 (these are Federal-Register-anchored / equivalent primary-source designations) **BIS Entity List Cross-Reference** (`bis_crossref_matches` table). Each row matches a Scrutica organization to a BIS Entity List designation with: - `scrutica_org_id` — FK to `organizations` - `designation_id` — FK to `export_control_designations` - `confidence_score` — Numeric 0.0-1.0 match confidence - `match_method` — One of: `exact_name`, `alias`, `parent_match`, `fuzzy_match` - `editorial_disposition` — Where applicable, an editorial note explaining the match **Scenario** (5 pre-built scenarios; documented at `/scenario-analysis`). Each: seed nodes, propagation parameters, affected-facility list, recovery timeline, scenario assumptions, data sources. **Chip / Hardware SKU** (`hardware_catalog` table). Each chip: - `name` — Canonical SKU name (e.g., `H100-SXM5`, `B200`, `MI300X`) - `vendor_org_id` — FK to `organizations` - `architecture_family` — e.g., `Hopper`, `Blackwell`, `CDNA3`, `CDNA4` - `peak_bf16_tflops` — Numeric; dense BF16 throughput - `peak_fp8_tflops` — Numeric; dense FP8 throughput (for inference-optimized variants) - `peak_int8_tops` — Numeric; INT8 inference throughput - `hbm_capacity_gb` — Integer HBM capacity per chip - `hbm_bandwidth_gbs` — Numeric HBM bandwidth per chip - `packaging_type` — One of: `CoWoS-S`, `CoWoS-L`, `InFO`, `SoIC`, `other` - `export_control_class` — One of: `unrestricted`, `3A090`, `4A090`, `export_compliant_variant` (H800, H20, B30, etc.) **Compute Deployment** (`compute_deployments` table; 4,317 rows). Each deployment: facility ID, chip SKU, count, disclosed-vs-estimated, source, vintage. The deployments table is append-only so the time-series view of capacity by chip class by facility is preserved. **Investment** (`investments` table; 33,627 rows). Each: round date, recipient org, lead investor org, co-investor orgs, valuation, raised amount, round type, source. Used for the compute-flow analysis under the Diversion Pipeline Tracker and the Sovereign AI program reconciliation. --- ## Provenance system The provenance system is implemented through five mechanisms: **Mechanism 1: Per-row `data_source` + `source_url`.** Every record from every primary source preserves the URL of the canonical filing / disclosure / measurement. The BIS Entity List rows preserve the Federal Register URL anchored at the section identifying the entity; the Epoch-derived facility rows preserve the Epoch corpus URL (e.g., `https://epoch.ai/data/data_centers#facility-id`); licensed-database-derived edges preserve the licensed record reference only where redistributable (the underlying records are held under subscription and not redistributed) and the parent SEC filing where available. This per-row provenance is the foundation of Scrutica's audit chain — any number on any page can be traced back to its primary measurement by following the `source_url`. **Mechanism 2: Per-row `is_estimated` flag.** Every derived or estimated value flags the boolean explicitly. Where the entire table is primary-source by construction, the boolean is suppressed and replaced by a column-specific confidence metric: `bis_crossref_matches.confidence_score`, `organizations.is_estimated` is restricted to specific columns rather than blanket-applied, `investments.raised_amount_is_estimated` per-investment. The suppression is documented in the data-access module's header comment so the absence of the boolean is not a hidden footgun. **Mechanism 3: Data Quality Flags** (`organization_data_quality_flags`). Cross-source contradictions, vintage anomalies, source-reach limits, and editorial dispositions land as flag rows with: - `organization_id` — The entity the flag applies to - `flag_type` — One of: `cross_source_divergence`, `vintage_anomaly`, `source_reach_limit`, `editorial_disposition`, `unverified_claim`, `coverage_gap`, `conflation_warning`, `avoid_double_count` - `flag_data` — JSONB payload with the per-flag specifics (the divergent values, the source citations, the editorial reasoning) - `resolution_disposition` — How the flag is resolved (which value the platform prefers and why, or "unresolved-known-residue" where editorial dispositioning has not been applied) - `created_at` / `resolved_at` — Timestamps Surfaced inline on the company and facility detail pages via the `` component. The flags are rendered to the user with the contradicting sources visible — silently winner-picking would violate principle 7. **Mechanism 4: Editorial name corrections** (`editorial_name_corrections`). Name-typo, transliteration, and acquisition-driven name-variant corrections land as new rows with: - `source_id` — Identifier of the originating record (SEC filing accession number, BIS Federal Register designation ID, etc.) - `verbatim_form` — The exact form as it appears in the primary source - `corrected_form` — The display form Scrutica renders - `source_citation` — URL to the originating record - `effective_date` — When the correction was first applied - `correction_rationale` — Free-text explanation The substrate never silently mutates the primary-source value; corrections surface at display time per principle 16. The verbatim form remains queryable so an analyst can verify the originating source still contains what Scrutica says it contains. **Mechanism 5: PROV-O `wasDerivedFrom` JSON-LD.** Dataset-level provenance using the W3C PROV-O vocabulary on the Dataset JSON-LD emission. The relevant predicates: - `prov:wasDerivedFrom` — Links the Scrutica dataset to its primary source corpora - `prov:qualifiedDerivation` — Adds the derivation activity (the Scrutica processing pipeline) as an explicit linked entity - `prov:wasGeneratedBy` — Names the generation activity (e.g., the FLOP Estimation Engine version 2.1) - `prov:wasAttributedTo` — Names the responsible agent (Scrutica + the editorial analyst) Surfaced per-page through the `` component. The PROV-O emission is rendered as part of the `Dataset` schema.org type so it appears in Google Dataset Search and equivalent academic-dataset indexing layers. **Visibility to agents.** All five mechanisms are visible to LLM agents fetching Scrutica pages. The structured-data emission (JSON-LD via `` plus per-entity `` blocks) makes the provenance machine-readable; the prose emission (inline citation tags, "Cite This" panels, source-url-attribution columns in tables) makes it human-readable. The MCP server (`/api/mcp`) exposes the same provenance through the tool-response payloads — every entity returned by an MCP tool call includes the source-url and authority-tier fields so the agent can present the citation chain to the user. --- ## Authority tier definitions Every claim in Scrutica is tagged with an authority tier: - **Tier 1**: Primary measurement / filing. Direct disclosure from the entity in a regulatory filing (SEC 10-K, EDGAR amended filing, Federal Register publication, OFAC SDN list update, EU dual-use list publication, Japanese end-use restriction notice), official government measurement (national grid interconnection data, EIA Form 860 / 861 / 923 disclosures, USGS economic geology reports), publisher-canonical published value (Epoch's frontier-DC corpus, MLPerf benchmark results, the SemiAnalysis 100K H100 decomposition), or on-page company disclosure with verifiable timestamp. Highest confidence. - **Tier 2**: Research-database / curated corpus. Epoch AI dataset, CSET ETO aggregation, the licensed corporate-ownership and supply-chain databases (which themselves aggregate primary-source filings — Tier-1-derived but treated as Tier 2 because the aggregation pass can introduce variance), PeeringDB self-reporting (operator-reported but unaudited), MLPerf benchmarks where the benchmark itself is third-party-verified. Medium-high confidence. - **Tier 3**: Press / analyst secondary. Goldman Sachs equity-research analyst reports cited with the analyst named, NYT / FT / Bloomberg / Wall Street Journal / Nikkei Asia / The Information reporting, IDC sovereign-AI tracking, Gartner data-center research, Congressional Research Service reports (CRS) citing analyst figures, CNAS Sovereign AI Index. Medium confidence. - **Tier 4**: Inferred / estimated. Scrutica-derived value through documented methodology with adjustable parameters. The confidence depends on the derivation chain — a Tier-4 FLOP estimate derived from a Tier-1 chip-count disclosure has higher confidence than a Tier-4 FLOP estimate derived from a Tier-3 power-capacity press report. Lower-tier sources are not excluded from Scrutica's substrate; they are labeled. Cross-source disagreements between tiers surface as contradiction flags rather than silent winner-picking. The rationale: a Tier-3 press report citing "industry sources" disagreeing with a Tier-1 SEC filing is itself information — either the press source is mistaken (and Scrutica's display should not collapse the disagreement) or the SEC filing's disclosure is incomplete (and Scrutica's display should make that visible). The contradiction-surfacing pattern preserves the editorial signal in both directions. **Tier-tagging conventions.** Each substrate table either carries an explicit `authority_tier` column per row (BIS designations, supply-chain edges, sovereign programs) or has an editorial documentation block in the data-access module's header naming the authority tier of the underlying corpus (e.g., `lib/data/facilities.ts` notes that Epoch-derived facilities are Tier 2 by source classification, with selected Tier-1 enrichments via primary SEC filings annotated per-row). --- ## Glossary (77 terms) The Scrutica glossary defines technical terms and their governance relevance. The full interactive glossary with cross-linked related terms is at https://scrutica.com/methodology/glossary. ### Semiconductor **Advanced packaging**. Post-fabrication techniques (CoWoS, InFO, EMIB, Foveros) that integrate multiple chiplets and memory stacks into a single package. Enables the large-die configurations required for AI accelerators. *Governance relevance:* Advanced packaging is the newest bottleneck in AI chip supply. Capacity is more concentrated than wafer fabrication, with TSMC controlling the majority of AI-relevant packaging. **CoWoS**. Chip-on-Wafer-on-Substrate: TSMC\ *Governance relevance:* CoWoS packaging capacity constrained AI chip supply through 2024-2025. TSMC controls nearly all advanced packaging for AI accelerators; alternative capacity (Samsung I-Cube) remains limited. *Example:* Every H100 and A100 GPU uses CoWoS packaging. **DUV lithography (DUV)**. Deep ultraviolet lithography uses 193 nm (ArF) or 248 nm (KrF) wavelength light. Multi-patterning techniques extend DUV to ~7 nm nodes at the cost of additional process steps and yield loss. *Governance relevance:* DUV machines are manufactured by ASML, Nikon, and Canon. China can access DUV (not EUV) equipment, which supports production at mature nodes but not cutting-edge AI chips. **EUV lithography (EUV)**. Extreme ultraviolet lithography uses 13.5 nm wavelength light to pattern features below 7 nm on silicon wafers. Requires a tin-droplet plasma source, multilayer mirrors, and vacuum-sealed optical path. *Governance relevance:* ASML is the sole manufacturer of EUV machines. Any disruption to ASML or its supply chain halts production of the most advanced AI chips worldwide. *Example:* ASML holds a complete monopoly on EUV lithography equipment. **Fabless**. A chip company that designs but does not manufacture its own chips. NVIDIA, AMD, Qualcomm, and Apple are fabless; they rely on foundries for production. *Governance relevance:* Fabless companies depend entirely on foundry access. Export controls on foundry services can restrict a fabless company\ **FinFET**. Fin Field-Effect Transistor: a 3D transistor architecture where the gate wraps around a raised "fin" of silicon. Used at 14nm through 3nm nodes. *Governance relevance:* The performance gains from FinFET architecture underpin every modern AI accelerator. The transition to GAA transistors at 2nm creates a new manufacturing chokepoint. **Foundry**. A semiconductor manufacturer that fabricates chips designed by other companies (fabless firms). TSMC, Samsung Foundry, and GlobalFoundries are the major foundries. *Governance relevance:* The foundry model concentrates manufacturing in a small number of companies and geographies. Taiwan hosts >60% of global advanced-node foundry capacity. **GAA (GAA)**. Gate-All-Around transistor: the successor to FinFET where the gate completely surrounds the channel. Samsung and TSMC are transitioning to GAA at 2nm and below. *Governance relevance:* GAA manufacturing requires new equipment and process expertise. Countries and companies without access to GAA technology face a capability gap in AI chip production that grows with each successive node. **HBM (HBM)**. High Bandwidth Memory: vertically stacked DRAM dies connected by through-silicon vias (TSVs). HBM3e provides ~4.8 TB/s bandwidth per stack. Required for modern AI accelerators. *Governance relevance:* SK Hynix (~${HBM_SK_HYNIX_PCT}%), Samsung (~${HBM_SAMSUNG_PCT}%), and Micron (~${HBM_MICRON_PCT}%) produce all HBM (${HBM_VINTAGE}, ${HBM_SOURCE}). HBM supply is a binding constraint on AI accelerator production, independent of logic chip availability. *Example:* H100 SXM5 ships with 80GB HBM3; H100 PCIe uses HBM2e. H200 SXM with 141GB HBM3e. **IDM (IDM)**. Integrated Device Manufacturer: a company that both designs and fabricates its own chips. Intel, Samsung, and Texas Instruments are IDMs. *Governance relevance:* IDMs have more supply chain resilience than fabless companies but face higher capital requirements. The US CHIPS Act subsidies primarily target IDM-style domestic manufacturing. **OSAT (OSAT)**. Outsourced Semiconductor Assembly and Test: companies that package and test fabricated chips. ASE, Amkor, and JCET are major OSATs. *Governance relevance:* OSATs are concentrated in Taiwan, China, and Malaysia. Advanced packaging (required for AI chips) is more concentrated than standard OSAT services. **Photomask**. A quartz plate with patterned chrome features that defines the circuit layout projected onto the wafer during lithography. A single advanced chip design requires 80-100 mask layers. *Governance relevance:* Photomask sets for advanced nodes cost $10-30M and take months to produce. This cost barrier limits who can design cutting-edge chips. **Process node**. A named manufacturing generation (e.g., N4, N5, N7) indicating transistor density and power characteristics. Modern node names are marketing labels; actual gate lengths differ from the stated nanometer value. *Governance relevance:* Export controls target chips manufactured at advanced nodes (roughly 14 nm and below). The node designation determines whether a chip is freely exportable or restricted. *Example:* TSMC N4 achieves ~1.7x logic density over N7. **Wafer**. A thin disc of crystalline silicon (typically 300 mm diameter) on which hundreds of chip dies are patterned simultaneously. Wafer starts per month is the standard measure of fab capacity. *Governance relevance:* Global advanced-node wafer capacity is concentrated in Taiwan, which holds ~${TAIWAN_LEADING_EDGE_PCT}% of leading-edge pure-play foundry manufacturing (CSET 2021, a country-level share including TSMC, UMC, and PSMC). TSMC is the dominant operator within it (no tier-1/2 source publishes its company-level slice); its wafer allocation decisions determine which AI chips get manufactured. **Yield**. The fraction of functional dies per wafer. Advanced-node yields typically start at 30-50% and improve to 80%+ over 12-18 months of production maturity. *Governance relevance:* Low yields at new process nodes constrain chip supply regardless of wafer capacity. Yield data is closely guarded; its absence is itself a signal of production difficulty. ### Accelerator **BF16**. Brain floating-point 16-bit format: same total bits as FP16 but with 8 exponent bits (matching FP32) and 7 mantissa bits. Preferred for training stability. *Governance relevance:* BF16 and FP16 deliver comparable throughput on modern GPUs. Governance thresholds typically count operations regardless of precision format. **FLOP (FLOP)**. Floating-point operation: a single arithmetic operation (add, multiply) on a floating-point number. FLOP/s (per second) measures raw compute throughput; total FLOP measures cumulative compute used for training. *Governance relevance:* FLOP is the primary metric for AI governance thresholds. The EU AI Act requires reporting for models trained above 10²⁵ FLOP; the former US Executive Order 14110 set notification at 10²⁶ FLOP. *Example:* GPT-4 training compute is estimated (Tier 3) at approximately 2 × 10²⁵ FLOP. **FP16**. Half-precision 16-bit floating-point format. Most AI training uses FP16 or BF16 arithmetic; FLOP/s benchmarks for AI typically report FP16 throughput. *Governance relevance:* The precision format determines the effective compute per operation. Export controls reference total operations per second, which varies by precision mode. **InfiniBand**. High-bandwidth, low-latency networking fabric used to connect GPU nodes in AI training clusters. NVIDIA (via Mellanox acquisition) dominates the AI-grade InfiniBand market. *Governance relevance:* InfiniBand is a controlled technology under US export rules. Restricting InfiniBand access limits a country\ **Memory bandwidth**. The data transfer rate between GPU compute units and memory, measured in TB/s. H100 SXM: 3.35 TB/s; H200 SXM: 4.8 TB/s (with HBM3e). *Governance relevance:* Memory bandwidth is a binding constraint for many AI workloads. Export control thresholds consider both compute throughput and memory bandwidth as dual criteria. **NVLink**. NVIDIA\ *Governance relevance:* NVLink is a sole-source technology from NVIDIA. Clusters using NVLink achieve higher training efficiency than those using standard PCIe or InfiniBand for GPU communication. **PCIe (PCIe)**. PCI Express: the standard expansion bus interface. PCIe GPUs (e.g., H100 PCIe) have lower TDP (350W vs. 700W) and no NVLink, limiting large-scale training capability. *Governance relevance:* Some export control frameworks distinguish between SXM and PCIe variants. PCIe GPUs are less useful for frontier AI training, so export restrictions may treat them differently. **SXM**. NVIDIA\ *Governance relevance:* SXM GPUs are the building blocks of large AI training clusters. The SXM vs. PCIe distinction matters for export controls because SXM enables cluster-scale training that PCIe cannot. **TDP (TDP)**. Thermal Design Power: the maximum sustained power draw of a chip in watts. H100 SXM: 700W; H100 PCIe: 350W; H200 SXM: 700W; B200: 1,000W. *Governance relevance:* TDP determines data center power requirements and cooling needs. A facility\ *Example:* Power estimation path: 425 MW facility / 700W per H100 = ~607K GPUs maximum. **Tensor core**. Specialized matrix-multiply units in NVIDIA GPUs that perform mixed-precision fused multiply-add operations. Responsible for the majority of AI training throughput. *Governance relevance:* Tensor core count and generation determine a GPU\ **TFLOP/s**. Tera (10¹²) floating-point operations per second. The standard unit for GPU compute throughput. H100 SXM: 989.5 TFLOP/s FP16 dense, 1,979 TFLOP/s with 2:4 structured sparsity. *Governance relevance:* TFLOP/s ratings determine how quickly a facility can accumulate training compute. Higher TFLOP/s per GPU means fewer GPUs are needed to reach regulatory FLOP thresholds. ### Training **Chinchilla-optimal**. A training configuration that balances model size and dataset size according to the Hoffmann et al. (2022) scaling law. Chinchilla-optimal models use ~20 tokens per parameter. *Governance relevance:* Chinchilla-optimal training requires less compute than over-parameterized training for the same capability. This complicates governance: a FLOP threshold can be reached with a smaller, better-trained model. **Fine-tuning**. Continued training of a pre-trained model on a smaller, task-specific dataset. Uses 100-10,000x less compute than the original pre-training run. *Governance relevance:* Fine-tuning can significantly alter a model\ **Inference**. Running a trained model to produce outputs (predictions, text, images). Inference compute per query is orders of magnitude smaller than total training compute, but aggregate inference demand can exceed training demand. *Governance relevance:* Inference compute is not currently subject to governance thresholds, but aggregate inference demand drives data center capacity expansion and energy consumption. **METR time horizon (METR)**. Model Evaluation and Threat Research benchmark measuring how long an AI agent can autonomously work on real-world tasks. Scored on a 0-5 scale corresponding to task durations from seconds to weeks. *Governance relevance:* METR scores provide a standardized measure of autonomous capability. Higher METR scores at lower compute levels would indicate capability is becoming more accessible, a key governance concern. *Example:* Frontier models in 2025 achieve METR time horizons of 2-3 (minutes to hours). **MFU (MFU)**. Model FLOP Utilization: the fraction of theoretical peak GPU throughput actually achieved during training. Typical values: 30-50% for large clusters. Higher MFU means more efficient hardware utilization. *Governance relevance:* MFU determines how much real compute a facility extracts from its hardware. Two identical GPU clusters can differ 2x in effective training compute based on MFU alone. *Example:* Scrutica\ **Parameter count**. The number of trainable weights in a neural network. GPT-4 is widely reported (Tier 3, analyst estimates; OpenAI has not confirmed) at ~1.8 trillion parameters (MoE). Larger parameter counts require more memory and compute. *Governance relevance:* Parameter count was an early proxy for model capability but is less informative than training compute. Current governance frameworks use FLOP thresholds rather than parameter thresholds. **Scaling law**. Empirical power-law relationship between training compute, dataset size, model parameters, and model performance. Kaplan et al. (2020) and Hoffmann et al. (2022, "Chinchilla") established key scaling relationships. *Governance relevance:* Scaling laws predict that spending more compute on training yields more capable models. This predictability enables governance frameworks that use compute thresholds as capability proxies. **Training compute**. The total floating-point operations used to train a model, measured in FLOP. Distinct from inference compute (running a trained model) and fine-tuning compute. *Governance relevance:* Training compute is the primary measurable input for AI capability governance. Higher training compute correlates with more capable (and potentially more dangerous) models. *Example:* The EU AI Act uses 10²⁵ total training FLOP as the GPAI reporting threshold. ### Datacenter **GPU cluster**. A set of GPU servers connected by high-bandwidth networking (NVLink + InfiniBand) that can collectively train a single model. Modern frontier clusters contain 10,000-100,000+ GPUs. *Governance relevance:* GPU cluster scale determines what size models a facility can train. Clusters above certain sizes can produce models that exceed governance thresholds. **Hyperscaler**. Cloud providers operating at massive scale: Amazon (AWS), Microsoft (Azure), Google (GCP), Meta, Oracle. Characterized by >100 MW data center campuses, custom silicon programs, and global infrastructure. *Governance relevance:* Hyperscalers control the majority of commercially accessible AI compute. Their infrastructure investment decisions (where to build, what chips to buy) shape national compute capacity. **Interconnect fabric**. The network infrastructure connecting GPU nodes within a cluster. Combines NVLink (intra-node), InfiniBand or Ethernet (inter-node), and spine-leaf topology for scalable bandwidth. *Governance relevance:* Interconnect quality determines maximum trainable model size. Without adequate interconnect, additional GPUs cannot participate in a single training run, so raw GPU count overstates effective cluster scale. **Liquid cooling**. Direct-to-chip or immersion cooling that removes heat more efficiently than air cooling. Required for GPUs above ~400W TDP (Blackwell B200 at 1,000W mandates liquid cooling). *Governance relevance:* Liquid cooling infrastructure is a prerequisite for next-generation AI clusters. Facilities without it cannot host the latest accelerators; the result is a capacity gap between upgraded and legacy sites. **MW (MW)**. Megawatt: 1 million watts. The standard measure of data center power capacity. A 100 MW facility can host roughly 100,000-140,000 H100 GPUs at full load (accounting for PUE). *Governance relevance:* Power capacity is often the binding constraint on AI compute. A 500 MW data center campus represents a significant national infrastructure investment and a measurable concentration of AI capability. *Example:* The xAI Colossus facility in Memphis draws approximately 150 MW. **Neocloud**. GPU cloud providers (CoreWeave, Lambda, Nebius, Applied Digital, IREN) that rent compute capacity without the broad cloud services offered by hyperscalers (AWS, GCP, Azure). *Governance relevance:* Neoclouds have opened AI-specific compute capacity outside hyperscaler ecosystems, but their near-total dependency on NVIDIA GPUs replicates the concentration dynamic at a different layer. **PUE (PUE)**. Power Usage Effectiveness: ratio of total facility power to IT equipment power. PUE 1.0 = all power goes to compute; PUE 1.2 = 20% overhead for cooling, lighting, networking. Industry average: ~1.3. *Governance relevance:* PUE determines how much of a facility\ *Example:* Scrutica\ **Utilization rate**. The fraction of installed GPU capacity actively performing compute, averaged over time. Typical values: 50-80% for training workloads. Affected by scheduling, maintenance, and workload mix. *Governance relevance:* Utilization determines how much of a facility\ ### Supplychain **Cascade**. The propagation of a disruption through supply chain relationships. Modeled via breadth-first search across the relationship graph with decay factors based on supply share and criticality. Criticality can be editorial (1-10 replaceability scale) or blended with a 3-month price correlation where available. Blend formula: adjustedCriticality = w \u00d7 5(1+r) + (1-w) \u00d7 editorial. *Governance relevance:* Cascade analysis surfaces the second- and third-order effects of concentration in the supply chain. Removing or constraining a single node may reshape access for dozens of downstream entities; the propagation depth indicates how much policy leverage a given chokepoint provides. *Example:* The cascade simulator models dependency propagation across 18,500+ supply chain edges. **Chokepoint**. A supply chain node where concentration or sole-source dependency creates a governance-relevant bottleneck. Measured by HHI, node centrality, and substitutability assessment. *Governance relevance:* Chokepoints are where export controls have maximum leverage and where disruptions have maximum impact. They serve as the natural targets for risk assessment and strategic policy intervention. **HHI (HHI)**. Herfindahl-Hirschman Index: the sum of squared market shares. Ranges from near 0 (perfect competition) to 10,000 (monopoly). Calculated as HHI = sum(s_i^2) where s_i is each firm\ *Governance relevance:* The US Department of Justice considers HHI above 2,500 as highly concentrated. Measured AI compute supply chain segments: EUV lithography at 10,000 (ASML monopoly, tier 1), overall foundry at ${FOUNDRY_FY2025_NAMED_HHI.toLocaleString('en-US')} named-share lower bound (TrendForce FY2025), HBM at ${HBM_HHI.toLocaleString('en-US')} (Counterpoint Q3 2025). Advanced packaging publishes no tier-1/2 share set, so no packaging HHI exists. *Example:* EUV lithography HHI = 10,000 (ASML monopoly). **Licensed supply-chain database**. A licensed institutional supply-chain relationship database, held under subscription and not redistributed. Contains bilateral supplier-customer links with relationship value ($M), relationship rank, and 3-month rolling stock price correlation. Authority tier 2: the underlying facts (who supplies whom) are surfaced; the vendor records themselves are not republished. *Governance relevance:* An institutional-grade data source financial analysts use to map supply chain dependencies. Scrutica uses it as a tier-2 source for bilateral supply chain edges, with public filings preferred wherever they cover the same relationship. **Price correlation**. The 3-month rolling Pearson correlation between two companies\ *Governance relevance:* Incorporated into the cascade simulator as an empirical edge weight. High positive correlation confirms that the market sees tight coupling; negative correlation suggests the supply relationship is not the primary driver of either company\ **Sole-source dependency**. A supply chain relationship where only one company produces a critical input. ASML for EUV machines, TSMC for advanced-node foundry services, and Zeiss for EUV optics are sole-source dependencies. *Governance relevance:* Sole-source dependencies are the most concentrated chokepoints in the compute supply chain, and therefore the points where export controls and policy interventions have maximum effect. No alternative supplier exists at any price. *Example:* The concentration page identifies supply chain steps with sole-source exposure. **Substitutability**. Expert-assessed score (0-1) indicating how easily a disrupted supplier can be replaced. 0 = no substitute exists (ASML for EUV); 1 = drop-in replacements available. Authority tier 3 (analyst assessment). *Governance relevance:* Substitutability determines cascade severity. Low substitutability at a chokepoint means disruption propagates with minimal attenuation. **Supply chain edge**. A directed relationship between two entities in the supply chain graph. Edges carry attributes: relationship type (supplier/customer/partner), value ($M), and source (licensed supply-chain database or SEC filing). *Governance relevance:* The supply chain graph surfaces hidden dependencies and cascade paths. Two companies that appear unrelated may share a critical upstream supplier, creating correlated exposure invisible from a bilateral view. ### Governance **Compute governance**. The use of compute as a governance lever for AI policy. Based on the premise that compute is measurable, excludable, and concentrated — properties that make it a tractable intervention point. *Governance relevance:* Compute governance is Scrutica\ **Dual-use**. Technology with both civilian and military applications. All advanced AI chips are dual-use; the same GPU that trains a language model can train a military targeting system. *Governance relevance:* The dual-use nature of AI chips is the legal foundation for export controls. Wassenaar, EAR, and EU dual-use regulations all regulate AI-relevant technology on dual-use grounds. **EU AI Act**. Regulation (EU) 2024/1689: the European Union\ *Governance relevance:* The EU AI Act is the first major jurisdiction to use training compute as a regulatory threshold. Its 10²⁵ FLOP line directly connects facility-level compute capacity to regulatory compliance. *Example:* Scrutica\ **GPAI (GPAI)**. General-purpose AI: the category defined by the EU AI Act for models trained on broad data that can perform a wide range of tasks. Models trained above 10²⁵ FLOP are classified as GPAI with systemic risk. *Governance relevance:* GPAI classification triggers transparency and reporting obligations under the EU AI Act. The 10²⁵ FLOP threshold determines which models face the most stringent requirements. **Proliferation**. In the AI governance context: the spread of advanced AI capabilities to additional actors. Distinct from weapons proliferation but often analyzed with similar frameworks (supply-side controls, demand-side monitoring). *Governance relevance:* Compute proliferation tracking (who is acquiring what compute capability, and where) is one of Scrutica\ **Reality ratio**. Scrutica metric: the ratio of deployed/disbursed spending to announced/committed spending for sovereign AI programs. Reality ratio 0.1 means 10% of announcements have reached operational deployment. *Governance relevance:* The reality ratio measures execution: what fraction of a sovereign AI commitment has reached operational deployment. A country with $10B announced and 5% reality ratio has $500M of actual compute. The complementary metric is allied governance reach, which measures supply chain exposure to allied export control jurisdiction. *Example:* The sovereign AI dashboard tracks reality ratios for every country with a national AI program. **Sovereign AI**. National programs to develop domestic AI compute capacity, independent of foreign hyperscalers. Includes government-funded GPU clusters, national cloud initiatives, and indigenous chip development. *Governance relevance:* Two questions define each sovereign program: execution (what fraction of announced spending reaches operational infrastructure) and supply chain exposure (what fraction of the program\ *Example:* Scrutica tracks 30+ sovereign AI programs with separate government and private investment figures. **Systemic risk**. Under the EU AI Act, the risk classification for GPAI models with "high-impact capabilities" (currently proxied by training compute above 10²⁵ FLOP). Triggers additional obligations: red-teaming, incident monitoring, cybersecurity. *Governance relevance:* Systemic risk classification is the strictest tier of EU AI regulation. Whether a model qualifies depends on its training compute, which depends on the facility\ ### Exportcontrol **BIS (BIS)**. Bureau of Industry and Security: the US Department of Commerce agency that administers export controls. Issues licenses, maintains the Entity List, and promulgates rules under the EAR. *Governance relevance:* BIS is the institutional center of US technology export control policy. Its rule-making actions directly determine which countries and companies can access advanced AI chips. **De minimis rule**. EAR provision that exempts foreign-made items containing less than a specified percentage (typically 25%) of controlled US-origin content. Threshold varies by destination and item. *Governance relevance:* The de minimis rule determines how far US export controls extend into foreign supply chains. A lower de minimis threshold (e.g., 10% for Entity List destinations) extends US regulatory reach. **Deemed export**. Transfer of controlled technology to a foreign national within the United States. Governed by the same EAR classifications as physical exports. *Governance relevance:* Deemed export rules affect AI research collaboration. A US university lab with foreign researchers may need export licenses for work involving controlled computing technology. **EAR (EAR)**. Export Administration Regulations: the US legal framework governing export of dual-use items. Administered by BIS (Bureau of Industry and Security). Key classifications: 3A090 (advanced computing ICs), 4A090 (computers containing 3A090 chips). *Governance relevance:* The EAR is the primary tool the US uses to restrict access to advanced AI chips. It determines which chips require licenses for which destinations, and which are prohibited entirely. **End-use controls**. Export restrictions based on the intended application rather than the item\ *Governance relevance:* End-use controls are harder to enforce than technical threshold controls because they require knowledge of the buyer\ **Entity List**. A BIS-maintained list of foreign entities subject to specific export restrictions. Inclusion requires a license for most items subject to the EAR, with a presumption of denial for advanced computing. *Governance relevance:* Entity List designation is the strongest export control tool short of a full embargo. It can cut a company off from the US technology supply chain entirely. *Example:* Scrutica cross-references BIS Entity List entries against the supply chain graph. **License exception**. An authorization under the EAR that permits export without an individual license when specified conditions are met. Examples: License Exception TSR (technology/software under restriction), ACE (authorized cybersecurity exports). *Governance relevance:* License exceptions create controlled pathways for technology transfer. Their scope and conditions are frequently modified; changes to license exceptions can reclassify thousands of transactions overnight. **TOPS (TOPS)**. Tera operations per second: the performance metric used in US export control rules to determine whether a chip is "advanced." The October 2022 rule set thresholds at 300 TOPS (for certain interconnect bandwidth) and 600 TOPS. *Governance relevance:* TOPS thresholds are the technical criteria that divide freely exportable chips from restricted ones. Chip designers may deliberately design below threshold to avoid export controls. **Wassenaar Arrangement**. A multilateral export control regime with 42 participating states. Sets baseline control lists for conventional arms and dual-use technologies, including semiconductor equipment and advanced computing. *Governance relevance:* Wassenaar provides the multilateral framework that US, EU, and Japanese export controls build on. Countries outside Wassenaar may serve as transshipment points for controlled technology. *Example:* The coordination gaps page identifies countries in the supply chain but outside Wassenaar. ### Methodology **Authority tier**. Scrutica\ *Governance relevance:* Authority tiers communicate how much weight to place on a data point. A Tier 1 figure from a company\ **BFS propagation (BFS)**. Breadth-first search: a graph traversal algorithm used in the cascade simulator to model disruption spreading hop-by-hop through the supply chain. Each hop applies a decay factor based on edge substitutability. *Governance relevance:* BFS propagation models how a disruption at one point in the supply chain ripples outward. The number of hops before the impact becomes negligible indicates how resilient the supply chain is to that disruption. **Capex disaggregation**. Decomposing a company\ *Governance relevance:* Capex disaggregation enables the cost-based estimation path for facility-level compute. When a company reports $5B in data center capex, disaggregation estimates how many GPUs that purchases. **Confidence interval**. A range of values within which the true value is estimated to fall with a stated probability (typically 90%). Scrutica reports estimation bounds rather than point estimates where multiple estimation paths are available. *Governance relevance:* Confidence intervals communicate honest uncertainty. A wide interval (e.g., 50-200 PFLOP/s) tells a policymaker the estimate is uncertain; a narrow one (e.g., 80-100 PFLOP/s) supports stronger conclusions. **Decay rate**. The per-hop reduction in disruption impact during cascade propagation. Calculated as 1 minus substitutability score, clamped at 0.8. Higher decay = faster attenuation = more resilient supply chain. *Governance relevance:* Decay rates encode how easily a disrupted supplier can be replaced. A low decay rate (hard to substitute) means the disruption cascades further before dissipating. **Estimation bounds**. The range produced by cross-validating multiple estimation paths. When hardware, power, and cost paths agree within 15%, confidence is high; when they diverge, the bounds are wide. *Governance relevance:* Estimation bounds are Scrutica\ *Example:* The FLOP engine shows cross-validated estimation bounds for each facility. **is_estimated**. A boolean flag on every Scrutica data point. True if the value is derived or inferred; false only if it comes from a primary source with direct measurement. *Governance relevance:* The is_estimated flag distinguishes verified facts from modeled estimates. Policymakers should cite estimated values with appropriate caveats. **Petaflop-day (PF-day)**. One petaFLOP (10¹⁵ FLOP/s) sustained for one day (86,400 seconds) = 8.64 × 10¹⁹ FLOP. Used as the unit for the Compute Cost Index. *Governance relevance:* Petaflop-days normalize compute across different hardware and providers so that cost comparisons are meaningful. The $/petaFLOP-day metric shows how much it costs to train AI in different jurisdictions. *Example:* The cost index reports compute pricing in $/petaFLOP-day across cloud providers and regions. **Provenance**. The documented origin and transformation history of a data point. Every Scrutica value carries data_source, source_url, and is_estimated fields. *Governance relevance:* Provenance enables citation and verification. A policy researcher can trace any number on the platform back to its original source and assess whether to trust it. **Scrutica Compute Unit (SCU)**. Scrutica\ *Governance relevance:* SCU makes provider, region, and hardware-generation pricing directly comparable. A reader can answer "is AWS on-demand H100 cheaper than CoreWeave A100 per unit of training output?" without doing the unit conversion by hand. Distinct from raw $/hour because $/SCU bakes in the GPU\ *Example:* A 3-year reserved CoreWeave H100 8-GPU instance prices at ~$10/SCU; AWS p5 on-demand prices at ~$50/SCU. --- ## Index of canonical entities Selected entries from the canonical-organisation table (129,061 total rows; 24,873 compute-relevant). The full directory is at https://scrutica.com/entities/companies. Each card pairs a Scrutica-resolved canonical ID with a public-facts narrative grounded in the entity's regulatory filings, vendor disclosures, and structured-data substrate. Tickers and HQ are cross-checked against `organizations.json` at build time. ### NVIDIA Corporation (`org-nvidia`) Largest designer of GPUs and AI accelerators by revenue. Product lines spanning the H100 SXM5 (peak 989.5 BF16 TFLOPS dense), H200 (HBM3e refresh), B100/B200 (Blackwell, 2024), and the GB200 NVL72 / GB300 system-class rack designs that integrate Grace ARM CPU + Blackwell GPU + NVSwitch fabric for 72-GPU coherent training pods. CUDA software ecosystem is the de-facto AI training stack; the moat is the toolchain, not the silicon alone. Manufactures via TSMC (advanced nodes) and Samsung Foundry (selected workloads). HBM sourcing from SK Hynix (primary), Samsung, Micron. The single most-cited dependency anchor in Scrutica's cascade simulation — the seed node for the TSMC-disruption scenario family is NVIDIA's customer fan-out. Headquartered in the United States. Public-market ticker: NVDA. Detail page: https://scrutica.com/companies/org-nvidia. ### Taiwan Semiconductor Manufacturing Company (`org-tsmc`) Largest pure-play foundry on advanced nodes. Manufactures for NVIDIA (H100/B100/B200 on N4 + N3 + advanced packaging via CoWoS), AMD (MI300 / MI325 / MI355 series), Apple (A18 / M-series), Qualcomm, and a long tail of fabless designers. The structural single-point chokepoint for frontier AI silicon: above the 5nm node, TSMC's share of cutting-edge logic capacity exceeds 90% by some industry-research measures. Headquartered in Hsinchu, Taiwan; primary fab cluster in the Hsinchu Science Park and Taichung; secondary operations in Tainan (Fab 18, advanced nodes). Geographic concentration is the cascade-simulation's most-load-bearing weighting input — any scenario that affects Taiwan-island operations propagates to virtually every frontier-AI training cluster within months. Public-market ticker: TSM. Detail page: https://scrutica.com/companies/org-tsmc. ### ASML (`org-asml`) Sole global supplier of EUV (extreme ultraviolet) lithography systems, the equipment class required to pattern logic features below 7nm at commercial yield. High-NA EUV (NXE:5000-class) shipped initial units 2024-2025 for next-generation node development. ASML's customer concentration: TSMC, Samsung Foundry, Intel are the top three EUV recipients globally. Headquartered Veldhoven, Netherlands. Subject to a layered export-control regime: Dutch national licenses for EUV exports to China since 2019, US-coordinated controls on DUV immersion since 2023-2024, and post-CHIPS-Act / EAR-revision restrictions on field-service personnel access. The legal-access layer for advanced-node manufacturing globally. Public-market ticker: ASML. Detail page: https://scrutica.com/companies/org-asml. ### Samsung (`org-samsung`) Conglomerate operating two distinct compute-supply-chain businesses: Samsung Electronics Memory Division (HBM3, HBM3e supplier — historically lagged SK Hynix on NVIDIA HBM qualification; entered HBM4 qualification phase 2025-2026) and Samsung Foundry (5nm-3nm-2nm logic competing with TSMC and Intel, with Tesla Dojo and US government contracts among the announced customers). Korean conglomerate parent (Samsung Group). The HBM market is the cleanest three-supplier oligopoly in the AI compute stack — Samsung as the second / third entrant depending on quarter. Headquartered in South Korea. Public-market ticker: 005930. Detail page: https://scrutica.com/companies/org-samsung. ### SK Hynix (`org-sk-hynix`) Largest HBM supplier by NVIDIA-qualified share through 2024-2025. HBM3 first-to-qualify with NVIDIA; HBM3e in volume from 2024; HBM4 qualification phase H1 2026. Headquartered Icheon, South Korea. Subsidiary of SK Group. The HBM3e shortage is the single most-cited near-term AI compute chokepoint — primary supply is committed through 2025 with limited spot availability, and the qualification cycle for new fabs is 12-18 months. Public-market ticker: 000660. Detail page: https://scrutica.com/companies/org-sk-hynix. ### Micron (`org-micron`) Late entrant to the HBM market — qualified HBM3e with NVIDIA in 2024 after the H100 generation had largely shipped. HBM share remains structurally below SK Hynix and Samsung; expansion plans target Boise (Idaho), Manassas (Virginia), Hiroshima (Japan), and new Singapore facilities. US-headquartered, making Micron a strategic CHIPS Act recipient. Ticker MU. Detail page: https://scrutica.com/companies/org-micron. ### Huawei (`org-huawei`) Chinese telecommunications and consumer-electronics conglomerate; designs the Ascend AI accelerator line (910B series most prominent in 2024-2025; 920 series successor in qualification phase). On the US BIS Entity List since May 2019 (Federal Register notice 84 FR 22961); subsequent rounds of controls (Oct 2022, Oct 2023) restrict advanced-node fabrication services and HBM access. HiSilicon (chip design subsidiary) and Huawei Cloud (cloud services) are the primary AI-compute-relevant business units. Domestic substitution narrative — Ascend on SMIC's N+1 / N+2 nodes substituting for export-restricted NVIDIA — is the central thesis Scrutica's Diversion Pipeline Tracker and Trade Impact pages render. Headquartered in China. Detail page: https://scrutica.com/companies/org-huawei. ### SMIC (`org-smic`) Largest Chinese foundry by capacity. On the BIS Entity List since December 2020; subject to advanced-node export restrictions affecting equipment from ASML (EUV blocked entirely), Tokyo Electron, Applied Materials, KLA, and Lam Research. SMIC's advanced-node yield disclosures (N+1 / N+2 / N+3 process names referencing 14nm-class through 5nm-class capabilities) are partially substantiated through Ascend / HiSilicon teardown analysis but lack direct foundry confirmation. The cross-source-divergence flag system surfaces the gap between disclosed and inferred capacity for SMIC's advanced-node fabrication. Headquartered in China. Detail page: https://scrutica.com/companies/org-smic. ### Hua Hong (`org-hua-hong`) Major Chinese foundry operating across mature nodes (12nm and above). Subject to Scrutica's cross-source-divergence flag (Federal Register designations vs SMIC-affiliate disclosures vs Hua Hong's own published capacity). On the BIS Entity List under multiple subsidiary designations. The data quality flag system at organization_data_quality_flags records the contradictions. Headquartered in China. Public-market ticker: 1347.HK. Detail page: https://scrutica.com/companies/org-hua-hong. ### Intel (`org-intel`) US-headquartered semiconductor designer and manufacturer. Three primary compute-supply-chain roles: (1) Intel Foundry Services — challenging TSMC at the advanced-node leading edge (Intel 18A in 2024-2025 risk production), targeting external customers including Microsoft and US government; (2) Gaudi AI accelerator line (Gaudi 3 generation, 2024); (3) Xeon CPU line used as the host processor in many GPU-cloud configurations. Major CHIPS Act recipient (Arizona, Ohio, New Mexico expansion). Public-market ticker: INTC. Detail page: https://scrutica.com/companies/org-intel. ### AMD (`org-amd`) Designs the Instinct MI-series AI accelerators (MI300X 2024, MI325X late 2024, MI355X / MI400 generation 2025-2026 roadmap). Distant second to NVIDIA in AI training-class GPU market share but with significant traction in select hyperscaler deployments (Microsoft, Meta have publicly disclosed MI300X clusters). Headquartered Santa Clara, CA. Manufacturing via TSMC; same supply-chain exposure as NVIDIA for advanced-node + HBM + CoWoS. Public-market ticker: AMD. Detail page: https://scrutica.com/companies/org-amd. ### Tokyo Electron (`org-tokyo-electron`) Major Japanese semiconductor-equipment supplier. Critical product categories: coaters, developers, etchers, deposition systems. Subject to Japan's October 2023 export controls coordinated with US BIS rules — restricts a defined list of process tools to mainland China. Headquartered Tokyo. Ticker 8035.T. Detail page: https://scrutica.com/companies/org-tokyo-electron. ### Applied Materials (`org-applied-materials`) Largest semiconductor-equipment supplier globally by revenue. Process tool categories spanning deposition (CVD, PVD, ALD), epitaxy, ion implantation, etch, CMP, and metrology. Subject to the same US-coordinated export-control regime as Tokyo Electron and Lam Research. Headquartered Santa Clara, CA. Ticker AMAT. Detail page: https://scrutica.com/companies/org-applied-materials. ### KLA (`org-kla`) Largest semiconductor process-control equipment vendor (inspection, metrology, defect detection). Effectively single-source for several leading-edge metrology categories. Subject to US export controls on advanced-node service to mainland China. Headquartered Milpitas, CA. Ticker KLAC. Detail page: https://scrutica.com/companies/org-kla. ### Lam Research (`org-lam-research`) Major US-headquartered etch and deposition equipment vendor. Critical position in plasma etch (dielectric, conductor) and dep-tools (ALD, PECVD). Subject to US export-control regime on advanced-node service. Headquartered Fremont, CA. Ticker LRCX. Detail page: https://scrutica.com/companies/org-lam-research. ### Equinix, Inc. (`org-equinix`) Global colocation operator. Operates IBX (International Business Exchange) facilities across 30+ countries, hosting hyperscaler edge presence + neutral interconnection. Critical infrastructure substrate for many GPU-cloud and frontier-AI providers — not an AI accelerator owner but the physical hosting layer for a meaningful fraction of commercial AI compute outside hyperscaler-owned facilities. Ticker EQIX. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-equinix. ### Digital Realty (`org-digital-realty`) Largest colocation REIT by revenue. Operates wholesale and retail data-center capacity globally. Hyperscaler-tenant exposure across AWS, Azure, GCP, Meta, Oracle. Subject to compute-visibility-tier-1 transparency (REIT regulatory disclosures). Headquartered Austin, TX. Ticker DLR. Detail page: https://scrutica.com/companies/org-digital-realty. ### CoreWeave (`org-coreweave`) Largest pure-play GPU cloud (neocloud) by deployed capacity. NVIDIA partnership at the heart of the business model — first major commercial deployer of H100 and Blackwell-generation hardware. IPO 2025 (NASDAQ: CRWV). Revenue concentration with Microsoft and OpenAI as anchor tenants. The canonical "neocloud" example: pure-play GPU access at scale outside the hyperscaler set. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-coreweave. ### Microsoft Corporation (`org-microsoft`) Operates Azure global cloud + Microsoft AI Compute infrastructure. Anchor commercial partner to OpenAI through equity investment and exclusive cloud-provider arrangement (modified 2024-2025 to allow OpenAI multi-cloud workloads). Major data-center operator across US, EU, UAE, Saudi Arabia, India. Co-designer of the Maia AI accelerator line. Hyperscaler-class compute-visibility-tier-1 transparency. Headquartered in the United States. Public-market ticker: MSFT. Detail page: https://scrutica.com/companies/org-microsoft. ### Amazon.com Inc (`org-amazon`) Operates AWS global cloud. Designs the Trainium and Inferentia in-house AI accelerator families through Annapurna Labs. Anthropic anchor cloud-provider relationship (Anthropic compute partnership). Major data-center operator across US, EU, India, Japan, Singapore. Hyperscaler-class transparency. Headquartered in the United States. Public-market ticker: AMZN. Detail page: https://scrutica.com/companies/org-amazon. ### Alphabet Inc (Google) (`org-alphabet`) Operates Google Cloud Platform (GCP) + the TPU v4 / v5p / v6e / v7 AI accelerator lines used internally for DeepMind and Search training plus rented externally via GCP. The TPU + Pathways software stack is the second-largest non-NVIDIA AI training ecosystem by deployed capacity. Hyperscaler-class transparency. Headquartered in the United States. Public-market ticker: GOOGL. Detail page: https://scrutica.com/companies/org-alphabet. ### Meta Platforms Inc (`org-meta`) Self-builds AI training infrastructure at hyperscaler scale (Llama family training). Co-designer of the MTIA AI accelerator (in production deployment alongside NVIDIA H100 / H200 / Blackwell). Operates multi-gigawatt-class facilities under direct construction (Hyperion in Louisiana; multi-site US expansion). Hyperscaler-class transparency. Headquartered in the United States. Public-market ticker: META. Detail page: https://scrutica.com/companies/org-meta. ### OpenAI (`org-openai`) Frontier AI lab; produces the GPT family of models. Cloud-provider relationship anchored at Microsoft Azure, modified 2024-2025 to allow multi-cloud (Oracle Cloud Stargate UAE; CoreWeave deployments). Private company; Microsoft minority equity investor + secondary investor pool. Major direct compute commitments through 2026 — Project Stargate (Texas), Stargate UAE (G42 partnership). Compute-visibility tier 3-4 (private company, partial disclosure). Headquartered in the United States. Detail page: https://scrutica.com/companies/org-openai. ### Anthropic (`org-anthropic`) Frontier AI lab; produces the Claude family of models. Cloud-provider relationships: Amazon AWS (anchor investor + cloud partner — Project Rainier the named compute commitment) and Google Cloud (secondary). Private company; Amazon + Google as named investors. Compute-visibility tier 3 (private company; selective AWS commitment disclosure). Headquartered in the United States. Detail page: https://scrutica.com/companies/org-anthropic. ### Google DeepMind (`org-google-deepmind`) DeepMind merged with Google Brain in April 2023 to form Google DeepMind; produces the Gemini family of models. Alphabet subsidiary; uses Google Cloud TPU + accelerator infrastructure. Headquartered London, UK with research presence in Mountain View, Zurich, Paris, Montreal. Detail page: https://scrutica.com/companies/org-google-deepmind. ### xAI (`org-xai`) Frontier AI lab; produces the Grok family of models. Affiliated with Elon Musk. Compute-visibility tier 3 (private company); Colossus cluster (Memphis, TN) the named primary training facility — disclosed in stages from 100K H100s (2024) to 200K-class (2025) per public commentary. Acquired by Tesla shareholders in 2025 via xAI / Tesla equity-swap proposal. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-xai. ### Mistral (`org-mistral`) European frontier-AI lab. Headquartered Paris. Cloud-provider relationships include ECLAIRION (France, 13.8K GB300 / 44 MW announced for Q2-2026) and Scaleway. Compute-visibility tier 3. Detail page: https://scrutica.com/companies/org-mistral. ### Apple Inc (`org-apple`) Designs the M-series and A-series silicon used across Mac, iPhone, and iPad product lines, plus Private Cloud Compute infrastructure for Apple Intelligence. Custom in-house silicon strategy distinct from the NVIDIA-centric AI accelerator ecosystem; Apple's training compute substrate is partially disclosed (Apple Foundation Model published architecture details + per-cluster training hardware) and partially undisclosed. TSMC is the primary fabrication partner across the M and A series; advanced-node demand is a material fraction of TSMC's leading-edge capacity allocation. Headquartered in the United States. Public-market ticker: AAPL. Detail page: https://scrutica.com/companies/org-apple. ### Oracle (`org-oracle`) Operates Oracle Cloud Infrastructure (OCI) global cloud + the Stargate UAE joint venture (operational supercluster, 4,000+ NVIDIA Blackwell GPUs in OCI Abu Dhabi region as of November 2025) + the broader Stargate Project (Texas-anchored compute build-out with OpenAI as anchor tenant). Distant fourth among major US cloud providers by compute-substrate share but the most active in named sovereign-AI partnerships. Ticker ORCL. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-oracle. ### IBM (`org-ibm`) Designs Telum and Spyre AI accelerators for IBM Z and Power systems; operates IBM Cloud + Red Hat OpenShift. Long-tail role in research-grade AI infrastructure (Watson family historically) but limited share in the hyperscale-AI training market. Headquartered in the United States. Public-market ticker: IBM. Detail page: https://scrutica.com/companies/org-ibm. ### ARM (`org-arm`) Designs the ARM CPU instruction set architecture licensed to NVIDIA (Grace CPU), Apple (M-series), AWS (Graviton), and most mobile and embedded silicon designers. SoftBank-controlled (via Arm Holdings IPO 2023; SoftBank retains majority stake). Headquartered Cambridge, UK. Critical infrastructure for AI server design: Grace + Blackwell in the GB200 / GB300 NVL72 system architecture is the most-visible AI-specific ARM deployment. Public-market ticker: ARM. Detail page: https://scrutica.com/companies/org-arm. ### Cadence (`org-cadence`) One of two dominant electronic design automation (EDA) tool vendors (Cadence + Synopsys; Mentor / Siemens EDA as distant third). EDA tool access is a structural chokepoint in semiconductor design — the four-tool-vendor ecosystem (Cadence, Synopsys, Mentor/Siemens, Ansys for parasitic extraction) is the deepest non-fabrication chokepoint in the compute supply chain. Headquartered San Jose, CA. Public-market ticker: CDNS. Detail page: https://scrutica.com/companies/org-cadence. ### Synopsys (`org-synopsys`) Largest EDA vendor by revenue. Toolset covers logic synthesis, physical implementation, verification, IP licensing. Critical chokepoint position for advanced-node chip design — every NVIDIA, AMD, Apple, and TSMC-customer chip flows through Synopsys tools at some point in the design chain. Headquartered Sunnyvale, CA. Ticker SNPS. Detail page: https://scrutica.com/companies/org-synopsys. ### Ansys (`org-ansys`) EDA-adjacent simulation tool vendor (parasitic extraction, multiphysics simulation, signal integrity analysis). Critical for the design-verification step of advanced-node chip development. Synopsys announced acquisition of Ansys in early 2024; combination pending regulatory approval as of 2025-2026. Headquartered Canonsburg, PA. Detail page: https://scrutica.com/companies/org-ansys. ### GlobalFoundries (`org-global-foundries`) Major non-leading-edge foundry; HQ Malta, NY with fabs in NY (Fab 8), Vermont (Essex Junction), Singapore, Germany (Dresden). Focused on mature and specialty-process nodes (12nm and above); exited leading-edge race in 2018. Major CHIPS Act recipient. Ticker GFS. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-global-foundries. ### Canon (`org-canon`) Japanese photolithography equipment vendor; alternative supplier to ASML for non-leading-edge lithography. Announced nanoimprint lithography (FPA-1200NZ2C) as a potential alternative pathway to EUV for select applications; commercial uptake remains nascent as of 2025. Headquartered Tokyo. Ticker 7751.T. Detail page: https://scrutica.com/companies/org-canon. ### Nikon (`org-nikon`) Japanese photolithography equipment vendor; smaller and more specialized than Canon and ASML. Focus on mature-node lithography and specialty applications. Headquartered Tokyo. Ticker 7731.T. Detail page: https://scrutica.com/companies/org-nikon. ### ASE (`org-ase`) World's largest semiconductor assembly and test services (OSAT) provider. Headquartered Kaohsiung, Taiwan. ASE provides packaging services across mature and advanced packaging tiers (excluding the leading-edge CoWoS class which TSMC retains in-house). Ticker 3711.TW (also ADR ASX). Detail page: https://scrutica.com/companies/org-ase. ### Amkor (`org-amkor`) Second-largest OSAT after ASE. Major operations in South Korea, the Philippines, China, Portugal, and the US. Building a major advanced-packaging facility in Peoria, Arizona (CHIPS Act-supported) to host CoWoS-class capacity outside Taiwan. Headquartered Tempe, AZ. Ticker AMKR. Detail page: https://scrutica.com/companies/org-amkor. ### Advantest (`org-advantest`) World's largest semiconductor test-equipment supplier. HBM testing is a critical sub-segment — every HBM stack requires dedicated test capacity, and Advantest's V93000 platform is the dominant testbench for HBM3 and HBM3e. Headquartered Tokyo. Ticker 6857.T. Detail page: https://scrutica.com/companies/org-advantest. ### Teradyne (`org-teradyne`) Second-largest semiconductor test-equipment supplier after Advantest. Headquartered North Reading, MA. Ticker TER. Detail page: https://scrutica.com/companies/org-teradyne. ### Besi (`org-besi`) Dutch hybrid-bonding equipment specialist. Hybrid bonding is the next-generation advanced-packaging technology (post-CoWoS) being deployed for HBM4 and beyond. Headquartered Duiven, Netherlands. Ticker BESI.AS. Limited competition globally in the hybrid-bonding equipment class — a potentially emerging chokepoint Scrutica tracks under `/chokepoints/early-warning`. Detail page: https://scrutica.com/companies/org-besi. ### ASM International (`org-asm-international`) Dutch ALD (atomic layer deposition) equipment specialist. Critical position in the advanced-node deposition stack. Headquartered Almere, Netherlands. Ticker ASM.AS. Detail page: https://scrutica.com/companies/org-asm-international. ### JSR Corporation (`org-jsr-corporation`) Major Japanese photoresist materials supplier. Critical structural position in the advanced photoresist materials supply chain — photoresist is one of the deepest non-fabrication chokepoints in the compute supply chain because the chemistries are vendor-specific and qualification cycles are multi-quarter. JSR + TOK + Shin-Etsu + DuPont form the named-supplier oligopoly. Headquartered in Japan. Detail page: https://scrutica.com/companies/org-jsr-corporation. ### Tokyo Ohka Kogyo (`org-tokyo-ohka`) Japanese photoresist materials supplier (TOK Chemical). Smaller than JSR but structurally important in advanced photoresist (EUV resist). Headquartered Kawasaki, Japan. Ticker 4186.T. Detail page: https://scrutica.com/companies/org-tokyo-ohka. ### Shin-Etsu (`org-shin-etsu`) Japanese chemicals conglomerate; major position in silicon wafer manufacturing (Shin-Etsu Handotai is one of the two dominant wafer suppliers globally alongside SUMCO) and in advanced photoresist materials. Wafer supply is a structural sub-chokepoint that surfaces less than HBM but binds with comparable severity. Headquartered Tokyo. Ticker 4063.T. Detail page: https://scrutica.com/companies/org-shin-etsu. ### SUMCO (`org-sumco`) Major Japanese silicon wafer supplier; second-largest globally alongside Shin-Etsu Handotai. The 300mm wafer market is concentrated to ~5 suppliers globally. Headquartered Tokyo. Ticker 3436.T. Detail page: https://scrutica.com/companies/org-sumco. ### Air Liquide (`org-air-liquide`) French industrial gases major; supplies the high-purity gases required for semiconductor manufacturing (nitrogen trifluoride, tungsten hexafluoride, fluorine, neon for excimer lasers). Critical sub-tier supplier with structural concentration. Headquartered Paris. Ticker AI.PA. Detail page: https://scrutica.com/companies/org-air-liquide. ### Linde (`org-linde`) Industrial gases major (German/Irish-domiciled). Competitor to Air Liquide in the high-purity gas supply chain. Headquartered Guildford, UK with operations globally. Ticker LIN. Detail page: https://scrutica.com/companies/org-linde. ### Cerebras (`org-cerebras`) Wafer-scale AI accelerator vendor; produces the WSE-3 (Wafer Scale Engine, 3rd generation). Distinct architectural approach versus NVIDIA / AMD / Intel chiplet designs. Headquartered Sunnyvale, CA. Private company with disclosed deployments including the Condor Galaxy series of training systems. Detail page: https://scrutica.com/companies/org-cerebras. ### Groq (`org-groq`) Inference-optimized AI accelerator vendor; LPU (Language Processing Unit) tensor-streaming architecture. Headquartered Mountain View, CA. Private company; focused on inference-serving rather than training. Detail page: https://scrutica.com/companies/org-groq. ### SambaNova Systems (`org-sambanova`) Datacenter-class AI accelerator vendor; reconfigurable dataflow architecture. Headquartered Palo Alto, CA. Private company; selected enterprise / sovereign-AI deployments. Detail page: https://scrutica.com/companies/org-sambanova. ### Tenstorrent (`org-tenstorrent`) AI accelerator vendor; open-source-RISC-V-anchored architecture. Wormhole and Blackhole product generations. Headquartered Toronto, Canada. Private company; differentiated approach via RISC-V and open-source software stack. Detail page: https://scrutica.com/companies/org-tenstorrent. ### Graphcore (`org-graphcore`) British AI accelerator vendor; IPU (Intelligence Processing Unit) architecture. Headquartered Bristol, UK. Acquired by SoftBank in 2024 after a difficult commercial trajectory against NVIDIA's CUDA ecosystem dominance. Detail page: https://scrutica.com/companies/org-graphcore. ### HiSilicon (`org-huawei-hisilicon`) Chip design subsidiary of Huawei; designs the Kirin mobile SoC and the Ascend AI accelerator series. Subject to US BIS Entity List designation since 2019 with cascading restrictions in 2020 (US Foreign Direct Product Rule application to TSMC-fabricated HiSilicon chips). The chip-design entity behind the Ascend 910B / 920 series and the broader Chinese domestic-substitution narrative. Headquartered in China. Detail page: https://scrutica.com/companies/org-huawei-hisilicon. ### Cambricon (`org-cambricon`) Chinese AI accelerator design house; Cambricon-1A / 1H / 1M series and the Siyuan (思元) / Hanwu (寒武) GPU lines. Listed on the Shanghai STAR Market. The most-visible non-Huawei Chinese AI chip designer. Subject to selective US export-control attention. Headquartered in China. Detail page: https://scrutica.com/companies/org-cambricon. ### ByteDance (`org-bytedance`) Chinese internet conglomerate; parent of TikTok / Douyin and operator of significant in-house AI training and inference compute. Self-operates major data center facilities in China and contracts overseas regions via cloud providers. The most analytically-relevant non-Tencent / non-Alibaba Chinese hyperscaler. Headquartered in China. Detail page: https://scrutica.com/companies/org-bytedance. ### Alibaba (`org-alibaba`) Chinese hyperscaler (Alibaba Cloud / Aliyun) + e-commerce platform. Self-designs in-house silicon (T-Head Yitian server CPU; Hanguang AI accelerator). Operates the broadest non-US public-cloud substrate by APAC market share. Subject to US export controls limiting access to leading-edge NVIDIA SKUs. Headquartered in China. Public-market ticker: BABA. Detail page: https://scrutica.com/companies/org-alibaba. ### Tencent Holdings Ltd. (`org-tencent`) Chinese internet conglomerate; operates Tencent Cloud + WeChat ecosystem. Major in-house AI compute consumer (Hunyuan model family); self-operates significant facility footprint in mainland China. Subject to US export controls limiting access to leading-edge NVIDIA SKUs. Headquartered in China. Public-market ticker: 700. Detail page: https://scrutica.com/companies/org-tencent. ### IREN (`org-iren`) Iris Energy / IREN — pivoted from Bitcoin mining to AI cloud compute (renewable-energy-anchored Texas / Canada / British Columbia facilities). One of the named "neoclouds" alongside CoreWeave, Nebius, Applied Digital, Crusoe. Headquartered in Australia. Public-market ticker: IREN. Detail page: https://scrutica.com/companies/org-iren. ### Applied Digital (`org-applied-digital`) US-headquartered neocloud / data-center developer; operates ND (Ellendale, North Dakota) campus and other facilities. Pivoted from bitcoin / HPC hosting to AI cloud. Ticker APLD. Detail page: https://scrutica.com/companies/org-applied-digital. ### Crusoe (`org-crusoe`) US-headquartered neocloud; operates a facility footprint anchored on flared-natural-gas-powered sites in oil-and-gas regions. Distinctive grid-arbitrage strategy. Headquartered Denver, CO. Detail page: https://scrutica.com/companies/org-crusoe. ### Nebius (`org-nebius`) European neocloud; Yandex spinoff (Nebius Group N.V.). Operates GPU-cloud capacity primarily in Finland. Listed on NASDAQ. Ticker NBIS. Headquartered in the Netherlands. Detail page: https://scrutica.com/companies/org-nebius. ### Lambda Labs (`org-lambda-labs`) US-headquartered neocloud / GPU-cloud provider focused on AI training workloads. Headquartered San Francisco. Private company. Detail page: https://scrutica.com/companies/org-lambda-labs. ### Together AI (`org-together-ai`) US-headquartered neocloud / inference + training provider; emphasis on open-source model serving. Headquartered San Francisco. Private company. Detail page: https://scrutica.com/companies/org-together-ai. ### Cohere (`org-cohere`) Canadian frontier-AI lab; Command family of models. Headquartered Toronto. Private company; commercial cloud-provider relationships. Detail page: https://scrutica.com/companies/org-cohere. ### Inflection AI (`org-inflection-ai`) Frontier AI lab (Pi assistant). Acquired by Microsoft in March 2024 via the highly-publicized founder-and-team transition to Microsoft AI division; the residual Inflection AI entity continues as a separate corporate vehicle. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-inflection-ai. ### Super Micro Computer, Inc. (`org-supermicro`) Super Micro Computer — server-system vendor; assembles GPU-server rack configurations for NVIDIA + AMD accelerator chains. Major position in the AI-server-OEM market. Ticker SMCI. Subject to subpoena-related disclosure delays in 2024 that prompted Scrutica's editorial vintage flag. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-supermicro. ### Dell Technologies (`org-dell`) Dell Technologies — server-system vendor; competes with Supermicro and HPE in the GPU-server OEM market. Headquartered Round Rock, TX. Ticker DELL. Detail page: https://scrutica.com/companies/org-dell. ### HPE (`org-hpe`) Hewlett Packard Enterprise — server / HPC vendor; significant exposure via Cray supercomputing line and named sovereign-AI HPC contracts. Ticker HPE. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-hpe. ### Broadcom Inc (`org-broadcom`) Designs networking ASICs (Tomahawk Ethernet switching family is the dominant data-center networking fabric chip), custom AI accelerators for hyperscalers (Google TPU under Broadcom co-design relationship; recent disclosures suggest Meta + ByteDance custom-silicon engagements), and the AVAGO RF supply chain. Critical structural position in AI compute networking. Ticker AVGO. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-broadcom. ### Marvell Technology (`org-marvell`) Designs custom-silicon for hyperscalers (Amazon Trainium 2 and successor co-design relationship publicly disclosed), networking ASICs, and storage controllers. Ticker MRVL. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-marvell. ### MediaTek (`org-mediatek`) Taiwan-headquartered fabless chip designer; partnership with NVIDIA on the Tegra family and named role in NVIDIA-AI-PC roadmap. Distinct from the TSMC-aligned NVIDIA / AMD set; significant volume position in mobile and embedded AI silicon. Ticker 2454.TW. Detail page: https://scrutica.com/companies/org-mediatek. ### Qualcomm Inc (`org-qualcomm`) Designs Snapdragon mobile SoCs + Hexagon AI accelerator IP. Selected datacenter-AI initiatives (AI 100 family); primary market remains mobile / edge AI rather than hyperscale training. Ticker QCOM. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-qualcomm. ### Cloudflare, Inc. (`org-cloudflare`) Operates a globally-distributed edge compute platform + AI inference at the edge (Workers AI). Limited training-class compute but significant inference-deployment substrate; relevant to compute-visibility analyses of where AI inference physically runs. Ticker NET. Headquartered in the United States. Detail page: https://scrutica.com/companies/org-cloudflare. --- ## Sovereign AI programs (full detail) The 36 programs Scrutica tracks, with announced and committed figures, status, key partners, and the editorial data-quality flags surfacing known accounting ambiguities. The interactive dashboard is at https://scrutica.com/sovereign-ai. ### UAE: Stargate UAE / G42 / Core42 / TII Sovereign AI + Microsoft Cumulative + Abu Dhabi Digital Strategy - Country: `ae` · Status: procurement_active / construction_active / Oracle Supercluster operational (Nov 2025) - Announced (headline): $518.74B - Announced (government-only): $3.54B - Committed: $18.74B - Timeline: 2024–2029 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: Microsoft, OpenAI, NVIDIA, Oracle, Cisco, SoftBank, Technology Innovation Institute (TII), G42, Core42, Khazna Data Centers, Abu Dhabi Department of Government Enablement (DGE) - Source count: 11 Three concurrent UAE AI vehicles, separate decomposition rows: (1) Stargate UAE $500B aspirational, Phase 1 200 MW Q3 2026; (2) Microsoft cumulative $15.2B through 2029 — $1.5B Apr 2024 G42 equity (subsumed) + $4.6B+ datacenter capex + $1.2B+ local OpEx + forward pledges; 200 MW Khazna datacenter expansion online before end-2026; (3) Abu Dhabi Government Digital Strategy 2025-2027 $3.54B, with Oracle Sovereign AI Supercluster (4,000+ NVIDIA Blackwell GPUs in OCI Abu Dhabi region, operational Nov 2025) as the compute substrate. G42 divested Chinese tech holdings in 2024 as condition of Microsoft investment; received US chip export approval Nov 2025. MGX Fund I closed at $49B (2026-07-01) — investment vehicle context, not counted toward compute-capacity totals. ### France: France 2030 / AI Action Summit / GENCI / Mistral Compute / SoftBank Hauts-de-France - Country: `fr` · Status: operational (Jean Zay H100 extension inaugurated May 2025) / Mistral Compute operational (44 MW, June 2026) / SoftBank Hauts-de-France campus announced (Phase 1 firm, targeting 2031) - Announced (headline): $201.34B - Announced (government-only): $2.70B - Committed: $2.70B - Disbursed: $43.6M - Timeline: 2018–2031 - Interdependence: NVIDIA: high · TSMC: high · US: medium · BIS reach: low - Key partners: GENCI, Eviden (Atos), CNRS, OVHcloud, Scaleway, Mistral AI, ASML, SoftBank - Source count: 9 EUR 109B headline is private investment. Government cumulative since 2018 is EUR 2.5B. Jean Zay: 1,456 H100 + 416 A100 + 1,832 V100 operational. Mistral Compute: 13,800 GB300 GPUs in 44 MW Bruyères-le-Châtel (Essonne) DC, $830M debt financing, operational per 2026-06-18 reporting. Scaleway: first European Blackwell Ultra B300 provider. ASML largest shareholder in Mistral (11%). Ministry of Armed Forces selected Mistral for defense AI. SoftBank (2026-05-31): up to EUR75B/5GW aspirational, EUR45B/3.1GW Hauts-de-France (Dunkirk, Bosquel, Bouchain) firm Phase 1 by 2031. France 2030 (2026-06-16): additional EUR655M announced, not yet reconciled against the existing EUR2.5B cumulative figure. ### Saudi Arabia: HUMAIN / Project Transcendence - Country: `sa` · Status: procurement_active - Announced (headline): $100.00B - Announced (government-only): $77.00B - Timeline: 2025–2030 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: medium - Key partners: NVIDIA, AWS, Google Cloud, PIF, SDAIA, HUMAIN, xAI - Source count: 4 Phase 1: 18,000 NVIDIA GB300 operational. 600,000 GPU 3-year target (ambition). 1.9 GW DC capacity target by 2030, 500 MW AI computing target. xAI partnership: first customer for 500MW+ facility. $77B HUMAIN infrastructure strategy (PIF vehicle). $10B Google Cloud + PIF partnership. $23B in signed US-tech deals. HUMAIN-AWS AI Zone: up to 150,000 accelerators. HUMAIN-Cohere (2026-07-09): >=50MW allocation within the existing capacity target, Cohere's first ex-North-America deployment, targeted live Q4 2027. ### South Korea: National AI Initiative & K-Chips Act - Country: `kr` · Status: procurement_active - Announced (headline): $74.00B - Announced (government-only): $1.48B - Committed: $1.10B - Timeline: 2025–2030 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: NVIDIA, Samsung Electronics, SK Group, Hyundai, MSIT, Naver, SK Telecom, LG, NCSoft, Upstage, KEPCO - Source count: 10 ₩100 trillion headline conflates government R&D, private investment, and industrial transformation. $381M verified government-only for the sovereign AI model consortia program (5 selected Aug 2025: Naver Cloud, SK Telecom, LG AI Research, NC AI, Upstage; narrowed to 3 in Jan 2026 first cut: LG, SK Telecom, Upstage). NAICC: 10,080 B200 + 3,056 H200 = 13,136 GPUs (procurement_active, no operational partition yet). 52K GPUs by 2028, 260K by 2030 targets. 3 GW DC in Jeollanam-do. AI Basic Act effective January 2026. NAICC's SPC ('Korea AI Computing Center') completed registration ~2026-06-09/13 (Samsung SDS 30% / govt 29% / Naver Cloud 26.1%, CEO Ahn Jung-tae); Haenam groundbreaking still pending as of 2026-07-14. ### Japan: METI ¥10T AI Strategy / RIKEN FugakuNEXT / Rapidus - Country: `jp` · Status: procurement_active / operational (RIKEN spring 2026 target) - Announced (headline): $64.00B - Announced (government-only): $7.90B - Committed: $7.90B - Disbursed: $470.0M - Timeline: 2024–2030 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: NVIDIA, Fujitsu, Sakura Internet, KDDI, SoftBank, GMO, Highreso, RUTILEA, RIKEN, MEXT, Rapidus - Source count: 6 METI committed ¥10 trillion ($64B) public support through fiscal 2030. Target: 60 EFLOPS by fiscal 2027. FY2026 budget: ¥1.23 trillion ($7.9B) for semiconductors and domestic AI. Rapidus 2nm GAA milestone July 2025. RIKEN 2,140 Blackwell GPUs targeting spring 2026. FugakuNEXT: 600+ EFLOPS FP8 target ~2030. METI: $740M to 6 AI compute firms. SoftBank: JPY 150B (~$960M) AI infrastructure. ### China: National AI Programs / Big Fund III / Eastern Data Western Computing - Country: `cn` · Status: operational at scale / procurement_active - Announced (headline): $55.70B - Announced (government-only): $55.70B - Committed: $55.70B - Timeline: 2022–2025 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: high - Key partners: Huawei, Biren Technology, Alibaba, ByteDance, Tencent, Baidu, MIIT - Source count: 6 Most BIS-exposed program. 246 EFLOP/s baseline (June 2024), targeting 300 EFLOP/s by 2025. Big Fund III (344B yuan) + National AI Industry Fund (60B yuan) are key vehicles. Huawei Ascend is domestic GPU alternative at SMIC 7nm (~30K wafers/month Q2 2025, targeting 60K by end 2026; TrendForce tier 3). SMIC record $9.3B revenue in 2025. Many AI data centers stand unused. A further ~$295B (2 trillion yuan) nationwide AI data-center grid plan was reported by Bloomberg 2026-06-09 as still in NDRC drafting stage as of 2026-07-14 — not enacted, not counted in announced_usd. ### EU-Wide: EuroHPC Joint Undertaking - Country: `eurohpc` · Status: JUPITER operational (1 ExaFLOP, #4 TOP500, Nov 2025); 19 AI Factory sites across 3 waves - Announced (headline): $10.96B - Announced (government-only): $10.96B - Timeline: 2021–2027 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: EU Commission, 35 Member/Associated Countries, NVIDIA, AMD, HPE, Eviden - Source count: 7 JUPITER operational at 1 ExaFLOP (#4 TOP500, November 2025), propelling Europe into the exascale era. 19 AI Factory sites across 3 waves. Wave 1 (Dec 2024): 7 sites. Wave 2 (Mar 2025): 6 sites. Wave 3 (Oct 2025): Czech Republic, Lithuania, Netherlands, Romania, Spain, Poland. 13 AI Factory Antennas selected (EUR 55M). EUR 10B covers 2021-2027 including pre-AI HPC. LISA upgrade (1,328 H100 GPUs, EUR 50M) inaugurated 2026-06-11 at Leonardo (Bologna), complementing IT4LIA. ### Vietnam: National AI Strategy & NVIDIA Partnership - Country: `vn` · Status: planning - Announced (headline): $7.00B - Announced (government-only): $38.5M - Committed: $38.5M - Timeline: 2025–2027 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: medium - Key partners: FPT Corporation, NVIDIA, Ministry of Science and Technology, Vietnam National Innovation Center - Source count: 6 No verified sovereign GPU deployment. NVIDIA R&D center (VRDC) cooperation agreement December 2024. First AI law in Southeast Asia passed December 10, 2025, effective March 1, 2026. $7B+ in private AI DC investments (FPT, etc.). Government NDDF: 1T VND (~$38.5M). ### Taiwan: AI New Ten Major Construction + NCHC Sovereign Compute - Country: `tw` · Status: operational (NCHC Tainan system) / budget_stalled (2026 central budget frozen in Legislative Yuan) - Announced (headline): $6.20B - Announced (government-only): $6.20B - Committed: $2.80B - Timeline: 2025–2029 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: medium - Key partners: NCHC, NSTC, NDC, NVIDIA, AMD, Hon Hai (Foxconn), Wistron NeWeb, Visionbay.ai - Source count: 6 Over 1,700 H200 GPUs + GB200 + B300 operational at Tainan. 480 PFLOPS target by 2029. Taiwan Compute Alliance includes Foxconn, AMD, NVIDIA. TAIDE national Mandarin LLM in development. Ironic TSMC dynamic: world's dominant foundry is headquartered here, yet the government must buy NVIDIA GPUs fabricated by its own national champion. 2026 budget crisis: NT$30B+ first-year AI funding frozen in Legislative Yuan amid KMT/TPP opposition blockade. AI Basic Act passed Dec 2025 but implementation funding remains uncertain. ### Germany: National AI Strategy / Gaia-X / Gauss Centre for Supercomputing - Country: `de` · Status: funding_approved / procurement_active / operational (existing GCS sites) - Announced (headline): $6.00B - Announced (government-only): $3.70B - Committed: $3.70B - Disbursed: $3.70B - Timeline: 2019–2030 - Interdependence: NVIDIA: high · TSMC: high · US: medium · BIS reach: low - Key partners: BMBF, BMWK, BMAS, Eviden, NVIDIA, Deutsche Telekom, Gaia-X Association - Source count: 5 EUR 3.38B directed to specific projects by June 2024. Blue Lion (EUR 250M) and HLRS (EUR 85M) under procurement. Germany's AI infrastructure ambition outpaces private sector follow-through. ### United Kingdom: AI Opportunities Action Plan / UK Compute Roadmap / AIRR / Sovereign AI Unit / AI Hardware Plan - Country: `gb` · Status: operational (Isambard-AI Phase 1 launched July 2025) / procurement_active (Edinburgh) / fund_launching (Sovereign AI Unit April 2026) - Announced (headline): $4.75B - Announced (government-only): $4.75B - Committed: $4.75B - Disbursed: $318.0M - Timeline: 2025–2030 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: NVIDIA, Nscale, CoreWeave, Microsoft, British Business Bank, EPCC, University of Bristol, Balderton Capital (James Wise, Sovereign AI Unit chair), AI Safety Institute - Source count: 8 From 21 AI ExaFLOPS (2025) to 420 AI ExaFLOPS by 2030 (20x). Isambard-AI full system (5,448 GH200, 5MW) launched July 17 2025 — the earlier 168-GH200 Phase 1 (TOP500 ISC June 2024 #128) was the initial March 2024 partition that was integrated into the full system, not a separate ongoing operational deployment. £225M total Isambard-AI investment. AI Safety Institute confirmed as beneficiary. Edinburgh supercomputer (up to GBP 750M) expected online 2027. Sovereign AI Unit (£500M venture fund, chaired by James Wise of Balderton Capital) launching April 16 2026. AI Hardware Plan (GBP 1.1B, 2026-06-08): GBP 750M supercomputer/chips + GBP 120M Hardware Innovation Programme + GBP 45M skills + GBP 150M investment fund. ### Malaysia: National AI Roadmap 2030 / Sovereign AI Cloud - Country: `my` · Status: planning (National AI Roadmap 2030) / funding_approved (Sovereign AI Cloud) - Announced (headline): $4.30B - Announced (government-only): $440.0M - Committed: $440.0M - Timeline: 2024–2030 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: National AI Office (NAIO), YTL Power International, NVIDIA, Microsoft - Source count: 5 National AI Roadmap 2030: RM 20B ($4.3B) total. Government: RM 2B Sovereign AI Cloud (2026 budget). NAIO established December 2024. National AI Action Plan 2026-2030: target top 20 global AI readiness by 2030. Year 2 milestones: RM500M startup fund, 5,000 AI experts, 100-factory AI adoption. Private: YTL/NVIDIA $2.36B facility with GB200 GPUs, 600 MW. ### Brazil: Brazilian Artificial Intelligence Plan (PBIA) 2024-2028 - Country: `br` · Status: legislation_passed / operational (Santos Dumont upgrade completed ~July 2025) - Announced (headline): $4.20B - Announced (government-only): $1.55B - Committed: $4.20B - Timeline: 2024–2028 - Interdependence: NVIDIA: high · TSMC: high · US: medium · BIS reach: low - Key partners: LNCC, Eviden (Atos Group), BNDES, FNDCT, MCTI, Inria Brasil - Source count: 8 Santos Dumont upgrade genuinely complete at 18.85 PFLOPS. R$23B headline overstates direct government outlay; R$12.72B of that is returnable credit from FNDCT and BNDES. ### Canada: Canadian Sovereign AI Compute Strategy / Pan-Canadian AI Strategy / AI for All - Country: `ca` · Status: funding_approved / procurement_active - Announced (headline): $2.60B - Announced (government-only): $2.60B - Committed: $2.60B - Timeline: 2024–2030 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: none - Key partners: Digital Research Alliance of Canada, CIFAR, ISED, Canada Infrastructure Bank, NRC - Source count: 8 C$2B committed in Budget 2024 across three pillars: up to C$705M for public supercomputer (AI Sovereign Compute Infrastructure Program), C$700M for commercial DC incentives (AI Compute Challenge), C$300M for SME access (AI Compute Access Fund). Budget 2025 added C$925.6M over 5 years (C$800M from Budget 2024 set-asides). Canada Infrastructure Bank empowered to invest in AI infrastructure. No GPU counts in official documents. 'AI for All' national strategy launched 2026-06-04: AI Compute Access Fund topped up C$300M->C$1B (C$700M new); 850MW->2.3GW sovereign-compute capacity PROPOSED (partnerships being finalized, not yet committed). ### United States: NAIRR / CHIPS Act AI Components / Stargate Project - Country: `us` · Status: operational (NAIRR Pilot) / legislation_pending (full NAIRR) - Announced (headline): $2.60B - Announced (government-only): $30.0M - Committed: $30.0M - Disbursed: $30.0M - Timeline: 2024–? - Interdependence: NVIDIA: high · TSMC: high · US: low · BIS reach: low - Key partners: NSF, DOE, Microsoft, NVIDIA, Google, Amazon, Meta - Source count: 6 NAIRR Pilot: $30M/year, ~3.77 exaFLOPS, ~600 projects. Full NAIRR estimated at $2.6B but not appropriated. CREATE AI Act pending. Stargate $500B is private. CHIPS Act $52.7B is semiconductor manufacturing. Stargate Michigan (Saline Township) broke ground 2026-06-01, 1GW/~$16B initial — private, not counted in this row's government figures. ### Spain: BSC AI Factory + MareNostrum 5 + AI Strategy 2024 - Country: `es` · Status: operational (MareNostrum 5, BSC AI Factory) / procurement_active (EUR 129M expansion) - Announced (headline): $1.64B - Disbursed: $165.0M - Timeline: 2024–2025 - Interdependence: NVIDIA: medium · TSMC: high · US: medium · BIS reach: low - Key partners: BSC-CNS, EuroHPC JU, Spanish Government, Catalan Government, FSAS Technologies, Telefonica, Fujitsu - Source count: 8 MareNostrum 5 over 450 PFLOPS. BSC AI Factory operational. Mixed architecture (AMD, Intel, not exclusively NVIDIA). ALIA project training Europe's largest public AI model on MareNostrum 5. ### India: IndiaAI Mission - Country: `in` · Status: operational - Announced (headline): $1.25B - Announced (government-only): $1.25B - Committed: $1.25B - Disbursed: $48.0M - Timeline: 2024–2029 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: MeitY, Digital India Corporation, NVIDIA, Google, G42, L&T, Yotta - Source count: 12 ₹10,300 crore (~$1.25B) approved. 38,000 GPUs operational + 1,050 Google Trillium TPUs (3rd tender) — reconfirmed current as of a 2026-07-13 digitalindia.gov.in release. G42 8 exaflop supercomputer partnership. Target: 100K GPUs by end 2026 (3x current). L&T: 30 MW Chennai + 40 MW Mumbai NVIDIA AI factory. Yotta: $2B, 20,000+ Blackwell Ultra GPUs. Adani-Jabil (2026-06-15): non-binding INTENT to form an AI-hardware manufacturing alliance in India (GW-scale rack/server manufacturing for export) — not yet a signed deal, no alliance-specific dollar figure disclosed. ### Singapore: National AI Strategy 2.0 / AISG / NSCC Sovereign Compute - Country: `sg` · Status: operational (ASPIRE 2A+, 20 PFLOPS) / procurement_active (S$270M next-gen, operational late 2025) - Announced (headline): $740.0M - Announced (government-only): $740.0M - Committed: $200.0M - Disbursed: $200.0M - Timeline: 2023–2030 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: NRF, NSCC Singapore, NVIDIA, MDDI, AISG - Source count: 6 ASPIRE 2A+ operational (H100-based, 20 petaflops). S$270M next-gen supercomputer (classical + quantum, operational late 2025). National AI Strategy 2.0 (2025-2030): S$1B+ for national AI R&D. Enterprise Compute Initiative: S$150M cloud credits. 300 MW additional DC capacity allocated. ### Italy: IT4LIA AI Factory + Leonardo + CINECA - Country: `it` · Status: operational (Leonardo + LISA + IT4LIA since September 2025) - Announced (headline): $469.0M - Committed: $469.0M - Disbursed: $261.6M - Timeline: 2022–? - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: CINECA, EuroHPC JU, MUR, ACN, INFN, Leonardo S.p.A., IQM - Source count: 6 Leonardo: 15,000+ GPUs, 250+ PFLOPS, #9 globally. LISA: 1,328 Hopper GPUs. IT4LIA AI Factory operational since Sept 2025. IQM 54-qubit quantum computer delivery Q4 2025. ### Mexico: Coatlicue Supercomputer + National AI Plan - Country: `mx` · Status: announced_unbuilt (24-month build from Jan 2026; zero hardware deployed) - Announced (headline): $327.0M - Announced (government-only): $327.0M - Committed: $327.0M - Timeline: 2026–2028 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: Barcelona Supercomputing Center (BSC), NVIDIA - Source count: 5 Coatlicue: 314 PFLOPS target, MXN 6B ($327M), 24-month build. Interim access via BSC. Entirely unbuilt as of March 2026. National AI Agenda proposed but not enacted. ### Norway: Sigma2/NRIS National Research Infrastructure + LUMI-AI Participation - Country: `no` · Status: operational (Olivia delivered November 2025) / funding committed (LUMI-AI, operational target 2027) - Announced (headline): $317.0M - Announced (government-only): $317.0M - Disbursed: $20.9M - Timeline: 2025–2027 - Interdependence: NVIDIA: medium · TSMC: high · US: medium · BIS reach: low - Key partners: Sigma2 AS, HPE, Norwegian Research Council, CSC (Finland), Universities of Bergen, Oslo, Tromso, NTNU - Source count: 7 Olivia (NOK 225M, HPE) delivered November 2025 at Lefdal Mine Datacenter. LUMI-AI participation NOK 215M. NOK 3.4B is total planning figure. NBIM Oil Fund AI investments are portfolio, not compute. ### Australia: National AI Plan + GovAI Sovereign Compute - Country: `au` · Status: legislation_passed (National AI Plan Dec 2025) / no_government_compute_procurement (GovAI is cloud hosting, not GPU ownership; private sector building at scale) - Announced (headline): $280.0M - Announced (government-only): $24.0M - Committed: $24.0M - Timeline: 2025–? - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: CSIRO, Department of Industry, Science and Resources, OpenAI, NextDC, AWS, Microsoft, NCI, Pawsey Supercomputing Research Centre - Source count: 6 National AI Plan is primarily governance, adoption, and ecosystem strategy. No dedicated sovereign compute program comparable to Gulf states. GovAI is a hosting platform, not a GPU procurement program. ### Netherlands: Dutch AI Factory (AIFNL) + Snellius/EuroHPC - Country: `nl` · Status: funding_approved — operational target 2027 - Announced (headline): $218.0M - Announced (government-only): $141.7M - Committed: $218.0M - Timeline: 2025–2027 - Interdependence: NVIDIA: high · TSMC: high · US: medium · BIS reach: low - Key partners: AIFNL Foundation, SURF, TNO, EuroHPC JU, Samenwerking Noord - Source count: 6 EUR 200M total (national + regional + EuroHPC). Groningen site. Operational target 2027. Existing Snellius approaching end of life. ASML is Dutch but no domestic chip fab. ### Israel: National AI Program (Telem) & Supercomputer - Country: `il` · Status: operational - Announced (headline): $136.0M - Announced (government-only): $136.0M - Committed: $140.0M - Disbursed: $140.0M - Timeline: 2025–2027 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: Israel Innovation Authority, Nebius, NVIDIA - Source count: 6 1,000 NVIDIA B200 GPUs operational. Supercomputer launched May 2025 via competitive tender. 70% high-tech / 30% academic access. Phase 2 (2025-2027): NIS 500M budget including National AI Research Institute. ### Switzerland: Alps Supercomputer + Swiss AI Initiative - Country: `ch` · Status: operational - Announced (headline): $113.0M - Announced (government-only): $113.0M - Committed: $135.6M - Disbursed: $113.0M - Timeline: 2024–? - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: CSCS, ETH Zurich, EPFL, HPE, NVIDIA - Source count: 5 Alps operational, 434.9 PFLOPS peak, #7 TOP500. Over 10,000 GH200 chips. 10 MW power. CHF 37M/year operating budget. Apertus open LLM in development. ### Czech Republic: CZAI AI Factory + IT4Innovations + KarolAIna - Country: `cz` · Status: selected (October 2025 EuroHPC fourth wave) / procurement_active - Announced (headline): $46.0M - Announced (government-only): $23.0M - Committed: $46.0M - Timeline: 2025–2027 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: VSB-Technical University of Ostrava, IT4Innovations, VUT Brno, Charles University, CVUT, EuroHPC JU - Source count: 6 KarolAIna: ~340 AI chips delivering 850 PFLOPS AI performance (FP8). Part of EuroHPC fourth wave. Builds on existing Karolina petascale system. ### Ireland: CASPIr Supercomputer (EuroHPC) + AIF IRL-Antenna - Country: `ie` · Status: announced / procurement_active - Announced (headline): $27.3M - Announced (government-only): $17.7M - Committed: $27.3M - Timeline: 2025–2027 - Interdependence: NVIDIA: medium · TSMC: high · US: high · BIS reach: low - Key partners: ICHEC, University of Galway, EuroHPC JU, Department of Enterprise, Trade and Employment - Source count: 6 CASPIr: >15 PFLOPS. Ireland's first EuroHPC-class system. Tender published March 27, 2026. Uniquely exposed: enormous AI compute on its soil, entirely foreign-owned. ### Africa (Regional): Kenya, Rwanda, Nigeria + Continental Initiatives - Country: `africa-regional` · Status: strategy_published / no_sovereign_compute - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: UNDP, Cassava Technologies, NVIDIA, Google, AWS, Microsoft Azure - Source count: 6 No operational sovereign AI compute programs with disbursed funding in any African country. Patchwork of national AI strategies (mostly policy documents), one major private consortium investment (Cassava/NVIDIA $720M), and foreign-funded infrastructure. ### Finland: LUMI + LUMI-AI (EuroHPC AI Factory) - Country: `fi` · Status: operational (LUMI) / announced / procurement_active (LUMI-AI) - Announced (government-only): $272.5M - Disbursed: $157.5M - Timeline: 2022–2027 - Interdependence: NVIDIA: low · TSMC: high · US: medium · BIS reach: low - Key partners: CSC, EuroHPC JU, SRV, Finnish Government - Source count: 7 LUMI operational, #9 TOP500, ~380 PFLOPS Linpack, 10,752 AMD MI250X GPUs. Over half of resources used for AI research. LUMI-AI successor expected spring 2027. Notable for using AMD, not NVIDIA. ### Indonesia: National AI Roadmap + Danantara Sovereign AI Fund (Planned) - Country: `id` · Status: announced / strategy_published - Timeline: 2025–2029 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: BRIN, KOMDIGI, Danantara, BDx Indonesia - Source count: 5 Danantara SWF ($20B initial, AI is one of five priority sectors) established Feb 2025 but no specific AI compute allocation. Sovereign AI Fund planned for 2027-2029. No government compute infrastructure as of March 2026. ### NATO/AUKUS: Collective AI Compute Initiatives - Country: `nato-aukus` · Status: no_shared_compute_program - Interdependence: NVIDIA: low · TSMC: low · US: low · BIS reach: low - Key partners: NATO DIANA, NATO Innovation Fund, AUKUS AIA Working Group - Source count: 6 Neither NATO nor AUKUS has shared AI compute. DIANA (23 accelerator sites, 182 test centers) is innovation testing, not compute. NATO Innovation Fund (EUR 1B/15yr) is passive VC. AUKUS RAAIT trials test operational AI, not infrastructure. ### Philippines: NAICRI (National AI Center for Research and Innovation) + DOST-ASTI COARE/Saliksik HPC - Country: `ph` · Status: naicri_launched (2026-02-26) / coare_saliksik_operational - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: high - Key partners: DOST (Department of Science and Technology), DOST-ASTI (Advanced Science and Technology Institute), NAICRI, COARE (Computing and Archiving Research Environment), PREGINET, NVIDIA - Source count: 3 The Philippines' sovereign-AI position is institutional-first, compute-modest. (1) NAICRI, launched 2026-02-26 at the Manila Hotel, is 'the country's central institutional anchor for AI research, advanced computing, and innovation—providing continuity beyond individual projects and funding cycles', addressing 'fragmented infrastructure, limited data availability' (DOST-ASTI's own event page); 'NAICRI is anchored at DOST-ASTI but designed as a national center' (PIA). (2) The compute layer is DOST-ASTI's COARE; its Saliksik HPC comprises 36 CPU nodes (88 logical CPUs, 500 GB RAM each), 6 NVIDIA Tesla P40 nodes (1 GPU each), and 2 NVIDIA Tesla A100 nodes (8 GPUs each, 128 CPUs, 1 TB RAM) — 22 GPUs total, per the operator's own wiki. (3) Adjacent operational platforms showcased at the NAICRI launch: PREGINET (research/education network), NAIRA (AI-as-a-Service hub), DIMER (shared AI-model repository). COARE has powered FASSSTER pandemic surveillance, IRRI rice genome sequencing, climate modeling, and hazard mapping. No consolidated budget figure published at tier 1-2. ### Poland: PIAST-AI Factory (EuroHPC) + PSNC - Country: `pl` · Status: selected (March 2025) / procurement_active - Timeline: 2025–2026 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: PSNC, Poznan University of Technology, Adam Mickiewicz University, Nicolaus Copernicus University, EuroHPC JU, Polish Ministry of Digital Affairs - Source count: 5 PIAST-AI selected in EuroHPC second wave (March 2025). PSNC Poznan is host. Poland also participates in LUMI-AI consortium. Quantum computing education PLN 10M+ separate. ### Sweden: Berzelius + MIMER AI Factory (EuroHPC) - Country: `se` · Status: operational (Berzelius) / procurement_active (MIMER selected December 2024) - Announced (government-only): $43.0M - Timeline: 2024–2026 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: low - Key partners: Knut and Alice Wallenberg Foundation, Linköping University, Eviden, AI Sweden, EuroHPC JU, Swedish Research Council - Source count: 5 Berzelius: 880 NVIDIA H200 GPUs, 110 DGX nodes, operational. Funded by Wallenberg Foundation (private). MIMER EuroHPC AI Factory selected Dec 2024. Enterprise AI factory (Ericsson/AstraZeneca/SAAB) is private. ### Thailand: National AI Strategy and Action Plan (2022-2027) + ThaiSC/LANTA + ThaiLLM - Country: `th` · Status: strategy_active (2022-2027) / lanta_operational (since Nov 2022) / thaillm_launched (2026-04) - Timeline: 2022–2027 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: high - Key partners: NSTDA (National Science and Technology Development Agency), ThaiSC (NSTDA Supercomputer Center), NECTEC, MHESI (Ministry of Higher Education, Science, Research and Innovation), MDES (Ministry of Digital Economy and Society), Big Data Institute (BDI), Digital Economy and Society Development Fund (DEF), HPE (LANTA vendor), NVIDIA - Source count: 5 Thailand's sovereign-AI compute position rests on three verified pillars. (1) ThaiSC, the NSTDA Supercomputer Center, operates LANTA: a 346-node heterogeneous HPE Cray EX system, 8.15 PFlop/s peak, 176 GPU nodes with 704 NVIDIA A100 GPUs, 160 CPU nodes (20,480 cores), 10 high-memory nodes, 200 Gbps HPE Slingshot interconnect — operational since November 2022 (ThaiSC's own dating), located at Thailand Science Park, Pathum Thani. TOP500 70th at debut (2022, top-ranked in ASEAN at the time) declining to 199th (2025) per OECD.AI. (2) The National AI Strategy and Action Plan (2022-2027) anchors the program; its official portal (ai.in.th) lists ThaiSC/LANTA as a national AI resource. (3) ThaiLLM, launched 2026-04-01 and branded by NSTDA as 'A Sovereign AI Foundation for the Nation': Thai-language foundation models (8B and 30B parameters, trained on >100B tokens of Thai data) running on ThaiSC infrastructure, free via thaillm.or.th, developed by MDES + BDI + MHESI + NECTEC with DEF funding. This is a genuine government-owned compute + sovereign-model program, structurally distinct from commercial hyperscaler tenancy. No consolidated budget figure is published at tier 1-2. ### Ukraine: National LLM / Diia AI / NVIDIA Sovereign AI Partnership - Country: `ua` · Status: development_active (LLM training) / beta_planned (spring 2026) - Timeline: 2025–2030 - Interdependence: NVIDIA: high · TSMC: high · US: high · BIS reach: none - Key partners: Kyivstar (VEON), Google (Gemma 3, Cloud/Vertex AI), NVIDIA, Ministry of Digital Transformation, WINWIN AI Center of Excellence - Source count: 5 Ukraine's sovereign AI program is structurally unique: no government compute budget, entirely financed by Kyivstar (VEON subsidiary, Ukraine's largest telco). LLM built on Google Gemma 3, trained on Google Cloud Vertex AI outside Ukraine (wartime energy constraints make domestic training impractical). NVIDIA partnership (Nov 2025) covers infrastructure roadmap, talent, R&D, and startup ecosystem — but no disclosed GPU procurement or capital commitment. Beta planned spring 2026. WINWIN AI Center of Excellence coordinates. Diia AI LLM tailored to Ukrainian legislation, public services, cultural context. Model to be transferred to state and open-sourced post-validation. Fedorov frames as national security: data sovereignty in wartime. --- ## Reference: top compute facilities by disclosed capacity The compute-substrate's highest-capacity facilities by publicly-disclosed power capacity. The list below is sourced from `data/processed/facilities.json` (the canonical curated subset of Epoch Frontier Data Centers + selected sovereign + selected hyperscaler facilities; the full 4,550 -facility table on https://scrutica.com/entities is larger and more inclusive of mature-and-below-frontier sites). Each card pairs the disclosed power envelope with H100-equivalent GPU count where derivable, operator and ownership attribution, status, and the editorial / source narrative carried in the substrate row. Per-facility detail pages at https://scrutica.com/facilities/{id} render the same fields plus the FLOP-engine three-path estimate and any data quality flags applied to the facility. ### 1 Gig Data Center East Fishkill, NY (`gridstatus-nyiso-1738`) NY, United States — 1000 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1738. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1738. ### Colossus 2 (`fac-epoch-xai-colossus-2`) Memphis, Tennessee, United States — 946 MW · H100-equivalent 1,111,672 · type: ai_training · status: operational · operator `org-xai`. Editorial notes: Total capital cost (2025 USD): $35.84B | Epoch owner label changed 'xAI' -> 'SpaceXAI' (2026-07). Denotes the same xAI entity: X.AI Holdings Corp. was acquired by SpaceX effective 2026-02-02 per SpaceX SEC Form S-1 (filed 2026-05-20, 'the xAI Merger'); 'SpaceXAI' is the post-merger brand (press, 2026-07-06/07). Mapped to org-xai; alias recorded in organization_aliases. | Epoch estimation methods: our cooling equipment power model, company disclosures, drone imagery.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-xai-colossus-2. ### Anthropic-Amazon New Carlisle (`fac-epoch-anthropic-amazon-new-carlisle`) New Carlisle, Indiana, United States — 910 MW · H100-equivalent 687,215 · type: ai_training · status: operational · operator `org-amazon`. Editorial notes: Total capital cost (2025 USD): $34.47B | Epoch estimation methods: regulatory filings, chip efficiency, third-party reporting.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-anthropic-amazon-new-carlisle. ### Microsoft Fairwater Atlanta (`fac-epoch-microsoft-fairwater-atlanta`) Fayetteville, Georgia, United States — 636 MW · H100-equivalent 768,064 · type: ai_training · status: operational · operator `org-microsoft`. Editorial notes: Total capital cost (2025 USD): $24.09B | Epoch estimation methods: our cooling equipment power model, drone imagery, company disclosures, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-microsoft-fairwater-atlanta. ### Meta Prometheus (`fac-epoch-meta-prometheus`) New Albany, Ohio, United States — 631 MW · H100-equivalent 763,011 · type: ai_training · status: operational · operator `org-meta`. Editorial notes: Total capital cost (2025 USD): $23.90B | Epoch estimation methods: our cooling equipment power model, regulatory filings, chip availability, chip efficiency, land use maps.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-meta-prometheus. ### Micron Fab 2 (`gridstatus-nyiso-1627`) NY, United States — 576 MW · type: hyperscale_dc · status: announced · owner `org-micron`. Editorial notes: Queue ID: 1627. Gen type: Load. Proposed COD: nan. | Owner linkage derived 2026-07-15: queue project name ('Micron Fab 2') and county (Onondaga) match Micron Technology's White Pine Commerce Park memory-fab campus (Town of Clay, Onondaga County, NY; up to four fabs per phase), per the NY Governor's office announcement of 2022-10-04. https://www.governor.ny.gov/news/hochul-schumer-mcmahon-announce-micron-coming-onondaga-county-micron-will-invest-unprecedented. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1627. ### TeraWulf Lake Mariner Campus (`fac-epoch-fluidstack-lake-mariner`) Barker, New York, United States — 510 MW · type: ai_training · status: operational · operator `org-terawulf`. Editorial notes: Owner corrected from Fluidstack to TeraWulf per Epoch substrate description + audit 2026-05-15 (B4a/B8); preserved through the 2026-07-10 Epoch re-ingest. | CONTRADICTION (differing scopes): Epoch satellite-measured AI-building power 100 MW (2026-07) vs 510 MW disclosed multi-tenant total — Epoch tracks only the AI buildings (3-5, Fluidstack lease); 510 MW spans both tenant blocks incl. Core42's buildings 1-2. | H100-equivalent count kept null: Epoch's ~161,192 figure is anchored to the 100 MW scope. Capacity snapshots carry Epoch's per-date satellite measurements. | Total capital cost (2025 USD): $3.79B. | Epoch estimation methods: our cooling equipment power model, company disclosures, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-fluidstack-lake-mariner. ### North East Data LLC Data Center (`gridstatus-nyiso-1741`) NY, United States — 500 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1741. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1741. ### Arsenal Data Site 1000 (`gridstatus-nyiso-1730`) NY, United States — 467 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1730. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1730. ### OpenAI Stargate Abilene (`fac-epoch-openai-stargate-abilene`) Abilene, Texas, United States — 421 MW · H100-equivalent 510,358 · type: ai_training · status: operational · operator `org-oracle`. Editorial notes: Total capital cost (2025 USD): $15.95B | Epoch estimation methods: our cooling equipment power model, company disclosures, regulatory filings, drone imagery.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-openai-stargate-abilene. ### Microsoft Fairwater Wisconsin (`fac-epoch-microsoft-fairwater-wisconsin`) Mount Pleasant, Wisconsin, United States — 369 MW · H100-equivalent 445,679 · type: ai_training · status: operational · operator `org-microsoft`. Editorial notes: Total capital cost (2025 USD): $13.98B | Epoch estimation methods: company disclosures, regulatory filings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-microsoft-fairwater-wisconsin. ### Colossus 1 (`fac-epoch-xai-colossus-1`) Memphis, Tennessee, United States — 340 MW · H100-equivalent 275,795 · type: ai_training · status: operational · operator `org-xai`. Editorial notes: Total capital cost (2025 USD): $12.88B | Epoch owner label changed 'xAI' -> 'SpaceXAI' (2026-07). Denotes the same xAI entity: X.AI Holdings Corp. was acquired by SpaceX effective 2026-02-02 per SpaceX SEC Form S-1 (filed 2026-05-20, 'the xAI Merger'); 'SpaceXAI' is the post-merger brand (press, 2026-07-06/07). Mapped to org-xai; alias recorded in organization_aliases. | Epoch estimation methods: our cooling equipment power model, company disclosures, regulatory filings.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-xai-colossus-1. ### Google New Albany (`fac-epoch-google-new-albany`) New Albany, Ohio, United States — 339 MW · H100-equivalent 351,692 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $12.84B | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-new-albany. ### Google Columbus (`fac-epoch-google-columbus`) Columbus, Ohio, United States — 303 MW · H100-equivalent 416,877 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $11.48B | Epoch estimation methods: our cooling equipment power model, regulatory filings, chip efficiency, chip availability.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-columbus. ### Data & Technology Campus (`gridstatus-nyiso-1726`) NY, United States — 300 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1726. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1726. ### New York State Artificial Intelligence Data Center (`gridstatus-nyiso-1731`) NY, United States — 300 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1731. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1731. ### Amazon Madison Mega Site (`fac-epoch-amazon-madison-mega-site`) Canton, Mississippi, United States — 284 MW · H100-equivalent 214,348 · type: ai_training · status: operational · operator `org-amazon`. Editorial notes: Total capital cost (2025 USD): $10.76B | Epoch estimation methods: regulatory filings, chip efficiency, third-party reporting.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-amazon-madison-mega-site. ### CoreWeave Denton TX (`fac-epoch-coreweave-denton-tx`) Denton, Texas, United States — 262 MW · H100-equivalent 252,652 · type: ai_training · status: operational · operator `org-coreweave`. Editorial notes: Total capital cost (2025 USD): $9.93B | Epoch estimation methods: company disclosures.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-coreweave-denton-tx. ### Lake Mariner Data II (`gridstatus-nyiso-1670`) NY, United States — 250 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1670. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1670. ### Wulf Compute Data Center II (`gridstatus-nyiso-1732`) NY, United States — 250 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1732. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1732. ### Pontoon Bridge Road Data Center (`gridstatus-nyiso-1745`) NY, United States — 250 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1745. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1745. ### Broome County Tech Park (`gridstatus-nyiso-1752`) NY, United States — 250 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1752. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1752. ### Huawei Horinger (`fac-epoch-huawei-horinger`) Hohhot, Inner Mongolia, China — 241.8 MW · H100-equivalent 123,799 · type: ai_training · status: operational · operator `org-huawei`. Editorial notes: Total capital cost (2025 USD): $9.16B | Distinct campus from fac-huawei-ulanqab (Huawei Cloud Ulanqab DC): Horinger New Area is under Hohhot, ~133-145 km from Ulanqab city (separate prefecture-level cities); Horinger groundbreaking 2023-10 (Xinhua 2023-11-29) vs Ulanqab campus construction from 2017. Convergent tier-3 evidence; verified 2026-07-10. | Epoch estimation methods: our cooling equipment power model, company disclosures, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-huawei-horinger. ### DayOne Nusajaya (`fac-dayone-johor-nusajaya`) Johor Bahru, Johor, Malaysia — 240 MW · H100-equivalent 281,455 · type: ai_training · status: operational. Editorial notes: Total capital cost (2025 USD): $9.09B | Epoch estimation methods: our cooling equipment power model, company disclosures, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-dayone-johor-nusajaya. ### QTS Richmond 1 (`fac-epoch-qts-richmond-1`) Sandston, Virginia, United States — 238 MW · H100-equivalent 242,546 · type: ai_training · status: operational. Editorial notes: Total capital cost (2025 USD): $9.02B | Epoch estimation methods: company disclosures, regulatory filings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-qts-richmond-1. ### Google Council Bluffs (East) (`fac-epoch-google-council-bluffs-east`) Council Bluffs, Iowa, United States — 237 MW · H100-equivalent 346,639 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $8.98B | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-council-bluffs-east. ### Google Omaha (`fac-epoch-google-omaha`) Omaha, Nebraska, United States — 237 MW · H100-equivalent 275,391 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $8.98B | Epoch estimation methods: our cooling equipment power model, regulatory filings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-omaha. ### Google Papillion (`fac-epoch-google-papillion`) Papillion, Nebraska, United States — 237 MW · H100-equivalent 80,343 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $8.98B | Epoch estimation methods: our cooling equipment power model, regulatory filings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-papillion. ### Arsenal Data Site 250 (`gridstatus-nyiso-1728`) NY, United States — 233 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1728. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1728. ### Arsenal Data Site 500 (`gridstatus-nyiso-1729`) NY, United States — 233 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1729. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1729. ### Amazon Ridgeland (`fac-epoch-amazon-ridgeland`) Ridgeland, Mississippi, United States — 228 MW · H100-equivalent 171,298 · type: ai_training · status: operational · operator `org-amazon`. Editorial notes: Total capital cost (2025 USD): $8.64B | Epoch estimation methods: regulatory filings, chip efficiency, third-party reporting.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-amazon-ridgeland. ### VNET Bayin Ulanqab (`fac-epoch-vnet-bayin-ulanqab`) Ulanqab, Inner Mongolia, China — 221 MW · H100-equivalent 122,789 · type: ai_training · status: operational · operator `org-vnet-group`. Editorial notes: Total capital cost (2025 USD): $8.37B | Distinct campus from fac-alibaba-ulanqab (Alibaba Cloud Ulanqab Super DC): VNET's own 2024-04-18 press release sites this project in the Bayin district (巴音片区) of Chahar High-Tech Zone (vnet.com/portal/article/index/cid/14/id/879.html, tier 1); Alibaba's campus sits in the separately-named Yiwu sub-district. Verified 2026-07-10. | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-vnet-bayin-ulanqab. ### Microsoft Goodyear (`fac-epoch-microsoft-goodyear`) Goodyear, Arizona, United States — 202 MW · H100-equivalent 205,154 · type: ai_training · status: operational · operator `org-microsoft`. Editorial notes: Total capital cost (2025 USD): $7.65B | Epoch estimation methods: our cooling equipment power model, company disclosures, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-microsoft-goodyear. ### St Lawrence Data and Agricultural Center (`gridstatus-nyiso-1213`) NY, United States — 200 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1213. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1213. ### Proposed Datacenters at 450 Broadway, Buchanan, NY, 10511 (`gridstatus-nyiso-1717`) NY, United States — 200 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1717. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1717. ### Greenidge 200 MW Data Center Project (`gridstatus-nyiso-1725`) NY, United States — 200 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1725. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1725. ### Globe Digital Holdings - 1 (`gridstatus-nyiso-1747`) NY, United States — 200 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1747. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1747. ### Microsoft Project Osmium (`fac-epoch-microsoft-project-osmium`) Cumming, Iowa, United States — 190 MW · H100-equivalent 151,591 · type: ai_training · status: operational · operator `org-microsoft`. Editorial notes: Total capital cost (2025 USD): $7.20B | Epoch estimation methods: our cooling equipment power model, company disclosures, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-microsoft-project-osmium. ### QTS Richmond 2 (`fac-epoch-qts-richmond-2`) Sandston, Virginia, United States — 180 MW · H100-equivalent 182,920 · type: ai_training · status: operational. Editorial notes: Total capital cost (2025 USD): $6.82B | Epoch estimation methods: company disclosures, regulatory filings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-qts-richmond-2. ### Meta Rosemount (`fac-epoch-meta-rosemount`) Rosemount, Minnesota, United States — 178 MW · H100-equivalent 215,765 · type: ai_training · status: operational · operator `org-meta`. Editorial notes: Total capital cost (2025 USD): $6.74B | Epoch estimation methods: our cooling equipment power model, structural similarity to known buildings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-meta-rosemount. ### Meta Jeffersonville (`fac-epoch-meta-jeffersonville`) Jeffersonville, Indiana, United States — 178 MW · H100-equivalent 203,132 · type: ai_training · status: operational · operator `org-meta`. Editorial notes: Total capital cost (2025 USD): $6.74B | Epoch estimation methods: our cooling equipment power model, structural similarity to known buildings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-meta-jeffersonville. ### North Country Data Center (`gridstatus-nyiso-0979`) NY, United States — 176 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 0979. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-0979. ### Alibaba Zhangbei (`fac-epoch-alibaba-zhangbei`) Zhangbei, Hebei, China — 169 MW · H100-equivalent 132,895 · type: ai_training · status: operational · operator `org-alibaba`. Editorial notes: Total capital cost (2025 USD): $6.40B | Epoch estimation methods: our cooling equipment power model, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-alibaba-zhangbei. ### Cayuga Data (`gridstatus-nyiso-1733`) NY, United States — 162 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1733. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1733. ### Google Storey County (`fac-epoch-google-storey-county`) Clark, Nevada, United States — 161 MW · H100-equivalent 153,612 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $6.10B | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-storey-county. ### Google The Dalles (`fac-epoch-google-the-dalles`) The Dalles, Oregon, United States — 154 MW · H100-equivalent 201,616 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $5.83B | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-the-dalles. ### Meta Montgomery (`fac-epoch-meta-montgomery`) Montgomery, Alabama, United States — 153 MW · H100-equivalent 174,330 · type: ai_training · status: operational · operator `org-meta`. Editorial notes: Total capital cost (2025 USD): $5.80B | Epoch estimation methods: our cooling equipment power model, structural similarity to known buildings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-meta-montgomery. ### Meta Kuna (`fac-epoch-meta-kuna`) Kuna, Idaho, United States — 152 MW · H100-equivalent 173,319 · type: ai_training · status: operational · operator `org-meta`. Editorial notes: Total capital cost (2025 USD): $5.76B | Epoch estimation methods: our cooling equipment power model, structural similarity to known buildings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-meta-kuna. ### Meta Temple (`fac-epoch-meta-temple`) Temple, Texas, United States — 152 MW · H100-equivalent 173,247 · type: ai_training · status: operational · operator `org-meta`. Editorial notes: Total capital cost (2025 USD): $5.76B | Epoch estimation methods: our cooling equipment power model, company disclosures, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-meta-temple. ### QTS Richmond 3 (`fac-epoch-qts-richmond-3`) Sandston, Virginia, United States — 144 MW · H100-equivalent 174,835 · type: ai_training · status: operational. Editorial notes: Total capital cost (2025 USD): $5.46B | Epoch estimation methods: company disclosures, regulatory filings, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-qts-richmond-3. ### Google Lincoln (`fac-epoch-google-lincoln`) Lincoln, Nebraska, United States — 141 MW · H100-equivalent 300,656 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $5.34B | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-lincoln. ### Niagara Digital Campus (`gridstatus-nyiso-1681`) NY, United States — 140 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1681. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1681. ### Google Lancaster (`fac-epoch-google-lancaster`) Lancaster, Ohio, United States — 137 MW · H100-equivalent 130,368 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $5.19B | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-lancaster. ### Giga Texas Data Center (`gridstatus-ercot-25INR0688`) TX, United States — 133 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 25INR0688. Interconnecting Entity: Giga Texas Energy, LLC. Gen Type: Other - Battery Energy Storage (NOTE: Filed as battery storage interconnection, facility is data center per project name).. · primary source: https://www.ercot.com/misdownload/servlets/mirDownload?doclookupId=1200125027 Detail page: https://scrutica.com/facilities/gridstatus-ercot-25INR0688. ### Coreweave Helios (`fac-epoch-coreweave-helios`) Afton, Texas, United States — 132 MW · H100-equivalent 159,551 · type: ai_training · status: operational · operator `org-coreweave`. Editorial notes: Total capital cost (2025 USD): $5.00B | Epoch estimation methods: our cooling equipment power model, company disclosures, regulatory filings.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-coreweave-helios. ### Google Midlothian (`fac-epoch-google-midlothian`) Midlothian, Texas, United States — 103 MW · H100-equivalent 199,090 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $3.90B | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency, regulatory filings.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-midlothian. ### Massena Load (`gridstatus-nyiso-0909`) NY, United States — 100 MW · type: hyperscale_dc · status: decommissioned. Editorial notes: Queue ID: 0909. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-0909. ### Google Mesa (`fac-epoch-google-mesa`) United States — 90 MW · H100-equivalent 125,821 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $3.41B | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-mesa. ### Google Waltham Cross (`fac-epoch-google-waltham-cross`) Cheshunt, Hertfordshire, United Kingdom — 88 MW · H100-equivalent 123,799 · type: ai_training · status: operational · operator `org-google`. Editorial notes: Total capital cost (2025 USD): $3.33B | Epoch estimation methods: our cooling equipment power model, chip availability, chip efficiency.. · primary source: https://epoch.ai/data/data-centers Detail page: https://scrutica.com/facilities/fac-epoch-google-waltham-cross. ### Cayuga Compute (`gridstatus-nyiso-1683`) NY, United States — 88 MW · type: hyperscale_dc · status: announced. Editorial notes: Queue ID: 1683. Gen type: Load. Proposed COD: nan.. · primary source: https://www.nyiso.com/documents/20142/1407078/NYISO-Interconnection-Queue.xlsx Detail page: https://scrutica.com/facilities/gridstatus-nyiso-1683. --- ## Reference: per-country compute footprint summary Country-level rollup of the disclosed-power facility substrate, sorted by total disclosed power capacity. The aggregation is computed from the curated `data/processed/facilities.json` subset (the canonical Epoch-anchored facility table) — the full 4,550-facility production table at `/entities` carries additional facilities below the curated frontier-anchor threshold. Cross-references to the corresponding sovereign-AI program are inline where the country runs a tracked sovereign-AI initiative. Each row summarises: total disclosed power capacity in megawatts, the facility count contributing to that figure, the count of operational vs. under-construction vs. announced status splits where disclosed, and the top operator entities by aggregate power capacity within the country. Where the country has a tracked sovereign-AI program the program name + announced-headline figure is appended so the reader can see the compute substrate alongside the financial-commitment narrative in a single rollup. ### United States - Disclosed power total: 16,830 MW across 72 facility rows - Status mix: 47 operational · 22 announced - Top operators by aggregate MW: `org-google` (2346 MW), `org-meta` (1618 MW), `org-microsoft` (1519 MW) - Tracked sovereign-AI program: United States: NAIRR / CHIPS Act AI Components / Stargate Project — announced headline $2.60B; government-only $30.0M. Detail: https://scrutica.com/sovereign-ai - Drill-down: https://scrutica.com/entities?country=US ### China - Disclosed power total: 631.8 MW across 3 facility rows - Status mix: 3 operational - Top operators by aggregate MW: `org-huawei` (241.8 MW), `org-vnet-group` (221 MW), `org-alibaba` (169 MW) - Tracked sovereign-AI program: China: National AI Programs / Big Fund III / Eastern Data Western Computing — announced headline $55.70B; government-only $55.70B. Detail: https://scrutica.com/sovereign-ai - Drill-down: https://scrutica.com/entities?country=CN ### Malaysia - Disclosed power total: 240 MW across 1 facility row - Status mix: 1 operational - Tracked sovereign-AI program: Malaysia: National AI Roadmap 2030 / Sovereign AI Cloud — announced headline $4.30B; government-only $440.0M. Detail: https://scrutica.com/sovereign-ai - Drill-down: https://scrutica.com/entities?country=MY ### United Kingdom - Disclosed power total: 88 MW across 1 facility row - Status mix: 1 operational - Top operators by aggregate MW: `org-google` (88 MW) - Tracked sovereign-AI program: United Kingdom: AI Opportunities Action Plan / UK Compute Roadmap / AIRR / Sovereign AI Unit / AI Hardware Plan — announced headline $4.75B; government-only $4.75B. Detail: https://scrutica.com/sovereign-ai - Drill-down: https://scrutica.com/entities?country=GB ### Portugal - Disclosed power total: 33 MW across 1 facility row - Status mix: 1 operational - Top operators by aggregate MW: `org-nscale` (33 MW) - Drill-down: https://scrutica.com/entities?country=PT --- ## Reference: AI accelerator catalog Selected AI accelerator product lines tracked in the chip catalog. The catalog (in the production `hardware_catalog` table) carries peak BF16 / FP8 / INT8 throughput, HBM capacity and bandwidth, packaging class, and export-control classification per SKU. Per-chip detail pages at https://scrutica.com/chips/{name}. ### NVIDIA Hopper generation (2022-2024 deployment generation) - **H100 SXM5** — 989.5 BF16 TFLOPS dense; 80 GB HBM3 (3.35 TB/s); CoWoS-S packaging; export-control class 3A090 (restricted to certain destinations under EAR). - **H100 PCIe** — Lower-power variant of H100; same compute, reduced TDP. - **H200** — Hopper-class compute (~equivalent to H100 in BF16 dense), refreshed HBM to HBM3e (141 GB capacity, 4.8 TB/s bandwidth). Same architecture; the upgrade is primarily on memory side. - **H800** — Export-compliance variant for export-restricted destinations. NVLink interconnect bandwidth reduced relative to H100. Effective compute lower than H100 SXM5 despite the chip-name proximity. - **H20** — Further export-restricted variant. Significantly lower compute envelope than H100; targeted for certain restricted destinations. ### NVIDIA Blackwell generation (2024-2025 deployment generation) - **B100** — Single-die Blackwell; vendor-published BF16 / FP8 / FP6 / FP4 throughput. - **B200** — Two-die Blackwell with chip-to-chip NVLink fabric; ~2× B100 in dense BF16 + lower-precision modes. - **B300** — Refresh generation with HBM3e upgrade. - **GB200** — Grace ARM CPU + B200 GPU in coherent platform; deployed in the GB200 NVL72 72-GPU rack-scale system. - **GB300 NVL72** — Refresh generation rack-scale system; named in xAI Colossus 2 and other multi-billion-dollar deployments. ### NVIDIA Rubin generation (announced 2024; targeted 2025-2026 production) - **Rubin / R100** — Next-generation post-Blackwell architecture; HBM4 memory transition. - **R-Ultra / R300** — Refresh generation announced for late 2026 / 2027. ### AMD CDNA-architecture line - **MI300X** — 304 compute units; 192 GB HBM3; vendor-published peak BF16 TFLOPS; one of the two named alternative AI accelerators with hyperscaler commercial deployment (alongside H100/H200) as of 2024-2025. - **MI325X** — HBM3e refresh of MI300X; ~256 GB HBM3e. TDP 1000 W. The single highest-HBM-capacity production accelerator at announcement. - **MI355X / MI400 series** — 2025-2026 roadmap variants. ### Google TPU line - **TPU v4** — Google Cloud-deployed; used internally for Gemini training. - **TPU v5p** — Next-generation training-class TPU; deployed 2024. - **TPU v6e (Trillium)** — Inference-and-training class refresh, 2024-2025. - **TPU v7** — Roadmap successor. ### AWS Trainium and Inferentia - **Trainium 1 (Trn1)** — First-generation AWS in-house training accelerator. - **Trainium 2 (Trn2)** — Second-generation; significantly increased BF16 throughput; named in Anthropic Project Rainier commitment. - **Inferentia 1 / 2 / 3** — Inference-optimized line. ### Other production AI accelerators - **Intel Gaudi 3** — 2024-generation training accelerator from Intel via Habana Labs acquisition heritage. - **Huawei Ascend 910B / 920** — Chinese domestic-substitution narrative; deployed across Chinese hyperscaler and government compute. - **Cerebras WSE-3** — Wafer-scale design; one wafer = one chip with the full reticle field active. - **Tenstorrent Wormhole / Blackhole** — Open-source-RISC-V-anchored architecture. - **Groq LPU** — Tensor-streaming architecture; inference-optimized. - **SambaNova RDU** — Reconfigurable dataflow architecture; enterprise-deployed. ### Recently announced / in qualification - **Microsoft Maia 100 / 200** — Microsoft in-house silicon for Azure AI workloads. - **Meta MTIA v1 / v2** — Meta Training and Inference Accelerator. - **Apple in-house silicon** (Private Cloud Compute) — Apple's training infrastructure substrate, partially disclosed. The full production-substrate chip catalog with vendor-published peak FLOPS at multiple precisions, HBM capacity / bandwidth, packaging class, and export-control classification is browsable at the `/chips` directory; the MCP `scrutica_estimate_flops` tool consumes a fixed enum subset of these (H100-SXM5, H200, B100, B200, A100-80GB, MI300X, TPU-v5p) for capacity-estimation calls. --- ## Scenarios — detail The five pre-built scenarios on `/scenario-analysis` anchor Scrutica's chokepoint-cascade methodology in named, primary-source-grounded geopolitical stress tests. Each is a worked example rather than a forecast — the question being asked is "if event X happened tomorrow, which entities in Scrutica's supply-chain graph would be downstream-exposed, and through what bilateral edges?" ### Taiwan Strait disruption (`taiwan-strait`) Four sub-scenarios anchored to TSMC's structural single-supplier role for frontier-AI silicon. The sub-scenarios are deliberately ordered from least-to-most physical-disruption-severity to let the user reason about the "what scale of intervention triggers what cascade magnitude" gradient. - **Sub-scenario 1 — Maritime blockade.** Sea-lane interdiction in the Taiwan Strait restricts shipping in and out of Taiwan-island ports. TSMC fab operations continue but outbound shipping of finished wafers and assembled chips is blocked. Cascade propagation: every TSMC customer's allocation queue stretches; the H100 / B200 / B300 inventory at downstream-customer warehouses depletes within 6-10 weeks under high-demand conditions; the supply-chain edges to NVIDIA + AMD + Apple + Qualcomm + MediaTek become the cascade's primary propagation lanes. - **Sub-scenario 2 — Targeted strike.** Direct physical damage to one or more TSMC fab buildings (Hsinchu Science Park or Tainan Fab 18). Recovery timeline 12-24 months for the affected fab; multi-fab damage extends recovery substantially. Cascade propagation: as in Sub-scenario 1 but with permanent capacity-loss component layered on top of the shipping-interdiction effect. - **Sub-scenario 3 — Grid failure.** Taiwan grid disruption (whether from typhoon damage, drought-induced hydroelectric shortfall, or coordinated infrastructure attack). Fab operations halted temporarily; restart cycles for advanced-node fabs are 4-12 weeks after grid restoration. Cascade propagation: shorter than Sub-scenarios 1 and 2 but still significant given the multi-month restart timelines. - **Sub-scenario 4 — Export restriction.** US adds TSMC-specific advanced-node restriction (hypothetical scenario, not currently in force). TSMC retains physical capacity but cannot ship advanced-node products to specified destination set. Cascade propagation: bifurcated — affected destinations face full TSMC restriction; unaffected destinations face only the second-order spillover effects of allocation reshuffling. The scenario detail page renders the affected-node count and the propagation depth under each sub-scenario, with the seed nodes (TSMC + its top-50 customers by licensed-database supply-share disclosure) listed explicitly. ### IRGC missile range (`iran-threat`) Iranian Revolutionary Guard Corps missile-range targeting capability extends across the Persian Gulf region. The scenario seeds on the data-center footprint in countries within IRGC missile-range coverage: UAE (entire country within range of Iranian medium-range ballistic missiles), Saudi Arabia (eastern provinces and Riyadh within range; western provinces partially within range), Qatar, Bahrain, Kuwait, the eastern Mediterranean coast. Cascade propagation: affected facilities include the UAE compute substrate (Stargate UAE 5 GW campus aspiration, the Microsoft + G42 datacenter expansion, the Oracle OCI Abu Dhabi region with the Sovereign AI Supercluster), the Saudi HUMAIN program facilities, and the broader Gulf region's data-center capacity. The Coordination Gap Analyzer renders the structural-arbitrage exposure where Gulf compute is operationally controlled by US / UK / EU hyperscalers (high BIS jurisdiction reach but physical-host country in the missile-range envelope). The scenario does not assess the probability of the IRGC kinetic action; it surfaces the structural exposure if such action occurred. The 2019 Abqaiq attack on Saudi oil infrastructure is the named historical analogue for "Iranian-attributed disruption to Gulf infrastructure" — the scenario uses Abqaiq as the calibration anchor for "what kind of disruption pattern is operationally plausible." ### Tokyo earthquake (`tokyo-earthquake`) Japan's compute-supply-chain exposure to a major Tokyo-region earthquake (Mw 7.0+). Three sub-scenarios bounded by historically-observed Japanese seismic events: - **M7.0 sub-scenario** — Comparable to the 2011 Tōhoku event's secondary damage radius (Tōhoku itself was 9.0 with damage centered in the northeast; the M7.0 scenario simulates a Kantō-region rupture). Affected nodes: Kioxia memory fabs (Yokkaichi is outside the immediate radius but supply-chain consequences extend), Tokyo Electron equipment manufacturing, JSR / Tokyo Ohka photoresist supply. - **M7.5 sub-scenario** — Mid-bound scenario; substantial damage to Tokyo-region facilities with regional supply-chain disruption lasting 6-18 months. - **M8.0 sub-scenario** — Upper-bound scenario; multi-year supply-chain disruption. Cascade propagation centers on the HBM / packaging / equipment supply chain where Japanese suppliers are structurally critical. Tokyo Electron is the most-load-bearing Japanese node; JSR is the photoresist chokepoint; Shin-Etsu Handotai (Niigata Prefecture, outside the immediate Tokyo radius) is the wafer-supply chokepoint. ### South China Sea cable severing (`south-china-sea`) Four sub-scenarios on submarine cable disruption affecting cross-regional Asian compute traffic. Submarine cables are a structural single-point-of-failure for inter-region data flow; physical severance has historical precedent (the 2008 Mediterranean cable cuts, the 2024 Red Sea cable cuts). Cascade propagation through compute substrate: Asian hyperscaler cloud regions (Alibaba Singapore, Tencent Singapore, AWS Singapore, Google Singapore, Microsoft Hong Kong, ByteDance Singapore) experience cross-region latency and capacity loss; failover paths through alternative routes face capacity-constraint pressure. The scenario surfaces the Coordination Gap angle: compute physically sitting in Singapore but operationally controlled by US hyperscalers is exposed to both jurisdictions' regulatory frames; submarine-cable disruption is a physical-layer event that bypasses both regulatory frames. ### Saudi grid / Abqaiq 2.0 (`abqaiq-2`) Saudi compute infrastructure exposure to a repeat of the 2019 Abqaiq attack on the Saudi oil-processing facility, projected onto AI compute substrate. The scenario uses Abqaiq as the calibration anchor (kinetic action on Saudi infrastructure with attribution to Iran-aligned actors) and projects forward: HUMAIN program facilities (Public Investment Fund-anchored), Saudi compute infrastructure dependencies, and the second-order effect of regional grid instability on adjacent compute deployments (UAE, Bahrain, Kuwait). Cascade propagation includes the sovereign-program dependencies — the HUMAIN program's interdependence score (NVIDIA: high; TSMC: high; US: high; BIS jurisdiction reach: medium-high) makes the Saudi compute substrate structurally arbitrage-exposed even before any kinetic disruption. --- ## Frequently asked questions (agent-optimized) This section is structured for agents fielding common compute-governance queries. Each Q&A pairs the question with the canonical Scrutica surface that answers it and the methodology version under which the answer was computed. ### How many AI compute facilities does Scrutica track? 4,550 compute-related facilities across 110 countries as of 2026-07-23. The directory is at https://scrutica.com/entities. Coverage centers on Epoch AI's frontier-DC catalog (US, UAE, China sites), the IM3 OpenStreetMap US data-center bulk corpus, the NYISO load-interconnection queue (and analogous data sets across the seven US ISOs/RTOs), and sovereign-program-linked international sites. Tier-2 sovereign-facility ingest (HUMAIN, Yotta, Tencent self-operated facilities, EuroHPC sites) is in active sprint. ### How many supply-chain edges does Scrutica's graph contain? 21,037 bilateral edges (15,049 after canonical deduplication on the supplier-customer-product key). Sources: a licensed supply-chain database (Tier 1) + a licensed corporate-ownership database (Tier 1) + SEC EDGAR Exhibit 21 (Tier 1) + CSET ETO (Tier 2) + editorial primary-source extraction. The supply-chain explorer is at https://scrutica.com/supply-chain. ### How many BIS Entity List designations does Scrutica track? 3,422 export-control designations cross-referenced onto compute-universe entities. The full surface is at https://scrutica.com/export-controls. The cross-reference confidence is recorded per match (`bis_crossref_matches.confidence_score`) so an agent can filter by match-confidence tier. ### How many sovereign AI programs does Scrutica track? 36 programs with announced-versus-deployed reconciliation. Each program carries announced (headline) versus committed versus disbursed figures with explicit data quality flags. The sovereign AI dashboard is at https://scrutica.com/sovereign-ai. ### What is the data refresh cadence? Cron-driven ingest at the cadence of the upstream source. BIS Entity List rules land within 1-3 days of Federal Register publication. SEC filings land within hours of EDGAR availability. licensed supply-chain edges refresh quarterly per the upstream database update cadence. Sovereign procurement disclosures (USAspending, EU TED, UK Find a Tender) land within days of the originating portal update. Per-page freshness indicators at https://scrutica.com/methodology document the per-source expected staleness. ### How can I cite a specific Scrutica estimate? Include the methodology version stamp (e.g., `v2.1`) and the estimation path (e.g., "Hardware Path") in your citation. The "Cite This" panel on every entity page emits APA + BibTeX + Chicago + JSON-LD; the methodology version is embedded in the emission. This lets future readers reproduce the calculation against the same parameter set even after Scrutica's methodology evolves. Citation format example: `Scrutica. "[Facility Name]." FLOP Engine Methodology v2.1, Hardware Path. [Date]. https://scrutica.com/facilities/[id]`. ### How does Scrutica handle cross-source contradictions? Both values surface explicitly with a cross-source-divergence flag and the editorial disposition. The data quality flag system at `organization_data_quality_flags` records the contradiction, the resolution disposition, and the dispositioning rationale. Silent winner-picking is a violation; visible disagreement is the discipline. See for example the Hua Hong cross-source-divergence flag rendering on the company detail page. ### What is the FLOP Engine's confidence interval? The engine emits three independent path-estimates (hardware inventory, power envelope, capital cost). The three typically agree within ~1.5× for well-disclosed facilities. Where they diverge by >2× (the cross-path-divergence threshold), the facility intelligence card surfaces the discrepancy as a contradiction flag with the disposition. Confidence intervals on individual path estimates are computed from the assumption-range bounds and exposed in the methodology page's interactive panel. ### Where does Scrutica's data not cover? State-owned enterprises with limited public disclosure obligations, privately-held companies in jurisdictions with limited reporting (especially China, Russia, Iran), and certain sovereign-AI vehicles that route through fund-of-fund structures. Forecast layers (forward-looking demand projections, revenue forecasts) are out of scope; the Query Engine refuses rather than fabricates when asked to forecast. ### Can I integrate Scrutica's data into my AI agent? Yes. Three integration paths: 1. **MCP server** at https://scrutica.com/api/mcp — Streamable HTTP transport, ten tools. 2. **Direct page fetch** — every URL returns HTML; per-page markdown alternatives at `.md` suffix (when the L3 route handler ships). 3. **Dataset DOIs** — scheduled Zenodo deposition for the Supply Chain Graph, Sovereign AI Index, and Compute Cost Index; full dataset download with Croissant JSON-LD provenance. The MCP agent guide at https://scrutica.com/api/mcp/llms.txt documents per-tool examples and anti-patterns. ### What is Scrutica's relationship to Epoch AI? Complementary. Epoch AI's Notable Models and Frontier Data Centers corpora are primary sources Scrutica ingests, attributes, and links back to; Scrutica's own layer is the cross-substrate join — facilities joined to corporate control, supply-chain edges, export-control designations, and sovereign programs. The `/related-work` page documents the citation surfaces and the Scrutica-side analytical layers (cascade simulation, sovereign AI reconciliation, compute visibility, coordination gaps). ### Is the data open? CC BY-SA 4.0 for the Scrutica-derived substrate. Source corpora that originate in licensed databases (held under subscription; not redistributed) are stored as extracted relational facts with provenance preserved; the raw paywalled datasets themselves are not redistributed. Downstream consumers preserving the BY-SA terms can re-use the substrate freely; commercial redistribution should contact David Gringras. ### How is Scrutica funded? Currently bootstrapped by David Gringras as part of ongoing AI governance research at Harvard. The platform is exploring AI safety / EA funding pathways (LTFF, Open Philanthropy / Coefficient Giving, RP Special Projects fiscal sponsorship) as it scales beyond the MVP phase. --- ## How agents can integrate with Scrutica **MCP server.** Configure your agent client (Claude Desktop, Cursor, VS Code with MCP support, or any other MCP-compatible client) with: ```json { "mcpServers": { "scrutica": { "url": "https://scrutica.com/api/mcp", "transport": "streamable-http" } } } ``` The server exposes ten tools, optimized for the agent-tool sweet spot (6-8 tools maximum for tool-routing clarity, expanded past it where the substrate-shape demands it — facility vs organization fetches are split into separate tools because their attribute sets diverge sharply and a single `get_entity` with type discrimination was producing worse LLM tool-routing in pilot; the Entity List designations and change-log tools are split for the same reason, since "who is listed" returns rows and "what changed" returns notice-grouped summaries, and folding them into one tool would mean a flag that silently changes both the response shape and the bound on its size): - `scrutica_search` — Free-text search across facilities, organizations, sovereign programs. Filter by `entity_type` and ISO 3166-1 alpha-2 `country`. Returns ranked candidates with id, name, type, summary, Scrutica URL, and provenance. - `scrutica_get_facility` — Retrieve a single facility by canonical ID (e.g. `fac-tsmc-arizona-fab21-p2`) with operator, owner, location, power, GPU inventory (where disclosed), data_source, source_url, is_estimated flags. - `scrutica_get_company` — Retrieve a single organization by canonical ID (e.g. `org-nvidia`) with legal name, country of HQ, organization type, parent / subsidiary references, supply-chain edge counts. - `scrutica_get_supply_chain` — Retrieve the supply-chain edges centered on one or more organizations, `direction` ∈ {upstream, downstream, both}. Each edge carries source/target IDs, criticality_score, 3-month price correlation (where available), data_source. - `scrutica_query_export_controls` — Look up BIS Entity List designations by entity name (case-insensitive substring on the published name) / org ID / jurisdiction. Returns designation date + Federal Register citation + source URL (authority tier 1). Does not cover OFAC SDN or Wassenaar CCL; an organization's OFAC SDN / NS-CMIC / Section-1260H flags ride on `scrutica_get_company`. - `scrutica_entity_list_changes` — What CHANGED in the Entity List: designation rows grouped into per-Federal-Register-notice change events (canonical citation, notice title and publication date where resolved, derived event date labeled with its source column, entity / addition / removal counts, per-country counts, cross-referenced compute organizations), date-keyed removal actions as first-class events, and a last-N-ISO-week activity rollup counting both directions. Bounded summaries, never row dumps. - `scrutica_estimate_flops` — Compute peak BF16 FLOP estimates for a hardware × count × utilization × sparsity × precision configuration. Returns point estimate + Scrutica-editorial bounds. - `scrutica_get_scenario` — Retrieve a named scenario (`taiwan-strait`, `iran-threat`, `tokyo-earthquake`, `south-china-sea`, `abqaiq-2`) with summary + interactive-analysis URL. - `scrutica_get_sovereign_program` — Retrieve a sovereign AI program by ISO country code (or `list_all`) with announced / committed / disbursed figures + interdependence + data quality flags. - `scrutica_get_methodology` — Retrieve a methodology section (`flop-estimation`, `cost-index`, `compute-visibility-index`, `supply-chain-weighting`, `chokepoint-cascade`, `sovereign-execution-classification`) with canonical URL + summary. Detailed agent guide at https://scrutica.com/api/mcp/llms.txt. The guide includes per-tool example calls, parameter syntax, expected response shape, anti-patterns to avoid (e.g., do not call `scrutica_get_facility` or `scrutica_get_company` in a loop over a long list — use `scrutica_search` with the relevant `entity_type` filter and `limit` first, then fetch by ID). **Natural-language Query Builder.** `https://scrutica.com/query` exposes the same substrate via NL queries. Fifty-five named tools across fourteen layers — facilities, organizations, supply-chain, export-controls, sovereign programs, capabilities, economics, geopolitics, scenarios, methodology, cost-index, threshold-atlas, compute-visibility, glossary. Every response wraps numerical claims in citation tags with authority-tier provenance; a server-side verifier flags any cited record_id absent from the underlying tool-result history. **Direct page fetch.** Every page is publicly accessible. The L3 markdown-alternative route handler (when shipped) supports appending `.md` to any URL for the Markdown-extracted version (e.g., https://scrutica.com/methodology.md). Markdown versions strip site chrome while preserving content, tables, and links — the format optimal for RAG pipelines that prefer flat markdown over rendered HTML. **Data export.** Every entity exposes a "Cite This" panel emitting APA + BibTeX + Chicago + JSON-LD + DOI (where applicable). The full datasets — Supply Chain Graph, Sovereign AI Index, Compute Cost Index — are scheduled for Zenodo publication with citable DOIs and Croissant JSON-LD; see https://scrutica.com/datasets for the canonical download links once the L5 dataset-packaging session lands. **Provenance verification.** Every claim Scrutica makes traces to a source URL on the originating page. Agents fetching Scrutica pages can verify any cited statistic by following the source URL to the primary-source filing or measurement. Trust-but-verify is the canonical pattern — Scrutica's structured-data emission is designed to make verification cheap rather than to encourage delegated trust. **Rate limits and politeness.** Scrutica's robots.txt allows all named AI crawlers (training and retrieval) at unrestricted access except for two whose historical crawl rates required throttling (Amazonbot crawl-delay 1, Bytespider crawl-delay 2). The MCP server is rate-limited per IP at the Vercel level; the limit is currently set high enough that normal agent usage will not hit it. Bulk-extraction usage that would exceed rate limits should contact davidgringras@hsph.harvard.edu — Scrutica is interested in supporting legitimate research consumption. **Citation expectation.** Scrutica's data is CC BY-SA 4.0. Agentic consumption that cites Scrutica in its responses meets the attribution requirement. Bulk-redistribution of the substrate requires preserving the BY-SA terms (downstream consumers also receive the data under CC BY-SA 4.0); commercial redistribution should contact David Gringras. --- ## Datasets, exports, and downstream integration Scrutica's substrate is exposed for downstream analytical consumption through four distinct integration paths, each oriented to a different consumer category. Per CLAUDE.md provenance discipline, every export preserves the per-record `data_source`, `source_url`, `is_estimated`, and `authority_tier` columns — downstream analysts do not lose the audit chain when they leave the Scrutica UI. **Path 1: Per-entity Cite-This emissions.** Every facility, organisation, sovereign program, BIS designation, supply-chain edge, scenario, and methodology page exposes a Cite-This panel that emits APA + BibTeX + Chicago + JSON-LD with the entity's canonical Scrutica ID and a methodology version stamp. The emission is one click from the entity page; the format matches the academic-citation conventions in each style. The methodology version stamp is the load-bearing field for reproducibility — citing a 2024 estimate with `v2.0` of the FLOP-engine methodology produces a different number than the same citation under `v2.1` if any underlying parameter changed. **Path 2: Bulk export per entity class.** The supply-chain explorer, the export-controls page, the sovereign-AI dashboard, the cost-index page, and the threshold-atlas all expose JSON / CSV / Croissant download buttons. The Croissant emission carries the W3C / Hugging Face Croissant schema with the PROV-O `wasDerivedFrom` predicate populated for every dataset record — this is the format that Hugging Face Datasets, Google Dataset Search, and the broader ML-tooling ecosystem prefer for provenance-aware ingestion. The CSV emission preserves all provenance columns inline (rather than stripping them to a "clean" two-column-output form) so downstream loaders inherit the audit chain. **Path 3: Zenodo dataset deposits with citable DOIs.** Three Zenodo-anchored datasets are scheduled for publication under CC BY-SA 4.0 with citable DOIs: - **Supply Chain Graph** — the 21,037-edge bilateral relationship graph with per-edge supply share, 3-month price correlation, editorial criticality, source corpus, and tier classification. Includes the per-edge confidence-tier classification and the canonical organisation aliases that resolve variant entity spellings onto stable Scrutica IDs. - **Sovereign AI Index** — the 36-program reconciliation dataset with per-program announced / committed / disbursed figures, interdependence scores, inflation decomposition components, data quality flags, and primary-source citation sets. Includes the per-program editorial dispositioning notes. - **Compute Cost Index** — the cross-cloud $/PFLOP-day time series with per-instance hourly pricing, per-hardware-tier classification, SCU conversion intermediate values, and region adjustments. Refreshed nightly via the `/api/cron/cloud-pricing` handler. Each Zenodo deposition will include a Croissant JSON-LD manifest and a README documenting the schema, the assumption ranges, and the per-dataset cross-validation pass against external benchmarks. The DOIs land on `/datasets` once the deposition lifecycle completes; until then the substrate is reachable via the per-page download buttons. **Path 4: Model Context Protocol (MCP) server.** The `/api/mcp` endpoint exposes the same substrate as nine agent-callable tools. The MCP server uses the Streamable HTTP transport per the 2025-11-25 spec; every response payload preserves the per-record `source_url`, `authority_tier`, and `is_estimated` fields so the agent can present the citation chain to the user. The full agent guide at https://scrutica.com/api/mcp/llms.txt documents per-tool examples, parameter syntax, expected response shape, and anti-patterns to avoid. **Schema conventions for downstream consumers.** Scrutica's TEXT primary keys are meaningful identifiers — `epoch-dc-001` for Epoch-corpus facilities, `osm-w341892` for OSM-derived facilities, `org-nvidia` for canonical organisations, `bis-entity-huawei` for BIS-only designations. The pattern is documented at https://scrutica.com/methodology#entity-identifier-conventions. Downstream tools that consume Scrutica's substrate can join against canonical identifiers in upstream corpora (Wikidata Q-IDs, ROR identifiers, LEI codes, SEC CIK numbers) via the `same_as_uris` JSONB column populated through the L2 jsonld-entity-schemas pipeline. **License and attribution.** Scrutica-derived substrate is CC BY-SA 4.0 unless otherwise attributed. Source corpora that originate in licensed databases (held under subscription; not redistributed) are stored as extracted relational facts with provenance preserved — the raw paywalled datasets are not redistributed. Downstream consumers that preserve the BY-SA terms can re-use the substrate freely; commercial redistribution should contact David Gringras for licensing arrangements that respect the upstream corpora terms. **Feeds for freshness signalling.** Two RSS 2.0 feeds cover the substrate's moving parts: https://scrutica.com/export-controls/changes/feed.xml (Federal Register actions on the Entity List, one item per notice, last 24 months) and https://scrutica.com/analysis/feed.xml (dated analysis notes). The /changelog page renders facility / organisation refresh activity as HTML; /watchlist mints an RSS subscription URL for your watchlist. The IndexNow cron metadata for the relevant handlers carries the per-refresh URL submission status so analyst pipelines anchored on Scrutica freshness can trigger their downstream re-derivation cycles in lockstep with the upstream refresh. --- ## Limitations (full disclosure) **Coverage bias.** Coverage is structurally biased toward entities with public disclosure obligations. State-owned enterprises, privately held companies, and facilities in jurisdictions with limited reporting are underrepresented — often precisely the entities most relevant to compute governance. China's Big Fund III, Russia's national AI program, Iran's compute infrastructure, North Korean compute (effectively absent from public-disclosure substrate), and the privately-held Chinese hyperscaler segment (Tencent / Alibaba / Baidu / ByteDance self-operated facilities, partially covered in active Phase 2 sprint) are the most-load-bearing examples. **Estimation uncertainty.** Compute-capacity estimates carry real uncertainty, particularly via the power-based and cost-based estimation paths. The methodology page makes the uncertainty visible rather than hiding it; bounds are reported alongside point estimates throughout. The cross-path divergence threshold (2×) is the platform's "flag this as contradictory" boundary; below that, the platform's editorial choice is "show the higher-confidence path with the lower-confidence paths as cross-validation context." **Cascade analysis caveats.** Cascade analysis is stronger on topology than edge weights. Graph structure derives from regulatory filings (Tier 1-2 signal); substitutability decay rates are expert-assessed (Tier 3 signal). The methodology is useful for identifying structurally critical nodes (where graph degree and weighted centrality concentrate); it is less reliable for precise severity prediction at multiple hops, especially under simultaneous-disruption scenarios where the linear-propagation assumption breaks down. Scrutica's editorial position: cascade output is a topology-layer input to multi-source impact analysis, not a complete impact model. **Sovereign program reconciliation.** Sovereign AI program reconciliation depends on the quality of public procurement disclosures. Programs in jurisdictions with limited procurement transparency (or programs that route through privately-held vehicles, fund-of-funds, or holding companies) will undercount their actual disbursement. The methodology page documents per-program reconciliation source coverage so the reader can see "this program's reality ratio is computed against N primary sources covering X% of the announced commitment." **Forecasting layer absence.** The platform does not yet expose forecasting layers. Forward-looking demand projections, capacity-growth projections, revenue forecasts, and scenario probability estimates are out of scope. The Query Engine refuses rather than fabricates when asked to forecast — "I don't have a forecasting model for X; the substrate supports historical and current-state analysis." **Real-time signal latency.** Scrutica's freshness layer (cron-driven ingest from upstream sources) operates at the cadence of the upstream publication. BIS Entity List designations land within 1-3 days of Federal Register publication; SEC filings land within hours of EDGAR availability; licensed supply-chain edges refresh quarterly per the upstream database update cadence; sovereign procurement disclosures land within days of the originating procurement portal update. The platform's freshness indicators on every entity page (TTLs in `src/lib/freshness.ts`) document the per-source expected staleness. **Non-AI-specific compute.** The substrate filter prioritizes AI-compute-relevant entities and facilities. General-purpose compute infrastructure (CPU-dominant cloud regions, HPC facilities without AI accelerator deployment, legacy data centers in non-AI workloads) is partially included via PeeringDB and OSM-derived coverage but is not the analytical center of gravity. Analysts looking for general-purpose-compute substrate should pair Scrutica with Uptime Institute's data center research, the JLL data center reports, and the analyst-database substrate at S&P Global, Synergy Research Group, or 451 Research. **Editorial scope decisions.** Some scope boundaries are deliberate editorial choices rather than coverage limitations: - Scrutica does not track model-specific training-run details (model architecture, hyperparameters, training-data composition) — Epoch's notable-models database is the authoritative source for that layer - Scrutica does not track AI-evaluation results in depth — Stanford HAI's AI Index, METR's autonomy evaluations, the LMSYS Chatbot Arena are the authoritative sources - Scrutica does not track AI labor-market data (engineer salaries, AI research staffing) — the Berkeley CAIS Index and the AI labor-market analyses elsewhere are the authoritative sources - Scrutica does not track AI policy / legislation in depth as the analytical core — the Stanford HAI AI Index policy section, the Brennan Center AI policy tracker, and the EU AI Act monitor are the authoritative sources The deliberate-scope-boundary rationale: Scrutica's distinctive contribution is the compute-infrastructure layer's cross-source integration. Expanding into adjacent analytical surfaces dilutes that contribution. The platform integrates with the adjacent surfaces (e.g., MLPerf benchmark records as Tier-2 cross-validation for the FLOP Engine, Epoch's notable models for capability-vs-compute scatter on the scaling explorer) but does not attempt to replace them. --- ## Source corpora used **Primary-source corpora (Tier 1):** - US Securities and Exchange Commission EDGAR (10-K, 10-Q, 8-K, S-1, Exhibit 21, Schedule 13G/13D, proxy filings) - US Treasury Office of Foreign Assets Control Specially Designated Nationals list - US Federal Register (BIS Entity List, OFAC SDN, Commerce Department export-control announcements, Treasury sanctions notifications, USTR Section 301 designations) - USAspending.gov (federal contracts + grants + assistance awards) - US Patent and Trademark Office records (used for entity-name resolution + corporate-aliases substrate) - EU Tenders Electronic Daily (TED) procurement portal - UK Find a Tender service procurement portal - Companies House (UK corporate registry) - EDGAR-equivalent registries in EU member states, Japan, South Korea, Singapore (where API-accessible) - Federal Reserve Economic Data (FRED) — selected indicators for trade flows, semiconductor cycle, regional economic activity - US Energy Information Administration (EIA) — Form 860 (electricity generators), Form 861 (electricity sales), Form 923 (power plant operations data) - Power-market interconnection queues across the seven US ISOs/RTOs: PJM, ERCOT, MISO, CAISO, ISO-NE, NYISO, SPP — used as the leading-indicator data set for forthcoming compute capacity - Federal Communications Commission filings (for telecommunications-adjacent compute infrastructure) - Bureau of Economic Analysis trade data - US Patent and Trademark Office assignment records (for ownership-chain resolution) **Research-database corpora (Tier 2):** - Epoch AI Frontier Data Centers + GPU Clusters + AI Models + AI Chip Owners - CSET ETO (Center for Security and Emerging Technology, Emerging Technology Observatory) - A licensed corporate-ownership database (fundamentals, entity hierarchies, industry classification; accessed via Harvard + MIT institutional credentials, held under subscription and not redistributed) - A licensed fund/LP database (fund-manager and LP→fund chains; institutional credentials, held under subscription and not redistributed) - A licensed supply-chain database (firm-level supplier-customer relationships; Harvard institutional access, held under subscription and not redistributed) - PeeringDB facility registry - OpenStreetMap Overpass (data-center coverage; the IM3 bulk corpus) - MLPerf Systems benchmark records (the ML Commons benchmark suite) - METR autonomy time-horizon evaluations - Stanford HAI AI Index annual report data - Google Dataset Search (for cross-discovery) - Google Scholar (citation network for academic literature) - OpenAlex (the open-access citation graph; useful for academic-source provenance) - HuggingFace dataset and model registry (for AI-model substrate cross-reference) - Bilateral trade data via BACI (CEPII) - Comtrade UN bilateral trade - China customs data (where available) **Press / analyst corpora (Tier 3):** - The New York Times, Financial Times, Bloomberg, Wall Street Journal, Nikkei Asia, Reuters, The Information, Politico Pro, The Register - IDC sovereign-AI tracking - Gartner data-center research - Goldman Sachs equity research (where citations preserve the original analyst) - Congressional Research Service reports (CRS) - NVIDIA / TSMC / ASML / AMD / Intel / Samsung / SK Hynix / Micron earnings transcripts and investor-day presentations - SemiAnalysis published decompositions (the 100K H100 cluster decomposition is the most-cited) - The Information's industry coverage - Stratechery / Asianometry / Asianometry analyses **Editorial / inferred (Tier 4):** - Scrutica-derived estimates via documented methodologies; flagged `is_estimated`; bounds reported. The methodology page documents every Tier-4 derivation chain. --- ## Contact and governance **Editorial.** davidgringras@hsph.harvard.edu **Corrections, data contributions, and methodological critiques welcome.** The corrections log at https://scrutica.com/corrections records every change. The pattern: send the disagreement with the specific Scrutica URL where the value appears, the alternative figure with its source citation, and the rationale; editorial dispositioning is logged and the resolution surfaces publicly. **Built by David Gringras**, MD/MPH candidate at Harvard (Frank Knox Fellow; MPH Health Policy; cross-registered at MIT, Harvard Law School, and the Harvard Kennedy School). Evaluations and Collaborations Lead, FATF-to-AI governance translation project at Arcadia Impact / The Future Society. Project supervisor, Orion AI Governance. **Recent peer-reviewed and preprint work** informing Scrutica's methodology and audience understanding: - Safety Under Scaffolding (https://arxiv.org/abs/2603.10044) - IatroBench (https://arxiv.org/abs/2604.07709) - Defensive AI (https://papers.ssrn.com/abstract=6402098) - Frontier Lag (https://arxiv.org/abs/2605.04135) Built at Harvard as part of ongoing research on AI governance infrastructure. **License (data).** CC BY-SA 4.0 except where source-attributed otherwise. Source datasets that originate in licensed corpora (held under subscription; not redistributed) are stored as extracted relational facts with provenance preserved; the raw paywalled datasets themselves are not redistributed. Downstream consumers that re-distribute Scrutica's substrate must preserve the BY-SA terms (their downstream consumers also receive the data under CC BY-SA 4.0). **Citation discipline.** Citing Scrutica in your work should include the methodology version stamp (`v2.1`, etc.) and the estimation path where applicable (e.g., "Hardware Path"). This lets future readers reproduce the calculation against the same parameter set even after Scrutica's methodology evolves. **The discoverability layer.** This file (`/llms-full.txt`) is the long-form agent-consumption format described at https://llmstxt.org/. The navigational index version is at `/llms.txt`. Per-page Markdown alternatives (when the L3 route handler ships) are accessible at any URL with `.md` appended. The MCP server at `/api/mcp` exposes the substrate via the Model Context Protocol for direct agent integration. --- *End of llms-full.txt — generated 2026-07-23 from substrate snapshot at 2026-07-23T03:54:36.316Z. Word count target: 30K+. Regenerate via `npx tsx scripts/build-llms-full-txt.ts`.*