Select an item in the diagram to open the relevant method. Use the worked examples to test how different assumptions affect the estimates.
versions.ts GPU_MODELS — per-chip dense throughput and TDP quoted from NVIDIA and AMD datasheetscost-index-section.tsx — per-GPU dense BF16 throughput sourced from the Epoch AI ML Hardware datasetvalidation-section.tsx VALIDATION_SOURCES — MLPerf hardware configurations cross-checked against the chip cataloguevalidation-section.tsx comparison approach — the 0.40 sustained default calibrated to PaLM 540B at 0.462 and LLaMA 3 405B at 0.384cost-index-section.tsx — the cost unit denominated in the same Epoch-sourced dense BF16 throughputcost-index-section.tsx — published hourly instance pricing normalised into the compute unitcascade-section.tsx — criticality derived from the supply-chain relationship rank, blended with market correlationownership-investigator-section.tsx — the chain-of-title walk reads ownership edges populated from filingsconfidence-level-guide.tsx — tier 1 is a primary regulatory filingownership-investigator-section.tsx — a designation, and every affiliate chain walked down from one, cites the Federal Register notice the ladder rates tier 1trade-flow-attribution-section.tsx — the HS 8542 filter and counterfactual window applied to declared customs valuesderivations.ts — peak throughput from GPU count times per-chip dense throughputderivations.ts — the training-throughput step multiplies effective throughput by the MFU assumptionderivations.ts deriveFromHardware — the bounds are the MFU range at its floor and ceilingderivations.ts deriveFromHardware — peak throughput is the common factor in both boundsflop-methodology-section.tsx — time-to-threshold divides the statutory threshold by the daily FLOP budgetflop-methodology-section.tsx — the same daily budget applies the per-chip throughputcost-index-section.tsx — the index is the compute unit priced per provider and regioncost-index-section.tsx — dense BF16 throughput is the denominator of the cost unitcascade-simulator.ts — each hop multiplies the impact arriving at it by the edge’s flow share, by its criticality (a 1–10 score, divided by ten), and by the scenario’s per-hop decaytrade-flow-attribution-section.tsx — corridor change measured against the pre-event baseline windowConnections omitted because no calculation is documented:
Record counts are from the snapshot dated 2026-09-09, including licensed records that cannot be browsed individually. Dense FP16/BF16 throughput is recorded for 23 of 24 accelerator records; one has no value in that field.
For regulatory and market records, the age labels change to “aging” after 7 days and “stale” after 30 days. Structural records have thresholds of 90 and 365 days; narrative records, 180 and 365 days. These labels describe the age of a record, not whether its contents have changed.
| Coverage | |||
|---|---|---|---|
| SEC EDGAR / XBRL Primary source for US-listed company financials. Capex, revenue, and subsidiary disclosures from regulatory filings. | T1: Government / Filing | Quarterly (10-Q, 10-K) | US-listed companies: financials, capex, segment data, Exhibit 21 subsidiaries |
| BIS Entity List US Department of Commerce list of entities subject to export license requirements. Primary source for export control status. | T1: Government / Filing | As amended | Entity List entries and associated licensing information |
| GAO / SIA CHIPS Act reports Government assessments and semiconductor-industry summaries of CHIPS Act implementation. | T1: Government / Filing | As published | Semiconductor manufacturing investments and federal funding |
| Epoch AI infrastructure datasets Research datasets assembled from public reports and other sources. Cluster upgrades may appear as separate records; location detail and measurement uncertainty vary by entry. | T2: Research Database | Monthly | GPU clusters and frontier data centres, including hardware, power, location and operating dates |
| CSET Supply Chain Data Georgetown CSET analysis of semiconductor supply chain relationships and market share data. | T2: Research Database | Annual | Semiconductor supply chain topology: equipment, foundry, packaging relationships |
| NVIDIA Earnings Transcripts Company disclosures on sales, customers and demand for AI hardware, including commentary on sovereign AI programmes. | T2: Research Database | Quarterly | Sovereign AI revenue, customer disclosures, GPU shipment context |
| Licensed corporate-ownership / fund database Licensed records of ownership, investors and fund commitments. Transaction amounts may be undisclosed; access to the underlying records is restricted. | T3: Industry / Analyst | Daily | Private capital deals: DC investments, AI infrastructure funding rounds |
| Licensed corporate-event feed Corporate announcements classified into event records, including capacity additions and investments. Access to the underlying licensed feed is restricted. | T3: Industry / Analyst | Daily | Corporate announcements: facility expansions, capacity additions, GPU deployments |
| Grid Queue Filings (PJM) Proposed connections to the electricity grid. These records concern electricity supply and transmission; they are not a register of data centres awaiting power. | T3: Industry / Analyst | Monthly | Generation and transmission service requests, with project capacity and status |
| Sovereign AI Research Files Compiled from government announcements, NVIDIA earnings transcripts, CRS reports, and press coverage. | T3: Industry / Analyst | As published | 36 country sovereign AI programs: announced investment, government vs. private split, deployed-vs-announced reconciliation |
| Press Reports / Industry Analysis Secondary reporting may contain both quoted company figures and a journalist’s or analyst’s estimates; the source type alone does not distinguish them. | T4: Press / Secondary | As published | Facility announcements, capacity rumors, deployment speculation |
Coverage depends on the source, geography and type of record. The counts below refer to the records held; a coverage percentage also requires a defined comparison set.
| Dataset | Records | Sources and scope | Coverage statement |
|---|---|---|---|
| Facilities | 4,573 across 110 countries |
| Facility records by source: 70 from the Epoch frontier-DC dataset, 437 from Epoch GPU clusters, 1,569 from PeeringDB, 1,173 from the expanded OpenStreetMap dataset, 391 from IM3-PNNL, and the balance from cloud-provider documentation, grid-queue discovery and direct research. These totals count the source recorded on each row; they do not measure overlap with each publisher’s full dataset. |
| Supply-chain edges | 18,980 supply-chain records across all sources |
| After duplicate source records are merged and records naming only one counterparty are excluded: 14,937 supplier–customer relationships, with each product or service counted separately. |
| Export-control designations | 3,435 records (3,420 without a recorded removal date) |
| The total includes historical records and related export-control authorities. 3,434 records are classified as explicitly named entities; 20 have no entry in the Federal Register citation field. These counts describe the records held, rather than completeness against the current Entity List. Organisation matching is recorded separately: 74 matches use name similarity and 102 use an affiliate chain. |
| Sovereign AI programs | 36 country / program profiles |
| The 36 profiles are individually researched, with sources listed for each programme. The count is not measured against a global programme list and should not be read as a coverage percentage. |
| FLOP / compute capacity estimates | 4,093 facility-level deployments |
| The platform-level rate counts tracked facilities where at least one path produces a bounded output; per-path coverage is stated in FLOP estimation. The hardware path produces narrow bounds, and the power and cost paths widen them. |
| Compute Cost Index | GPU rental prices and modelled ownership costs |
| Rental prices are attributed to the provider’s published rates or to a third-party price mirror. On-premises costs are model estimates. The Compute Cost Index includes the source and date for individual prices; provider and instance coverage varies. |
| Ownership chains | 103,614 directed edges; 130,394 organizations; 15,364 aliases |
| Public-filing research targets selected companies and is combined with licensed ownership data. Coverage has not been measured against a defined set of companies or ownership relationships. |
| Procurement and grant records | Awards, solicitations and funding opportunities |
| Source-specific filters select records for review of AI relevance. Solicitations and funding opportunities are distinct from awards; their advertised budgets are not evidence of money spent. |
| Semiconductor trade | Trade values by country pair, product code and reporting period |
| Customs product codes group many chip types together. Trade values cannot identify individual accelerator models or establish how many GPUs crossed a border. |
| Chip sales (Epoch AI) | Quarterly and cumulative chip-sales estimates |
| These estimates concern chip sales, with uncertainty intervals where supplied. They are separate from the customs trade records above and do not establish where the chips were installed. Source: Epoch AI. |
Every record carries data_source (which corpus it came from), source_url (the document), and is_estimated (whether the value was reported or derived).
Capacity values are versioned in an append-only table, compute_capacity_snapshots: a change writes a new row stamped with the date the new value was learned, and the old row stays.
An unknown value is set to null. A blank can be filled once a source turns up; a guess that has entered the table is afterwards indistinguishable from a measurement.
Facility records come from Epoch AI's Frontier Data Center catalog (Tier 2), the IM3 OpenStreetMap data-center bulk corpus (Tier 3), PeeringDB facility participation tables (Tier 2), a licensed corporate-ownership database (Tier 2), SEC EDGAR primary filings (Tier 1), gridstatus.io interconnection-queue data (Tier 2) and operator press releases (Tier 3). Supply-chain edges come from licensed supply-chain databases (Tier 2), SEC Exhibit 21 subsidiary filings (Tier 1) and CSET entity classification (Tier 2). Regulatory events are pinned to Federal Register documents (Tier 1).
Four authority tiers run from T1, a primary measurement or filing, down to T4, an estimate or an inference. A value takes the tier of the source it originated with: a CRS report citing a Goldman Sachs estimate carries the Goldman estimate’s tier, T3.
A derived value carries the tier of its weakest input: a GPU count estimated from power capacity (T3) combined with a disclosed power capacity (T1) is published at T3. In tier numbers that is the largest of them, T1 being the strongest tier.
aggregate_tier = max(input_1.tier, input_2.tier, ..., input_n.tier)
Where sources disagree on a value, Scrutica publishes both with their attributions, names the one it is treating as primary, and records possible explanations for the difference, such as the date, method or scope.
is_estimated FlagThe flag records whether a value was derived through a model or a proxy method. Uncertainty and missingness are recorded separately: bounds where an estimate has them, and null where a value is unknown.
Value from primary source with direct measurement. SEC 10-K capex, company-disclosed GPU count, government registry coordinates.
Value produced by inference: GPU count from power capacity, FLOP estimate from investment amount, market share from analyst reports.
Some supply-chain and ownership relationships first reached Scrutica through a licensed database and have since been found in a public primary document; those now cite the public document's URL in place of the vendor's. Re-sourcing one takes a public filing that confirms both the relationship and its direction, and a verbatim quote naming both parties, before the record is re-labelled. Relationships that no public document discloses keep their licensed provenance and stay off anonymous surfaces. Three families result, named for where the public document sits: sec-edgar-rederived (a filing on sec.gov; Tier 1), public-registry-disclosure-rederived (a foreign primary registry: SIRENE, the Handelsregister, Companies House, HKEX, DART, MOPS, and equivalents; Tier 1), and company-disclosure-rederived (a company's own press release, IR page, or current report; Tier 1). A re-sourced relationship counts as a public filing wherever the site reads provenance: it ranks at primary-source authority in the chain-of-title walk, and it renders the public document's own URL. Live counts per family appear in the source breakdown on the Supply Chain dependency graph and the cascade simulation.
Three paths run independently, one for each kind of disclosure a facility makes, and the bounds widen as the input becomes more indirect.
Highest Confidence
Counts the disclosed GPUs and reads their published throughput.
peak = gpu_count × fp16_tflops × 1e12
Medium Confidence
Works back from a facility’s disclosed power capacity to the number of GPUs that capacity supports.
gpus = power × 1000 / pue × gpu_frac / tdp
Lowest Confidence
Breaks a reported capital budget down into the hardware it buys.
gpus = invest × gpu_frac / cost_per_gpu
Throughput figures are dense BF16, with no 2:4 structured sparsity applied. NVIDIA publishes identical FP16 and BF16 tensor throughputs for every part in this table, so the two readings coincide; the engine specifies BF16, and the underlying field is still spelled fp16_tflops in the formulas below.
| Range | |||
|---|---|---|---|
Interconnect Efficiency Additional throughput multiplier applied by the estimator alongside utilization; its default is Scrutica's choice. Published end-to-end utilization already includes communication losses, so a measured utilization rate substituted here already carries the interconnect discount. | 0.85 | 0.60–0.95 | Scrutica modeling assumption |
Model FLOP Utilization (MFU) Utilization multiplier applied to dense peak throughput. PaLM 540B reports 46.2% model FLOP utilization and 57.8% hardware FLOP utilization; the latter includes rematerialization, and both figures describe that run. The default is Scrutica's choice, and the estimator additionally applies its interconnect multiplier. | 0.40 | 0.20–0.50 | Scrutica modeling assumption; benchmark context: PaLM, Table 3 |
Path A: Hardware-Based has no parameter of its own; it runs on the shared ones above.
| Range | |||
|---|---|---|---|
GPU Fraction of IT Load Assumed share of IT power allocated to accelerator chips. The remainder covers other IT equipment; cooling outside the IT boundary is handled through PUE. This fraction does not establish a facility’s rack composition or operating GPU count. | 0.49 | 0.35–0.65 | Scrutica modeling assumption |
Power Usage Effectiveness (PUE) Total facility power divided by IT-equipment power. The default converts a reported facility power figure to an assumed IT load; a site's own disclosure replaces it where one exists. Use a disclosure with matching facility scope and operating period when available. | 1.15 | 1.05–1.45 | Scrutica modeling assumption |
| Range | |||
|---|---|---|---|
GPU Fraction of Capex Assumed share of a project’s capital expenditure allocated to accelerator purchases. A project budget may cover other phases, land, buildings and infrastructure; matching the investment’s scope to the equipment being estimated is necessary before using this conversion. | 0.45 | 0.30–0.60 | Scrutica modeling assumption |
Highest Confidence. Counts the disclosed GPUs and reads their published throughput.
EU AI Act 1025 thresholdAt the selected daily budget, 1025 FLOPs equals 34 days of compute
Under Article 51(2), a general-purpose AI model trained with more than 1025 FLOPs is presumed to have high-impact capabilities. Article 52 requires provider notification and permits a substantiated argument against systemic-risk classification. Providers of models classified with systemic risk must assess and mitigate those risks under Article 55.
The index divides published rental prices by hardware throughput to compare the cost of compute across GPU configurations. Cloud rates price peak capacity; the on-premises calculator also accounts for assumed utilisation.
1 petaFLOP-day = 1 petaFLOP/s sustained for one day
= 8.64 × 1019 FLOP
The cloud calculation divides daily instance cost by GPU count and per-GPU dense BF16 throughput in PFLOP/s.
A lower hourly price can buy less compute if the instance contains fewer or slower GPUs. Price per petaFLOP-day accounts for that difference. Workload efficiency still determines how much training those rented hours accomplish.
Cloud converts a quoted rate at peak throughput into the price of a petaFLOP-day of purchased capacity; that is the conversion the Cost Index applies to every published rate. On-premise divides lifetime cost by output at MFU, so it prices a petaFLOP-day delivered. Purchased and delivered are different denominators, so the two are not competing estimates of one number.
BF16 TFLOP/s dense, no 2:4 structured sparsity (source: Epoch AI ML Hardware dataset). On-prem defaults: MFU , PUE , $5K networking + $3K facility share per GPU (editorial estimates, tier 4).
On-demand / spot pricing for GPU compute from major cloud providers (AWS, GCP, Azure, Lambda, CoreWeave).
Cloud provider pricing APIs · Historical spot price tracking
1-year and 3-year reserved instance pricing. Committed-use discounts.
Cloud provider pricing pages · Enterprise contract disclosures
Total cost of ownership for self-operated GPU clusters including hardware, power, cooling, facility, and labor.
Hardware vendor ASPs · Electricity tariff databases · Construction cost indices
Effective cost for government-funded sovereign AI compute including subsidies, grants, and below-market financing.
CHIPS Act disclosures · EU Chips Act awards · Sovereign AI program announcements
| Data source | Cadence | Method |
|---|---|---|
| Cloud Provider APIs (Azure, OCI; AWS spot) | Daily | Vercel cron |
| AWS Pricing Bulk API (on-demand) | Daily | GitHub Actions |
| Cloud Provider Pricing Pages (GCP, CoreWeave, Lambda) | Manual refresh | Manual snapshots of published pricing pages |
| Enterprise Contract Disclosures | Quarterly | SEC filing extraction |
| Hardware Vendor ASPs | Quarterly | Earnings transcript parsing |
| Electricity Tariff Data | Annually | EIA / regional utility databases |
| Construction Cost Indices | Annually | Turner / RSMeans indices |
The catalogue holds 20 model records, with 18 cost figures and 19 recorded ECI scores. The scatter sets training cost on a logarithmic x-axis against Epoch’s Capabilities Index. Each cost keeps its source’s dollar basis, 2023 USD for Epoch’s estimates, and covers the training run alone, with no multiplier applied for total development.
The model records are defined in MODELS in src/components/training-cost/TrainingCostCliff.tsx. trainingCost() returns each record’s training-cost field as the source recorded it. The cost trajectory retains successive maximum costs in release-date order across the 18 records with costs. The replication curves use only the 18 records with both a cost and an ECI score: at each release date, they take the lowest recorded cost among models meeting ECI 115, 120, 125, or 130. computePareto() sorts the scored records by cost and retains successive record-high ECI scores for the optional frontier line. Trinity Large stays in the table, the model selector and the export; its announcement gives no separate training-cost figure, and no ECI score is recorded here, so it enters neither cost nor capability curves.
One curve records each new high in training cost. The other curves show when a cheaper model reached a specified capability score. They use the costs attached to those historical models, and do not recalculate an old run on current hardware.
Epoch describes ECI as a scale with an arbitrary origin: score differences are meaningful, and an absolute score has no standalone interpretation. Dividing a raw score by cost would make the resulting ranking depend on where that origin was set, so cost and ECI stay separate operands, with no score-per-dollar ranking and no ratio isolines.
A ten-point ECI difference can be compared with another ten-point difference. A score of 140 does not mean twice the capability of 70.
Selecting a scored Chinese model finds the nearest ECI score among the non-Chinese records; selecting another scored model searches the Chinese records. The comparison reports the larger training cost divided by the smaller, on the source dollar bases above, together with the absolute score difference and the cheaper model’s name. Two models with similar scores may have been trained to different objectives and have different cost coverage, so the ratio is not evidence of an efficiency advantage.
MODELS in src/components/training-cost/TrainingCostCliff.tsx accompany cost-confidence labels, authority tiers and a separate estimated-value flag. It is that flag that records whether a cost was disclosed or estimated; the catalogue holds both, at the same authority tier.The simulation walks outward from the disrupted supplier through a weighted directed graph of supply-chain relationships, breadth-first, one hop at a time. The impact carried forward at each hop is parentImpact × flowShare × (criticality / 10) × propagationDecay. flowShare is the share of the customer’s input for that product that comes from this supplier, and defaults to 30% where no share is recorded. criticality on a licensed supply-chain database edge is derived from the source’s relationship-rank field (seven buckets: 2, 3, 4, 5, 6, 7, 8), and is editorial on curated-seed edges. propagationDecay is a per-scenario judgment of how readily the input can be sourced elsewhere. The licensed database’s annual_value_usd rides on each edge for provenance and never enters the formula. The reader sets shock severity, propagation delay and the threshold cutoff, and the walk stops when severity falls below 5% or after six hops from the origin, whichever comes first.
A disruption at one supplier propagates outward through the AI supply chain; the simulation traces the shock hop by hop and surfaces who loses access, to what, and for how long. Some inputs (ASML EUV lithography) have no substitute; others (power delivery, construction) source from many vendors. The decay rate at each hop reflects the asymmetry.
// BFS forward propagation — simulateForwardDisruption() in src/lib/compute/cascade-simulator.ts
function simulateForwardDisruption(graph, scenario) {
// Seed nodes get full severity — no decay applied to the source of disruption
for each seed in scenario.affectedNodes:
seed.currentImpact = scenario.severity
for each edge in seed.customers:
propagated = scenario.severity × edge.flowShare × (edge.criticality / 10)
if propagated > MIN_IMPACT_THRESHOLD: enqueue(edge.target, propagated, depth = 1)
while queue is not empty:
(node, incoming, depth) = queue.dequeue()
// Apply per-hop substitutability decay (non-seed nodes only)
decayed = incoming × scenario.propagationDecay
if decayed ≤ node.currentImpact: continue // multi-path: skip if no improvement
node.currentImpact = decayed
if depth < MAX_HOPS:
for each edge in node.customers:
propagated = decayed × edge.flowShare × (edge.criticality / 10)
if propagated > MIN_IMPACT_THRESHOLD: enqueue(edge.target, propagated, depth + 1)
}
Constants: MAX_HOPS = 6; MIN_IMPACT_THRESHOLD = 0.05; per-node visit cap = 3 (multi-path performance guard).propagationDecay is a per-scenario constant in the 0.65–0.95 range; the table below gives each scenario’s value and the reasoning behind it. The reverse walk, from customer back to supplier, uses approximateRevenueShare in place of flowShare to model demand loss propagating upstream.
| Parameter | Default | Range | Description |
|---|---|---|---|
propagation_delay | 30 days | 1–365 | Time for a disruption to propagate from one supply chain node to the next. |
initial_shock_severity | 1 | 0.1–1 | Severity of the initial disruption at the source node (1.0 = complete disruption). |
Complete loss of TSMC Taiwan fab capacity (natural disaster, blockade).
ScenarioTSMC Taiwan Disruption
Simulated at decay0.90Affected organizations423 hops
Results from BFS propagation over the live supply-chain graph (5 scenarios, 14 decay points each). For the full animated simulation, see the Cascade Simulation. Higher decay means more severe propagation: each hop retains a larger fraction of the disruption.
These per-scenario propagationDecay constants are editorial tier-4: hand-authored per scenario against market-share data, sole-source status and substitution-path qualification. No public dataset of semiconductor cascade propagation exists to calibrate them against. They are an ordering, anchored at 0.95 for a near-monopoly and 0.65 for a category with meaningful alternatives. Source of truth: PREBUILT_SCENARIOS in src/lib/compute/cascade-simulator.ts.
| Scenario | Decay | Rationale |
|---|---|---|
| TSMC Taiwan Disruption | 0.90 | TSMC is deeply embedded in advanced-node logic, but some customers maintain Samsung Foundry and Intel Foundry relationships at lower nodes — a non-trivial fraction of impact is absorbed at the first hop. |
| ASML Export Halt | 0.95 | EUV lithography is irreplaceable — no alternative supplier exists at any node ≤7nm, so impact propagates almost fully downstream. |
| US–China Full Decouple: Revenue Impact | 0.70 | Domestic substitution (Huawei Ascend, SMIC at advanced nodes, domestic EDA pilots) dampens each upstream hop as US revenue loss propagates back through the supply chain. |
| US–China Full Decouple: Access Denial | 0.80 | Upstream dependency on US tools (EDA, advanced equipment, foundry access) is structural for Chinese AI compute consumers — domestic substitutes exist but cover narrow node ranges, so impact propagates with little dampening per hop. |
| CoWoS Advanced Packaging Bottleneck | 0.70 | Multi-source packaging via OSAT partners (Amkor, SPIL handling 240–270K wafers/year overflow) now dampens downstream propagation; CoWoS capacity tripled to ~75K wafers/month by end-2025 per TrendForce. |
| SK Hynix HBM Disruption | 0.80 | HBM is required for every AI accelerator and SK Hynix has ~57% market share, but Samsung and Micron are pre-qualified for HBM3E and can ramp without re-qualification — multi-sourcing dampens propagation modestly per hop. |
| EDA Export Restriction | 0.80 | Synopsys and Cadence jointly control ~75% of the global EDA market; open-source alternatives (qflow, OpenROAD) exist at trailing nodes but not at production scale for advanced logic, so most downstream chip-design work cannot continue without them. |
| Chiplet Assembly Disruption | 0.75 | Compound packaging + EDA disruption hits chiplet-architecture chips (AMD MI300X, NVIDIA Blackwell) disproportionately; some single-die monolithic designs and trailing-node products are insulated, providing modest per-hop attenuation relative to the blanket scenarios. |
| Helium Supply Crisis (Hormuz) | 0.75 | US strategic helium reserve and alternative shipping routes (Australian, Algerian, US domestic LNG-byproduct streams) provide some buffer; helium boil-off caps stockpiling at roughly 45 days. |
| InP/EML Laser Shortage | 0.80 | Silicon photonics provides a partial alternative path for some 800G+ transceiver designs, but InP epitaxy is the binding bottleneck — Coherent / Lumentum / Source Photonics dominate the production base and pre-allocations push other buyers past 2027. |
| ABF Substrate Monopoly (Ajinomoto) | 0.95 | Ajinomoto holds 95–98% of global ABF substrate market with no meaningful second source; glass substrates are 2028–2030 at earliest, so impact propagates almost fully through every AI accelerator package. |
| Gallium/Germanium Restriction (Nov 2026) | 0.65 | Recycling, stockpiling, and alternative sourcing (allied gallium production restarts, germanium recovery from zinc smelting) dampen downstream impact; this is a binary deadline scenario, not a current disruption, so the per-hop attenuation reflects optionality rather than realized substitution. |
| CoWoS 2026 Allocation Constraint | 0.85 | Structural shortfall with no alternative CoWoS-L packaging at volume; OSATs cannot replicate Blackwell-class packaging at scale, and NVIDIA holds more than half of TSMC’s CoWoS allocation (above 60% on Morgan Stanley and Bernstein estimates, above 50% on TrendForce’s), so non-NVIDIA AI GPU production absorbs the constraint with limited dampening. |
| Japan Semiconductor-Input Seismic Shock (Korea coupling) | 0.82 | Silicon-wafer and advanced-photoresist supply is highly Japan-concentrated with thin qualified second sources, so the first hop from the input makers carries most of the shock; downstream memory customers hold some inventory buffer and partial multi-sourcing, which dampens each subsequent hop. |
| Allied Coordination: TSMC Advanced Node Restriction | 0.90 | Substitutability is unchanged from the blanket TSMC disruption (Samsung Foundry / Intel Foundry absorption capacity is identical regardless of who is restricted); what separates this from a blanket outage is a country test applied to every edge, which zeroes impact on allied-country customers and leaves everyone else at full impact, rather than anything in the per-hop decay. |
| EDA License Restriction: China Entity List Expansion | 0.80 | Substitutability is unchanged from the blanket EDA scenario; ~75% market concentration of Synopsys + Cadence holds whether the restriction targets China specifically or all customers, so per-hop attenuation tracks the underlying tool dependency rather than the policy framing. |
| Netherlands DUV Export Restriction to China | 0.85 | DUV is replaceable at mature nodes (Nikon and Canon retain ArF and KrF capacity for nodes ≥28nm), so per-hop attenuation is meaningfully greater than for EUV — but advanced-DUV ArF immersion remains effectively sole-source ASML for 7nm-class multi-patterning work. |
| China Critical Minerals Retaliatory Restriction | 0.65 | Substitutability tracks the gallium-germanium-deadline scenario (recycling, stockpiling, allied gallium production restarts dampen each hop); only the directional framing differs — China restricting allies vs. allies anticipating a deadline lapse. |
| Multilateral Cloud Compute Access Restriction | 0.75 | Restricted-jurisdiction customers retain some domestic cloud alternatives (Alibaba Cloud, Tencent Cloud, Yandex, regional sovereign clouds) but the GPU-density and frontier-model-availability gap with US hyperscalers is structural — partial substitution dampens propagation without closing it. |
| Netherlands EUV Restriction: Allied Countries Only | 0.95 | Substitutability is unchanged from the blanket ASML export-halt scenario — EUV remains zero-alternative-supplier regardless of which jurisdictions retain access; the allied-only carve-out is a country test applied to every edge, zeroing impact on allied-country customers, rather than anything in the per-hop decay. |
Category-level substitutability, scored 0 to 1: 0.02 for EUV lithography, where ASML is the sole source, up to 0.50 for data-centre construction, where contractors are interchangeable. These scores are a domain reference with their own provenance in market shares, sole-source status and capacity data; the simulator runs on the per-scenario propagationDecay constants in the table above, which are separate editorial assessments about the same subjects. The two scales point opposite ways: a low score here means no substitute exists, while a high decay there means a shock travels far.
| Category | Substitutability | Description |
|---|---|---|
| EUV Lithography | 0.02 | Near-zero substitutability. ASML sole source. |
| Advanced Foundry (<7nm) | 0.05 | TSMC dominant. Samsung limited alternative. |
| GPU Accelerators | 0.15 | NVIDIA dominant but AMD MI300X is a partial substitute. |
| HBM Memory | 0.10 | SK Hynix ~57% market share, Samsung ~22%, Micron ~21% (Q3 2025, Counterpoint Research memory market tracking; Samsung’s share recovered from ~15% in Q2 2025 on the HBM3E NVIDIA qualification cycle). HBM4 mass production began Feb 2026 (SK Hynix at M16/M15X; Samsung at Pyeongtaek). Concentrated but multi-source. Updated 2026-05-15. |
| Advanced Packaging (CoWoS) | 0.15 | TSMC dominant for CoWoS; capacity ramping 75-80K wafers/month (early-mid 2026) toward 120-130K by end-2026 (TrendForce + TSMC Q1 2026 earnings; CoWoS yield reported >98% May 2026). OSAT partners (Amkor ~180-190K, SPIL ~60-80K wafers/year) handle overflow. CoWoS-L/S reported fully booked through 2026 — the bottleneck analysts had called "effectively resolved" is now allocation-bound (NVIDIA >50% of capacity) rather than capacity-bound. Substitutability still meaningful from OSAT overflow. Source: TrendForce + TSMC IR (Tier 2). Updated 2026-05-19. |
| AI Networking (InfiniBand + Ethernet) | 0.30 | Ethernet surpassed InfiniBand in AI back-end network revenue in 2025 (Dell'Oro Group). UEC 1.0 (June 2025) achieves InfiniBand-class performance. Meta validated Ethernet RoCE at 24K GPUs. Higher substitutability than pre-2025 InfiniBand monopoly. Trajectory unchanged through Q1 2026; no material movement in InfiniBand-vs-Ethernet share. Updated 2026-05-19. |
| Data Center Construction | 0.50 | Many EPC contractors globally. Highly substitutable. |
| Power Delivery | 0.25 | Regional utilities, PPAs, on-site generation available in theory, but PJM capacity prices surged 10x ($329/MW-day for 2026-27 vs $29/MW-day for 2024-25). AEP Ohio paused new DC interconnections. FERC issued a Dec 18, 2025 final order directing PJM to file new co-location tariffs (compliance Jan 20 + Feb 16, 2026); reforms are progressing but the queue remains binding (TC1 process ~1.75y entry-to-cluster-end vs the 1y efficient-process target). Power remains a binding constraint on expansion. Substitutability lower than pre-2025 assessment. Source: PJM capacity auction (Tier 1); FERC final order (Tier 1). Updated 2026-05-19. |
Each supply-chain layer holds a different amount of inventory, which delays the point at which a disruption reaches the layer below it. Buffer figures come from 10-K and annual-report inventory line items (Tier 1), TrendForce and SemiAnalysis (Tier 2) and analyst estimates (Tier 3).
| Layer | Buffer | Range | Shape | Tier |
|---|---|---|---|---|
| EUV Lithography | None | 0–0 wk | cliff | T1 |
| Advanced Wafers (N3/N5) | 10 wk | 8–16 wk | linear | T1 |
| Packaged Chips (GPU) | 3 wk | 2–5 wk | cliff | T2 |
| Server Assembly | 14 wk | 8–26 wk | exponential | T3 |
| HBM Memory | 3 wk | 0–8 wk | cliff | T2 |
| ABF Substrates | 12 wk | 4–30 wk | linear | T2 |
Qualification time is the lag before a substitute produces anything; the capacity ceiling is the largest fraction of the disrupted supply it can absorb. Output ramps along an S-curve between the two.
| Disrupted supplier | Substitute | Qualification | Ceiling | Tier |
|---|---|---|---|---|
| TSMC Advanced Node | Samsung Foundry | 24mo | 15% | T2 |
| Intel Foundry (IFS) | 36mo | 5% | T2 | |
| ASML EUV | No viable substitute | |||
| SK Hynix HBM | Samsung HBM | 0mo | 45% | T2 |
| Micron HBM | 0mo | 25% | T2 | |
| NVIDIA Training GPUs | AMD MI300X/MI350 | 12mo | 15% | T2 |
| TSMC CoWoS Packaging | Amkor (OSAT) | 18mo | 20% | T2 |
| SPIL (OSAT) | 18mo | 10% | T2 | |
The model’s buffer-depletion and substitution-ramp parameters are calibrated against three observed disruptions: a sudden fab shutdown, a commodity shock absorbed by inventory, and a sustained packaging bottleneck.
Texas Freeze (Samsung Austin S2)
Ukraine Neon Gas Supply
CoWoS Packaging Bottleneck
Cascade propagation runs server-side over the whole dataset: 18,980 source rows across every source, licensed and public, entity-resolved and deduplicated to 14,937 supplier–customer–product keys. Rows naming only one endpoint (“this supplier sells to N unnamed customers”) are dropped before that key count is struck, because a cascade needs two named ends to travel between; the engine rebuilds the same graph live and walks 14,937 edges across 5,080 organizations, which is the figure the simulation API reports. These counts read from the dataset snapshot of 2026-09-09 and move with every refresh, so the engine’s live count and the snapshot’s can sit a refresh apart. Duplicate rows describing the same relationship merge keeping the strongest weight, so the propagation graph’s coverage runs higher than the raw rows’ (figures below give both).
Edge weights draw on three signals. Supply share (% of customer input from this supplier) was available on 0.27% of propagation edges (0.24% of raw rows); the remainder uses a default of 30%. Criticality (1–10 replaceability) was populated on 71.78% of propagation edges (51.71% of raw rows): the primary licensed supply-chain database derives it from the source’s relationship-rank field, which has seven distinct values, 2 through 8; two smaller re-derived disclosure sources contribute at lower coverage (12.7% and 64.8% of their raw rows respectively); curated-seed edges, under 1% of the graph, have hand-authored editorial scores; the rest (a second licensed database, CSET-ETO, and several structural sources with no source criticality) fall back to a default of 5. Sole-source edges are upgraded to criticality ≥ 9.
The third signal is optional market-correlation weighting (on by default, toggleable), which blends criticality with the 3-month stock-price correlation (Pearson r, available on 55.84% of propagation edges and 39.89% of raw rows, from the same 2026-07-11 measurement) via adjustedCriticality = w × 5(1+r) + (1−w) × current, where w defaults to 0.5. For edges without correlation data the unblended criticality stands alone, and sole-source edges retain criticality ≥ 9 regardless of correlation. Price correlations are from licensed market-data downloads held under subscription, not redistributed (vintage: April 2026).
The “total affected compute” metric blends two weightings: nodes with facility-level FLOP/day estimates are FLOP-weighted, and nodes without FLOP data are weighted by network degree centrality, the count of their supply-chain connections. The composite is therefore an approximation, and its degree-weighted half measures how connected a node is.
Every edge carries into the simulation payload the data_source it was ingested from, and the authority tier of that source. Neither enters the propagation formula, which runs on flow share, criticality and per-hop decay alone.
T2 · Equipment supplier → foundry, foundry → chip designer
Market-share-weighted edges. Topology only (no bilateral specificity).
T1 · Parent → subsidiary, supplier → customer
Both counterparties are named in the source.
T3 · Investor → facility, GP → infrastructure fund
Financial relationship edges. Deal amounts undisclosed ~45% of the time.
T2 · Organization → GPU cluster, cluster → facility
Training deployment edges with hardware type attribution.
Bilateral semiconductor trade (HS 8542 IC subheadings 854211/854219/854231/854232/854239/854290) is aggregated from five customs and statistical sources, then differenced pre and post each BIS rule date inside the bilateral comparison window. The headline change is the unadjusted difference; where a market-cycle confound applies (the 2020 COVID crash, the 2021 supercycle, the 2023 inventory correction), it prints as a caveat on that event’s card.
The table semiconductor_trade_flows ingests UN Comtrade, CEPII BACI (HS92, 231 reporters), Taiwan Customs DGCAS, China General Administration of Customs (GACC), and Japan e-Stat (Ministry of Finance customs statistics). Each row has its data_source, reporter_country/partner_country as ISO-3 codes, direction (export/import; China GACC uses “both”), value_usd, and year. The page reads via getTradeFlowImpactData() in src/lib/data/trade-flow-impact.ts; corridor-level analysis runs through analyzeCorridorImpact(). Reporter-code variants across vintages (US/USA, CN/CHN/CHINA) are normalized at query time. The HS-code filter is fixed in IC_HS_CODES.
For each US action inside the bilateral comparison window, the baseline is the average of the two calendar years before the event and the comparison is the calendar year after it (analyzeCorridorImpact()). Actions dated after the window closes, the January 2025 diffusion rule among them, are counted in the action tally and left undifferenced, because the Comtrade and BACI series end before their post-event year does. Where a comparison window straddles a known macro shock (the 2020 crash, the 2021 supercycle, the onset of the 2022 inventory correction, the 2023 trough), the event card carries the confound as a caveat, and same-year actions that share a window are listed together, so a corridor change can be read against every rule that shares its window.
Each BIS rule old enough to have a full year of customs data behind it is compared against a baseline of the two prior calendar years and the year right after; the more recent actions, the January 2025 diffusion rule among them, are listed and left undifferenced. The comparison applies no macro adjustment; where one of those years was a major macro event (the 2020 COVID crash, the 2021 post-COVID supercycle, the 2023 inventory correction), the event card says so. Where two policy actions land in the same year, both are listed against the corridor change they share.
Featured corridors per event are hand-curated in EVENT_CORRIDORS in src/lib/data/trade-flow-impact.ts and rendered first; on-demand discovery via discoverTopCorridors() returns the largest-magnitude changes outside the featured set, filtered to corridors with pre-event annual flow > $100M to suppress noise. Both layers render together so a reader can test the redirection question, whether Singapore, the UAE, Malaysia and other intermediaries absorbed what the US-to-China corridor shed: the corridors BIS named (Netherlands → China, Japan → China, Korea → China for the October 2023 expansion), and the auto-discovered corridors sharing the same window. Result rendering uses ISO-3 with a fallback English-name normalizer (normalizePartnerName()) for sources whose original-language names would otherwise read as encoding bugs (China GACC emits 韩国 for Korea).
Five derivation chains back the Investigator, one per section below. The chain-of-title walk climbs the parent graph a hop at a time and reports where it stopped. BIS designation reach walks downward from each Federal-Register-anchored designation to the affiliates it covers. Sovereign-LP attributions identify sovereign investors with commitments to funds whose portfolios include the company. Ownership-change events read ownership-transferring deals out of the private-equity deal tables. The licensed ownership database is the refresh path that fills the tables the other four read.
The Investigator answers five questions about an entity: who owns it (chain-of-title); whether an export-control designation reaches it through an affiliate (BIS designation reach); whether sovereign wealth funds hold indirect exposure to it through private-equity commitments (LP attributions); which mergers and acquisitions have touched it (ownership-change events); and which database the other four read from (the licensed ownership database).
Iterative parent walk over the ownership_upward_edges view, one Supabase round-trip per hop, capped at OWNERSHIP_WALK_MAX_HOPS = 8. Backing function: walkOwnership() in src/lib/data/ownership-investigator.ts. Returns a typed WalkResult with one of five terminal classes: reached_ubo, opacity, cycle_detected, depth_capped, contested. Alternates per hop are first-class: where two sources name different parents, the analyst sees both.
For any entity the Investigator can resolve, this chain walks upward through the corporate parent graph and reports where the walk terminated: reached a public-company or sovereign top, hit a structurally opaque shell, detected a cycle, ran out of hops, or found two documents naming different parents. When two sources disagree on a parent at the same step, both surface; the analyst decides which to cite.
ownership_upward_edges — SQL view fronting ownership_chain_edges with PARENT and SUBSIDIARY rows normalized to a uniform child → parent shape. Columns consumed: child_org_id, parent_org_id, data_source, confidence, as_of_date.ownership_chain_edges — source table joined per hop for source_url, the licensed-database relation id, and authority_tier (the upward-edges view drops these for breadth).organizations — canonical entity registry; per-hop parent display reads name, country_hq, cik, ticker, org_type for the public-company-vs-government terminal check.organization_aliases — name-resolution side table consumed by resolveEntity() (the resolver runs before this walk).canonical_org_slim.json — the map of which organization ids are variants of the same firm, read by getAllIdsForOrg(). It was built before the licensed ownership ingest, so at every hop the walker widens the id set by querying organizations for rows carrying the same licensed-source company id.// walkOwnership() in src/lib/data/ownership-investigator.ts
function walkOwnership(entityId) {
start = resolveCanonicalId(entityId)
visited = { start }
frontier = expandWithLicensedDbBridge(getAllIdsForOrg(start))
chain = [hop0(start)]
for hop in 0..MAX_HOPS-1:
raw = SELECT * FROM ownership_upward_edges WHERE child_org_id IN frontier
rows = filterPhantomSelfEdges(raw, frontier)
if rows.empty: terminal = looksOpaque(tail)? opacity : reached_ubo; break
edges = enrichEdges(rows) // join source_url + authority_tier + licensed-db relation id
candidates = groupBy(edges, canonicalParent).map(g => pickStrongestEdge(g))
candidates.sort(compareEdgeStrength) // HIGH→LOW conf, then tier, then recency
primary = candidates[0]; alternates = candidates[1..]
if primary.canonicalParent in visited: terminal = cycle_detected; break
visited.add(primary.canonicalParent)
chain.push({ primary, alternates })
if parent.cik && parent.ticker: terminal = reached_ubo; break
if parent.org_type == 'government': terminal = reached_ubo; break
if hop == MAX_HOPS-1: terminal = depth_capped
frontier = expandWithLicensedDbBridge(getAllIdsForOrg(primary.canonicalParent))
return { chain, terminal, terminal_detail }
}Each hop is a single Supabase round-trip, and it does double duty: the widening query returns both the next hop’s id set and the org row the parent renders from, so nothing has to be fetched again to draw the hop or to test the terminal condition on it.
compareEdgeStrength() ranks HIGH > MEDIUM > LOW on confidence, breaks ties on authority_tier (lower number = primary filing), and finally on as_of_date recency. pickStrongestEdge() reduces a per-canonical-parent group to a single representative; the rest become hop.alternates with full citation metadata so the page renders both. Both live in src/lib/data/ownership-investigator.ts.
filterPhantomSelfEdges()). Multi-source ingest periodically registers the same firm under two IDs and writes a parent edge between them; the canonical-org map merges them into one variant set, so the edge becomes a self-loop. Without the filter the walker would terminate cycle_detected on hop 1 for entities whose chains should walk cleanly. The 2026-05-02 sweep removed 13 such edges (Equinix, Broadcom, Lam Research, Amazon, Microsoft, Alphabet, Oracle, Digital Realty, NEC, Vornado, Credo, Kulicke & Soffa, SolarEdge).looksOpaque()). When parent rows run out, the walker checks whether the tail’s name or country matches an opaque-structure pattern (Cayman, BVI, Bahamas, Jersey, Guernsey, “HoldCo”, “SPV”, “nominees”, “blind trust”). Hits classify as opacity with the tail’s name as detail; misses classify as reached_ubo. The page renders the two terminals distinctly.organizations and adds whatever ids it finds, which is what lets an edge stored only against the synthetic form match when the walker asks the view for rows whose child is in that set.data_source, source_url, as_of_date, confidence and authority_tier, so the claim that lost the ranking can be weighed on the same evidence the winner was picked on.LOW in enrichEdges(). The default is dormant today (the licensed-database ingest set all 4,374 of its rows to HIGH on 2026-04-24) and takes effect the moment a source without confidence metadata is ingested.src/lib/data/stub-detection.ts. The synthetic id stays as the org’s canonical id, so deep links still resolve.HUMAIN, the Saudi sovereign-anchored AI infrastructure company launched May 13, 2025, walks upward through the ownership tables to PIF (Saudi Arabia’s Public Investment Fund) inside the hop budget. Every edge in that walk has a licensed corporate-ownership database data_source and the most recent as_of_date on file. The walk terminates on PIF as a state-linked sovereign vehicle: the public-company terminal predicate (CIK + ticker both populated) does not fire on PIF, so the chain rendering surfaces “reached UBO; sovereign vehicle”. The HUMAIN → PIF chain is the canonical cross-link to the sovereign-LP attributions section — PIF appears there as an LP across multiple PE funds whose portfolios touch other compute orgs.
HIGH by construction.HIGH but may be MEDIUM when the underlying disclosure is partial.HIGH, so on these edges the confidence label separates nothing and data_source with authority_tier carries the provenance.Two-stage match: a direct lookup on export_control_designations against the entity’s alias-expanded canonical id set, plus a downward affiliate-closure walk against designation_closures with three matching routes (direct entry / canonical-alias / one-edge extension into the compute universe). Backing function: getDesignationReachForEntity() in src/lib/data/ownership-investigator.ts; closure-walker at scripts/cascade-bis-designations.ts.
For any entity, this chain reports two findings separately: whether the entity itself appears on a BIS Entity List or equivalent designation, and whether a designated entity reaches it through a documented affiliate chain. Each renders with its own hop-by-hop chain and Federal-Register citation.
export_control_designations — Federal-Register-anchored designation registry (3,435 rows in the 2026-09-09 snapshot, 3,420 of them in force, alongside 4,408 closure rows in designation_closures).designation_closures — precomputed downward affiliate walks per designation, hop-stamped, with closure_chain JSONB with ordered {from_org_id, to_org_id, relation, hop} tuples.bis_crossref_matches — per-org cross-reference matches keyed on (designation_id, canonical_target_org_id), with match_type in {NAME_SIMILARITY, AFFILIATE_CLOSURE} and a confidence_score column. The data layer reads the live table at request time through the shared export-controls reader, excluding deprecated rows and taking NAME_SIMILARITY at ≥ 0.9 for the “directly named” bucket; the committed data/processed/bis_crossref_matches.json bundle is a canonical-map input to the offline pipeline. A backstop ILIKE on entity_name runs against the designation registry itself, catching entities the crossref pipeline never matched.BISCascadeHop interface (src/lib/data/export-controls.ts) has the per-hop hop_is_stub flag for downstream stub-aware rendering.// getDesignationReachForEntity() in src/lib/data/ownership-investigator.ts
function getDesignationReachForEntity(orgId) {
allIds = getAllIdsForOrg(orgId)
direct = []
for match in precomputed_crossref:
if match.confidence_score >= 0.9 && match.matched_org_id in allIds: direct.push(match)
// Belt-and-suspenders ILIKE on entity_name (catches entities the crossref pipeline missed)
direct ∪= SELECT * FROM export_control_designations WHERE entity_name ILIKE orgName
direct ∪= SELECT * FROM export_control_designations WHERE id IN allIds
closures = SELECT * FROM designation_closures WHERE canonical_closure_entry_id IN allIds
return direct.map(d => { ...d, hop_count: 0, is_direct: true, authority_tier: 1 }) ∪
closures.map(c => { ...c, is_direct: false })
}The closure walker that produces designation_closures rows runs offline (scripts/cascade-bis-designations.ts) with three matching routes:
// scripts/cascade-bis-designations.ts (offline; produces bis_crossref_matches AFFILIATE_CLOSURE rows)
(1) direct — closure entry IS a Scrutica compute-universe org
(2) canonical — closure entry resolves to one via canonical_org_map
(3) edge_extend — closure entry has a 1-hop PARENT/SUBSIDIARY/JV/AFFILIATE edge
// Edge-extension downgrades confidence one tier (HIGH→MEDIUM, MEDIUM→LOW, LOW→LOW)
// Cascade is DOWNWARD only — never walk a PARENT edge to find sibling subsidiaries.
// Per-hop hop_is_stub set when to_org_id matches the licensed-db stub prefix.Direct rows have authority_tier = 1 by construction (Federal Register primary). Closure rows have the row-stored authority tier (typically 2 for the licensed-database-anchored ownership graph; downgraded to 3 for edge-extension hops). Reconciliation against existing NAME_SIMILARITY matches: when both routes hit the same (designation_id, canonical_target_org_id), the name-match row stays canonical and a closure_coincides_with_name_match flag is set; cascade rows that no longer reproduce on a refresh get deprecated_at = now() with reason CHAIN_EDGE_DROPPED.
hop_is_stub = true. The renderer shows such a hop as a stub identified only by its id, so the id stays available for cross-reference while the generic placeholder label is withheld. Mirrors the structural prefix check in src/lib/data/stub-detection.ts.org-bis-{slug} row; cascading a designation back onto that row returns the designation the walk started from. The closure walker excludes these (isCascadable() in scripts/cascade-bis-designations.ts). The crossref pipeline does the same.removal_date is populated (Federal Register removal action) get all their bis_crossref_matches rows updated with deprecated_at + deprecation_reason = DESIGNATION_REMOVED. The active read path filters on deprecated_at IS NULL; the deprecated rows stay in the table, so what a since-delisted designation once reached is still queryable.is_direct boolean separates “this entity IS the designated entity” from “this entity is in an affiliate closure of a designated entity.” The page renders each with a distinct chip; direct designations have the “Direct” badge and authority tier 1, closure designations show their hop count and the first three relation labels.The cascade firing rate (designations producing at least one AFFILIATE_CLOSURE row reaching a compute-universe org) is 2.94% of designations on record: 101 designations out of 3,435. This is structurally constrained by the upstream licensed-database company-id resolution: 699 of 3,435 designations (20.35%) have an ownership-graph anchor (at least one closure entry in designation_closures) for the BFS to start from, and of those only 101 produce a closure walk that lands on a compute-universe org. The remaining 79.65% of designations on record are silent because that upstream resolution failed: the silence records the resolution, and carries no information about whether the designated entity has compute-relevant affiliates.
ecd-3330): no cross-reference row at all, because its Jaccard name similarity against canonical “HiSilicon” is 0.5 and the offline crossref pipeline only writes a match at 0.7 or above (a lower bar than the 0.9 the read path then requires before calling a match a direct designation). Matching on the discriminative token, the one word in a name that identifies the firm, would reach it.organizations registry at all, so there is nothing for a designation to reach. The gap is in the ingest.Huawei Technologies (added to the Entity List May 16, 2019, FR 84 FR 22961; expanded Aug 2019 in 84 FR 43495; Aug 2020 in 85 FR 51602) reaches HiSilicon (the canonical compute-universe match) via a documented closure walk. The closure rows have per-hop chains, the live crossref row clears the 0.9 gate the direct-designation bucket requires, and the page renders the chain end-to-end with FR citations on each hop. The Huawei→HiSilicon match is the most-fired single cross-reference in the set.
federal_register_citation renders inline with the chip.getSovereignAttributionsForEntity() in src/lib/data/ownership-investigator.ts retrieves sovereign investors’ fund commitments associated with a company. It matches the company’s canonical identifier and its variants against the investment chains, then orders the results by commitment date, newest first.
A sovereign investor can finance a private-equity fund whose portfolio includes the company. That connection identifies a source of capital at fund level; the fund’s investment in the company needs its own amount and terms.
sovereign_investment_chains links the sovereign investor through a fund to a portfolio company. The query selects commitment and fund identifiers, the investor and fund names, commitment date, amount, disclosure flag and status, plus the source URL, source identifier, vintage and authority tier.getAllIdsForOrg() supplies the company’s canonical identifier and known variants for the match on target_canonical_org_id.// getSovereignAttributionsForEntity() in src/lib/data/ownership-investigator.ts
function getSovereignAttributionsForEntity(orgId) {
allIds = getAllIdsForOrg(orgId)
return SELECT id, sovereign_lp_name,
intermediate_fund_name, intermediate_fund_id,
commitment_date, commitment_usd_m, is_dollar_disclosed,
commitment_status, data_source, source_url, data_vintage,
authority_tier
FROM sovereign_investment_chains
WHERE target_canonical_org_id IN allIds
ORDER BY commitment_date DESC
}National AI budgets and project spending are compiled separately in the sovereign AI programmes.
M&A-shaped deals from pe_investments joined to pe_deal_investors for parties, newest first. Backing function: getOwnershipChangeEventsForEntity() in src/lib/data/ownership-investigator.ts. The extraction query retains ten deal types that imply a transfer of ownership — M&A, buyout/LBO, secondary buyout, corporate divestiture, asset sale, reverse merger, joint venture, spin-off, platform creation and public-to-private — and drops seed, Series A–H, growth, PIPE and debt financings before ingest (the extractor’s OWNERSHIP_EVENT_TYPES filter).
This chain reports the deals that moved control of an entity. Funding rounds that raise capital without moving control (Series A, B, C and the rest) are left out, because the question it answers is who controlled whom, and from when.
pe_investments — the M&A deal table. Columns: the licensed-database deal id, deal_date, announced_date, deal_type, deal_subtype, deal_status, deal_size_usd_m, is_majority_control_transfer, data_source, source_url, authority_tier.pe_deal_investors — per-deal counterparties. Columns: pe_investment_id, party_role, canonical_party_org_id, fund_name, fund_id.canonical_org_map.json — resolves canonical_party_org_id to display names; the same map other chains consume.// getOwnershipChangeEventsForEntity() in src/lib/data/ownership-investigator.ts
function getOwnershipChangeEventsForEntity(orgId) {
allIds = getAllIdsForOrg(orgId)
deals = SELECT * FROM pe_investments
WHERE canonical_target_org_id IN allIds
ORDER BY deal_date DESC
if deals.empty: return []
investors = SELECT * FROM pe_deal_investors WHERE pe_investment_id IN deals.ids
partiesByDeal = groupBy(investors, pe_investment_id).map(invs =>
invs.map(inv => {
name = inv.canonical_party_org_id ? getCanonicalName(inv.canonical_party_org_id) : null
party_name = suppressPlaceholderStubName(name) // cid-only placeholder label → null
return { party_role, canonical_party_org_id, party_name, fund_name, fund_id }
})
return deals.map(d => { ...d, parties: partiesByDeal[d.id] ?? [] })
}Where a party resolves to no name, the panel falls back to the fund name, and to “— name not recorded” when there is not one of those either.
data_source label.terminal: 'contested' — the competing claims render side-by-side, each with its own onward chain, and no ultimate-owner conclusion is drawn. Same-document multi-parent structures (founder vehicles, split holdings, one filing naming two holdcos) are co-parents: document identity (edgesAreCoParentClaims in src/lib/data/ownership-investigator.ts) exempts them, and the walk goes on as it normally does: it climbs the strongest of the parents as its main line and attaches each of the others as a branch showing where that one leads.canonical_party_org_id resolves to a stub carrying a placeholder label, the suppressor at src/lib/data/stub-detection.ts returns null and the row falls back through the chain above, so the reader sees the fund name, or nothing. The canonical id stays populated, so deep links still resolve.pe_investments with no rows at all, and the panel names the deal types it looked for.deal_size_usd_m can be null, and the page renders “size undisclosed” in italics.deal_date column records the announce date, and there is no separate close-date column.Microsoft → Activision Blizzard (close date October 13, 2023; SEC 8-K): the pe_investments row has deal_type = Buyout, is_majority_control_transfer = true, deal_size_usd_m = 68700. The chain-of-title walk for Activision Blizzard before the deal terminates at the Activision public-company terminus; after the deal, the walk passes through Microsoft as parent. The page surfaces both the event and the chain transition; the chain rendering reads “ownership transferred 2023-10-13.”
source_url resolves to the SEC EDGAR document.Licensed rows are withheld at the database, where the public role cannot read them. What a subscription source contributes to a page is metadata: how complete a result is, whether a public document exists for a relationship there is reason to think is there, and which records were worth chasing into a registry.
A subscription database is used the way a librarian is used: it says where to look. What gets published is the document at the end of that search (a filing, a registry entry, a company’s own announcement) with its own link, readable without a subscription.
Scrutica serves 71,300 directed ownership edges, each citing a source a reader can open. A further 32,314 are established only in subscription sources and are withheld — so the published graph is 69% of 103,614 edges held in total, as of 2026-09-09.
The facility corpus splits the same way: 4,257 sites published of 4,573 held, with 316 established only in subscription sources. The site-wide facility count is of everything held: the 316 licensed sites are counted in that total but cannot be opened, listed, or drawn on the map.
Located via subscription source means the record beside it is a public primary, cited and linked and readable without a subscription, and that a commercial source is what pointed to it. Of the published ownership edges, 7,530 across 4,854 organisations were re-sourced this way: each was re-established against a public document naming both parties before it was re-tagged, and it is that document the page cites.
Both markers link here from wherever they appear.
The Compute Visibility Index measures, per country, what fraction of in-country compute capacity is attributable to a public-filing-grade ultimate beneficial owner. Four tiers plus an explicit coverage state. Tier 1: UBO identified through SEC / FCA / equivalent securities-regulator filings. Tier 2: UBO through corporate-registry filings (Companies House, Bundesanzeiger, METI EDINET, NRA registries). Tier 3: UBO through licensed analyst-database research (held under subscription, not redistributed) where primary filings are not available. Tier 4: the facility's linked organization is named in Scrutica's records, and the upward ownership walk found no documented ancestors above it. Data Gap is a separate state: the facility has no operator, owner, or hardware-owner record at all, so the walk has no organization to start from.
Tier 4 is a finding about the records: the upward walk ran and reached no documented ancestor. Data Gap is a finding about Scrutica's coverage: nothing was linked, so no walk could run. Tier 4 covers three situations at once: organizations no consolidation regime ever required to name a parent, organizations that could name one and do not, and relationships Scrutica has not yet joined. The Compute Visibility Index page splits that absence into those states, attributes each to the registrant, the registry or Scrutica, and gives the gap a column of its own.
Hosted vs Controlled. A facility physically located in country A whose ultimate beneficial owner is domiciled in country B counts toward country A on the “hosted” view and country B on the “controlled” view. The two diverge where jurisdictional arbitrage applies (e.g., AWS facilities in Ireland controlled from the US; G42-controlled facilities hosted in the UAE). The Allied Coordination Gap Analyzer, which lists the facilities where that split crosses an allied export-control boundary, treats the divergence as the policy-relevant signal, because export-control authority follows the controlling jurisdiction.
The Compute Under Control metric is the jurisdictional-attribution counterpart to the Compute Visibility Index. It answers a strict question: of the compute physically hosted inside country C, what fraction is operated by a C-incorporated firm whose ultimate corporate parent is also C-incorporated? Each facility falls into one of five named classes. Domestic: host country, operator country, and ultimate parent country all equal C. Foreign-concentrated: operator and ultimate parent agree on a non-host jurisdiction (the textbook hyperscaler-region pattern, Microsoft Azure Frankfurt operated by Microsoft Deutschland, ultimately Microsoft US). Bifurcated: operator and ultimate parent jurisdictions disagree (the 21Vianet-Azure-China pattern, a China-incorporated operator under a US-incorporated ultimate parent). Opaque: a jurisdiction the classifier needs is unresolved, either because the chain terminates without a country-attributable node or because the linked organization itself has no documented country. Data gap: the facility has no operator, owner, or hardware-owner record linked, so there is no organization to classify from.
A single-jurisdiction pick hides the bifurcated cases. The Hosted-vs-Controlled tab takes owner_org.country_hq at one hop, which smooths a facility whose operator and ultimate parent sit in different jurisdictions into whichever side that pick landed on. The strict criterion refuses the smoothing: such a facility appears in the drill-down for both countries as a jurisdictional-axis contributor, and in neither's “Under control” column. Bifurcation is one of the five reported classes, with its own count.
Ultimate-parent resolver. The classifier reads from mv_ownership_chain_closure (the transitive closure of the upward-walk view, bounded at six hops; the same view the Compute Visibility Index reads). For each facility's linked organization, the ultimate parent is the deepest reachable ancestor in the closure; ties at the same hop depth are broken on the ancestor's canonical id, which makes the pick deterministic. An organization present in the ownership graph with nothing above it resolves to itself, since the closure includes a hop-0 self row. The resolver does not infer beneficial ownership from indirect signals.
The archived gap-pattern overlay — the earlier version of this breakdown — recorded, for each sovereign program, whether a gap between announced and disbursed spending existed and what percentage of that gap fell to each of six named reasons. It recorded no dollar amount against any reason, and neither the announced nor the disbursed figure those percentages were struck over, so no percentage can be checked against the money it describes. The source links and confidence labels on each attribution show that the reason was cited somewhere; they do not show how much of a gap belongs to it. The per-program percentages, the cross-program matrix built from them, and every total derived from either are therefore withheld.
The archived allocations are held apart from the programme records. The code that reads them checks receipt structure and arithmetic only; whether a cited source says what the percentage beside it says is what the verification has to settle, so the allocations return only after an independent editorial verification of every amount against its source, in comparable units and on comparable dates, followed by a UI republication audit of each figure the page would display.
The withdrawal covers this cross-program breakdown only. Each program record still carries its own inflation_decomposition field, which splits a program by financing type; those records remain on Programs, and the separate capital ledgers remain on Funding.
The Capability Frontier places scores from Epoch, Artificial Analysis and LMArena beside available training-compute estimates and hardware attributions. Each index retains its own scale and recorded vintage. The curation combines selected model variants across the sources, so a difference between scores cannot be read as a comparison under matched inference budgets, tools or model endpoints.
The country groups are Scrutica’s editorial categories, assigned from the trainer organization’s country and open-weights field. Downloadable weights take precedence over country. Among closed-weight models, UAE and Saudi Arabia form one group; China, Russia, Iran and North Korea form another. The “allied” group uses the country set in the curator and also receives unmatched countries. A label records the trainer organization’s country and open-weights status, and nothing about where a model was trained or what licence its hardware needed.
Hardware evidence. The curator uses a joined Epoch hardware field or a supporting developer disclosure. Meta’s Maverick model card reports H100-80GB pretraining and identifies the FP8 release as a quantized variant. Google’s Gemini 2.5 report, §2.3 describes TPUv5p training for the family that includes Pro and Flash; it does not give a chip count for each model. Inferences from a sibling model alone are withheld, and an empty chip class means the snapshot records no chip for that model. A chip name and a release date do not carry the transaction facts an export-control assessment needs.
Date-and-score frontier. An outlined model has no competitor in the selected subset with a later or equal release date and a higher or equal score, with at least one strict advantage. This highlights recent, high-scoring models. The calculation has no training-compute operand and does not reconstruct which model led at its release date.
Available-metric rank. Each model receives a rank within every index where it has a score. Scrutica averages its available ranks, then ranks those averages; ties in the averages break by ECI rank, then model ID. Equal underlying scores receive successive ranks in input order. The calculation uses the full cohort before the country/weights filter; models with no score receive no rank. Because indices cover different populations and some models have only one score, this ordering can reflect missing coverage as well as performance.
Facility estimates reuse Epoch inputs, so agreement with those inputs cannot independently validate them. The checks below concern estimator assumptions, market-share calculations and recorded hardware configurations.
| Validation Source | Data Type | Coverage | Methodology Difference |
|---|---|---|---|
| Epoch AI GPU Clusters | Facility PFLOP/s estimates | Imported cluster records | Epoch supplies some of the facility inputs. Reusing its figures cannot provide an independent check; comparisons must use the same site or cluster scope and distinguish peak from sustained compute. |
| TrendForce / Counterpoint | Semiconductor market share | Quarterly updates | The concentration calculation can be reproduced from the cited shares. This checks the arithmetic; the shares’ market definition and date still govern what the result measures. |
| MLPerf Inference | Hardware specifications (accelerator type, memory, count) | Recorded v4.1 configurations | Submitted accelerator names and memory specifications provide a reference for the chip catalog. Matching must distinguish hardware variants before a difference is treated as an error. |
MLPerf Hardware Cross-CheckRecorded accelerator names: AMD Instinct MI300X, NVIDIA H100, NVIDIA L40S, NVIDIA Jetson AGX Orin 64G, NVIDIA H200, AMD MI300X, NVIDIA L4, NVIDIA B200, NVIDIA A100, NVIDIA GH200, TPU v5e, TPU v6, NVIDIA GH200 Grace Hopper Superchip 144GB, NVIDIA GH200 Grace Hopper Superchip 96GB, UntetherAI speedAI240 Preview, UntetherAI speedAI240 Slim. Names retain the submitters’ spellings, so different entries can refer to the same hardware.
No per-facility agreement score is published. The earlier matching procedure could pair a whole-site estimate with a cluster inside it, so the resulting difference did not measure accuracy. To examine how hardware counts, power and utilisation affect an estimate, adjust the Interactive FLOP estimator.
The bis_license_actions table ingests the BIS Office of Technology Evaluation per-country “Annual Country Licensing and Trade Analysis” PDFs (bis.gov · OTE). 2,587 rows × 33 jurisdictions × calendar years 2018-2022. Statutory authority Section 1765 of ECRA (50 USC §4824 / Pub. L. 115-232). Authority tier 1 (primary government measurement).
The trade series, CY2023–2026. Bilateral semiconductor trade-flow (HS 8542 / 854231 / 854232 / 854239) runs through 2026 on UN Comtrade, CEPII BACI (231 reporters), Taiwan Customs, China GACC, and Japan e-Stat. These measure what was declared under customs classifications; they do not measure what BIS licensed. The trade view and the diversion view show the two side by side and never combine them into one figure. CRS R48642 (Sutter, updated 2025-09-19, crs · R48642) is the leading public synthesis of the CY2023+ export-control posture in the absence of the OTE series; it cites press-release-conveyed CY2020-Q1-2022 license values and carries no CY2023 OTE-grade license data.
The licence-action grid, CY2023–2026. The per-country, per-ECCN, per-year license-action grid (approved / denied / returned-without-action / deemed-export) does not exist publicly. Reading a licensed-versus-banned-versus-grey-market split off trade-flow shifts is unsound at this vintage: trade-flow reflects customs declarations and re-export routing, license-action data reflects BIS decisions, and joining the two would credit routing changes to licensing decisions the data cannot see. That split stays uncomputed until the OTE data lands.
The PJM grid-context indicator appears on the US sovereign program, and it totals the generation capacity in PJM Interconnection's queue archive county by county across PJM's territory (13 Eastern-US states and DC). It appears there because several of the US program's largest US-sited resources sit inside that footprint — the NAIRR Pilot's Purdue Anvil node in Indiana, and CHIPS Act recipient sites in Pennsylvania and Ohio, Intel's New Albany fab among them — and the data-center load planned against that grid is what pre-empts the headroom new procurement needs. Others sit outside it, and the indicator says nothing about them: SDSC's Expanse is in San Diego, and the program's New York CHIPS recipients are on NYISO. The indicator renders only where a program's country code is us, because every other program powers its host facilities from a grid PJM does not reach — CAISO, ERCOT, ENTSO-E, K-EPCO and others.
What is populated: the generator queue. The pjm_interconnection_queue table indexes PJM's multi-decade Generator Interconnection Queue archive (~7,000+ rows; ~70% historically withdrawn before energization, the structural baseline for any interconnection-queue analysis). Generator-queue rows model new generation seeking to interconnect (gas turbines, solar farms, storage, wind); they do not model data-center loads requesting service. Per-facility resolution against a specific data center is structurally infeasible from the generator queue alone, so the indicator aggregates to county-level capacity totals.
The query engine answers questions typed in ordinary English against Scrutica's data and the analyses built on it. A question passes an injection-defence sanitiser, which strips a fixed list of override phrases and Unicode control characters, is wrapped in <user_query> delimiters, and reaches the configured model: Claude Sonnet 5 by default, served through Anthropic's API with prompt caching over the tool schemas and system prompt where that path is provisioned, and through OpenRouter otherwise. The model calls a think tool first to write out its plan, which the page draws as it arrives, and then runs data tools: in parallel where they are independent, in sequence where one tool's output feeds the next. The loop runs for at most five rounds and 8 data-tool calls in total. A call past that budget is not executed; its result slot tells the model to answer from what was gathered and to state which parts of the question it could not cover. The answer then streams back.
The public tool surface is 55 tools, covering the site's records, the analyses derived from them, and its own page content and corrections log; the tool reference lists each with every parameter and accepted value. An answer can draw only on what those tools return under their parameters and result limits.
Citation tags and the identifier check. The system prompt instructs the model to wrap a claim supported by an identifiable record in a <cite tool record_id tier> tag, copying the identifier from the tool result, and, where a result carries no identifier, to attribute the figure to the tool or method rather than invent one. The page parses the tags as the answer streams and renders a citations panel beneath it: for each tag, the tool, the identifier, the claim text and the authority tier, which comes from the record's provenance where a tool returned it and otherwise from the tag itself, labelled as the model's own assertion. Source URL, estimate flag and data vintage render where the provenance carried them. After synthesis completes, a server-side verifier parses the same tags and tests whether each cited record_id occurs as a substring of the run's concatenated tool-result JSON; an identifier it cannot find is flagged unverified provenance in the panel, and the claim text is left in place. The test is a string lookup: it does not read the record, so it establishes neither that the record supports the claim nor that a value was transcribed correctly, and untagged claims are outside it. If the verifier itself fails, the answer is saved with no verdict on record, which a saved run reports as distinct from a pass.
r100_rubin Aligned the display name with NVIDIA’s current product-facing name, Rubin, while preserving the stable internal key. Added NVIDIA’s preliminary per-GPU dense FP16/BF16 figure of 4,000 TFLOP/s; the separately published 50 PFLOP/s NVFP4 figure is sparse and is not substituted. Per-GPU TDP and unit price remain absent. 0 → 4,000 TFLOP/smaia_200 Added Microsoft’s published per-chip 750 W SoC TDP for Maia 200. Dense FP16/BF16 throughput and per-unit price remain absent; FP4/FP8 figures are not converted into the registry’s dense training-throughput field. 0 → 0.75 kWgpu-spec-availability Numeric callers now reject any estimate whose required accelerator fields retain the registry’s zero absence marker. Partially specified models remain available to calculations whose own operands are recorded: Rubin can run throughput-only hardware calculations, while its power and cost outputs remain withheld; Maia 200 remains withheld from throughput-dependent outputs.gb300 Added NVIDIA GB300 (Blackwell Ultra, per GPU, NVL72 configuration). Dense FP16/BF16 = 2,500 TFLOP/s — derived from NVIDIA’s GB300 NVL72 page (360 PFLOP/s with-sparsity per 72-GPU rack → 5,000 sparse / 2,500 dense per GPU under NVIDIA’s 2:4 convention) and listed directly by Epoch AI ml_hardware; 1,400 W per Epoch (NVIDIA publishes no per-GPU TDP). Before this, the Threshold Atlas had no GB300 spec and silently fell back to the H100 SXM5 assumption for GB300 sites — 743,001 GPUs across 6 facility inventories at the 2026-07-22 substrate probe, including Stargate UAE Phase 1 and HUMAIN Phase 1 — understating their per-second dense throughput by ~2.53× (2,500 vs 989.5).v100_sxm2,tpu_v4,tpu_v5e,tpu_v5p,ascend_910b Added five accelerators that are the dominant on-record hardware at live facilities but had no spec entry, so every one of their sites was silently H100-SXM5-assumed — in these cases OVERSTATING capacity: V100 SXM2 (125 dense tensor FP16, NVIDIA Volta datasheet; ~54 facilities incl. legacy supercomputers, overstated 7.9×), TPU v4 (275 BF16, Google docs; overstated 3.6×), TPU v5e (197 BF16; overstated 5.0×), TPU v5p (459 BF16; overstated 2.2×), Huawei Ascend 910B (375 dense FP16 per CSET "Pushing the Limits" 910B2 row, Epoch lists 376 — divergence surfaced; overstated 2.6×, a bad direction for an export-control-focused surface to err on a Chinese site). TPU TDPs are Epoch-curated where Google publishes none; TPU v4 power uses Google’s measured max (192 W) with Epoch’s 340 W divergence noted.mi325x,tpu_v6e Added AMD MI325X (CDNA 3: dense FP16/BF16 1,307.4 TFLOP/s per the AMD data sheet — compute-identical to MI300X, memory-upgraded to 256 GB HBM3E at 6 TB/s; 1,000 W) and Google TPU v6e Trillium (918 BF16 per Google Cloud docs; 380 W per Epoch), closing the 2025-generation coverage gap between MI300X and MI350X and between TPU v5 and v7. All values cross-checked against the live Epoch ml_hardware release of 2026-07-21.threshold-atlas-hardware-mapping The Threshold Atlas gpu_inventory → spec-key mapping now recognizes GB300, V100, TPU v4/v5e/v5p/v6e, MI325X, and Ascend 910B inventory strings (previously all fell to the H100 SXM5 fallback). Parts deliberately left unmapped, with the fallback intact, because no primary or research-database spec could be pinned: Trainium1 (AWS Neuron docs read as 190 TFLOP/s per NeuronCore-v2 × 2 cores/chip while Epoch lists 190 per chip — an unresolved 2× ambiguity), Gaudi3 (Intel materials 1,835 vs Epoch 1,678 BF16, unresolved), Intel Max 1550 (no tensor-FP16 figure in Epoch; Intel materials do not disambiguate), MI300A (980.6 dense ≈ 0.9% from the H100 fallback — immaterial), and pre-tensor-core / non-GPU-comparable architectures (Kepler/Pascal Teslas, Cerebras WSE, Groq LPU, PEZY, Tesla Dojo D1, Maia 100). A gap statement beats a guessed spec.trainium2 Added Amazon Trainium2. Dense BF16 = 667 TFLOP/s, ~500W TDP, 96GB HBM3e; cloud-only via EC2 trn2 (Anthropic Project Rainier). Verified against AWS Neuron Trainium2 architecture docs and the SemiAnalysis Trainium2 teardown; matches the value the Epoch GPU-cluster ingest already used (amazon trainium2 = 667). Before this, the Threshold Atlas had no Trainium2 spec and silently fell back to the H100 SXM5 assumption for Trainium2 sites, overstating their per-second dense FP16 by ~48% (989.5 vs 667).threshold-atlas-facility-scope The Threshold Atlas per-country capacity view now (a) excludes chip-manufacturing facility types (logic_fab, memory_fab, packaging) from the "sites that hold the compute to train a model" universe — a memory fab’s power draw is not model-training compute — and (b) reads each facility’s own dominant hardware_type from its gpu_inventory before falling back to the H100 assumption. Prior behavior counted powered fabs as threshold-capable and assumed H100 for every site regardless of the accelerator the platform already had on record.country-aggregate-mvs Country-level aggregate views (country stats, Compute Visibility rollup, Compute Under Control rollup) now apply the editorial numerical-corrections overlay that per-facility surfaces have applied since 2026-06-20 — a correction row supersedes the substrate value, including corrected-to-withheld (null) values for superseded build-phase capacity. Before this, database-side country totals silently disagreed with page-side totals by the corrected amounts (double-counting superseded Builds-Upon phases and keeping the uncorrected Stargate-UAE chip estimate). No parameter values changed; the same correction rows now bind everywhere.affiliate-rule-inferred-designations Removed four affiliate-rule-inferred export-control designations (undisclosed-entity surrogates; designations count 3,426 → 3,422). Two independent grounds, both verified: the legal basis for the inference — the BIS 50%-affiliates rule — was suspended effective 2025-11-10, eighteen days before the rows’ effective date, so the inferred designations were never operative; and the ownership evidence underlying the inference is no longer reproducible from the current substrate (verified 2026-07-11). Explicitly-named Entity List designations are unaffected. Version history is append-only: a full snapshot of the removed rows is retained for auditability, and the suspension status (stayed to 2026-11-09) remains tracked — if the rule takes effect, affiliate coverage will be re-derived from then-current evidence.methodology-overview-copy Corrects the freshness-TTL claim in v1.3.2, which drifted out of truth 10 days after publication and stood uncorrected for 53 days. v1.3.2 (2026-05-08) verified the then-live FRESHNESS_TTLS ranges of 1–6h / 6–24h / 24–72h / 24–168h (introduced by commit 306fee58, 2026-04-08); commit bf41b310 (2026-05-18) then recalibrated all four classes upward — structural data was surfacing as "aging" after a single day, a false-alarm staleness signal across most of the platform — and no correcting entry landed until this one (2026-07-10 audit). Current values, interpolated directly from FRESHNESS_TTLS in src/lib/freshness.ts at build so this entry cannot itself drift (the same binding discipline this registry's header mandates for parameter values): regulatory fresh <168h / stale >720h; market fresh <168h / stale >720h; structural fresh <2160h / stale >8760h; narrative fresh <4320h / stale >8760h. Version history is append-only: v1.3.2's text stands as published, with this entry as the correction of record. No parameter values changed.methodology-overview-copy Methodology overview copy calibrated against live Supabase substrate (probe 2026-05-08T22:46Z, paginated to avoid PostgREST’s default row cap). Three quantitative claims corrected for substrate drift: (a) facility country coverage 128 → 110 (the 110 distinct non-null countries that contain facilities; 157 facilities have NULL country and are not counted as a country); (b) export-control designation count 3,349 → 3,354 (substrate query: 3,354 total rows in `export_control_designations`, all with `removal_date IS NULL`, i.e. all active; the prior 3,349 reflected an older snapshot before five BIS-Entity-List additions landed); (c) cron-route count Ten → Fourteen (live `vercel.json` declares 14 entries; the Inngest function-definition file is retained for reference but is not the live scheduler, and pg_cron jobs are retention-only, not table-feeders). Five claims verified-and-kept-as-is: (~)4,500 facilities (substrate 4,529, ~ rounds correctly), (~)18,600 supply-chain edges (substrate 18,579, ~ rounds correctly), 34 sovereign AI programs (exact match), 13 GPU/accelerator generations + 2 placeholders (matches `GPU_MODELS` constant in this file: 13 entries with non-zero `fp16_tflops` plus R100 Rubin and Maia 200 placeholders), TTL ranges 1–6h / 6–24h / 24–72h / 24–168h (matches `FRESHNESS_TTLS` in `src/lib/freshness.ts` exactly). No parameter values changed; this is a copy-only calibration.mfu_assumption Description corrected after methodology validation. Prior copy attributed an MFU range of 0.46–0.57 to Chinchilla; Hoffmann et al. 2022 (arxiv 2203.15556) does not directly publish MFU for the 70B run. The 0.46–0.57 range is from PaLM (Chowdhery et al, arxiv 2204.02311, Table 3), and conflated MFU 0.462 with HFU 0.578 — these are distinct metrics (MFU strict vs hardware FLOP utilization including rematerialization). Updated description anchors the 0.40 default explicitly above Epoch AI 2022's 0.30 LLM recommendation, with PaLM 0.462 MFU and LLaMA 3 0.384 MFU as primary calibration points. No numeric default changed.interconnect_efficiency,mfu_assumption Source URL canonicalized from https://epochai.org/... to https://epoch.ai/blog/estimating-training-compute (Epoch AI publishes under epoch.ai; the older epochai.org redirects). Source attribution made specific: "Epoch AI (Sevilla et al, 2022)" replaces the bare "Epoch AI" string for interconnect_efficiency.gpu_fraction_of_it Default revised 0.70 → 0.49 to match SemiAnalysis's 100,000-H100 cluster power decomposition (the only published primary-adjacent component-level breakdown of a frontier AI training cluster). The decomposition: 70 MW GPU compute / 142.5 MW total IT load = 49% GPU share. The previous 0.70 default treated GPU TDP as covering server-level draw, which it does not — in-server CPUs, NICs, and PSUs draw an additional ~575 W per GPU; external storage and networking add another ~10% of total IT. Applied alongside the unchanged 700 W H100 TDP, the corrected default reduces predicted GPU count from 100 MW gross by ~27% (from ~83K to ~61K H100), which closes the discrepancy against the SemiAnalysis ground truth and aligns the Capacity Estimator with measured frontier facility decompositions. 0.7 → 0.49pue Default revised 1.20 → 1.15 to match empirical fleet disclosures from operators running frontier AI training clusters. Anchors: Microsoft FY25 fleet PUE 1.16 (Americas; 1.28 APAC), AWS 2024 fleet 1.15, xAI Colossus Memphis 1.18 (current; 1.10 planned for liquid-cooled expansion), Oracle Stargate Abilene 1.12 (with heat recovery). Outliers: Google fleet 1.09 (Central Ohio 1.04), Meta FY23 1.08 — both below the new default. No operator publishes AI-training-pod-specific PUE separately from fleet averages; the 1.15 default reflects the median of fleet-level disclosures with documented data-center cooling architecture. The max bound tightened from 1.60 (which captured legacy enterprise) to 1.45 (which captures hot-climate AI facilities). 1.2 → 1.15eu-ai-act-measurement-tolerance EU AI Act §51(2) compliance methodology now reflects the July 2025 GPAI Guidelines (European Commission, 18 July 2025). The guidelines establish a ±30% measurement tolerance on cumulative training FLOP, accept both hardware-based and architecture-based estimation paths, and exclude RLHF reward-model training from the cumulative count while including continued pre-training and supervised fine-tuning. The 10²³ FLOP floor introduced in the same guidelines defines the GPAI scope itself — below this, no GPAI obligations attach. Article 51(3) delegated-act authority remains unexercised as of April 2026.cascade-edge-weighting Cascade propagation now blends editorial criticality with a 3-month stock price correlation (Pearson r) when market correlation weighting is enabled. Edges without correlation data (the second licensed database, unlisted companies) keep editorial scores unchanged; sole-source edges retain criticality ≥ 9 regardless of correlation. Default blend weight w = 0.5.h100_sxm5,a100_sxm,b200 Canonicalized BF16 throughput on DENSE (no 2:4 structured sparsity) across the FLOP engine and Cost Index: 989.5 TFLOP/s per H100 SXM, 312 per A100, 2,250 per B200. Pre-Pass-46 the FLOP engine quoted dense and the Cost Index computed against sparsity-inclusive specs; both defensible from NVIDIA’s datasheet, but the two features disagreed on units. Dense matches Epoch AI ml_hardware (the cited source) and the throughput real training workloads achieve. Cost-per-petaFLOP-day values landed roughly 2x higher; provider ordering invariant.substitutability-decay-hbm HBM Memory decay description refreshed for HBM4 mass production (Feb 2026 at SK Hynix M16/M15X and Samsung Pyeongtaek) and the Q3 2025 share split SK Hynix ≈57%, Samsung ≈22%, Micron ≈21% (Counterpoint Research Q3 2025 memory market tracking, as reported through Semiecosystem and corroborated by Counterpoint’s own Q3 2025 press summaries). Samsung’s share recovered from ≈15% in Q2 2025 on the HBM3E NVIDIA qualification cycle. Concentrated but multi-source.mlperf-cross-check External validation now incorporates MLPerf Inference v4.1: 53 unique datacenter system configurations covering 16 accelerator families (NVIDIA A100/H100/H200/B200/GH200/L4/L40S, AMD MI300X, Google TPU v5e/v6, plus accelerator-startup submissions). Provides independent verification of accelerator type, count, and memory specifications against the Scrutica chip catalog.gb200 Added NVIDIA GB200 (Blackwell, per-GPU). Dense FP16 = 2,500 TFLOP/s, 1.2 kW TDP. Verified against Epoch AI ml_hardware.csv.mi350x Added AMD MI350X (CDNA 4). Dense BF16 = 2,309.6 TFLOP/s, 1.0 kW TDP. Shipped Q3 2025. Verified against Epoch AI ml_hardware.csv.trainium3 Added Amazon Trainium3. Dense BF16 = 671 TFLOP/s, FP8 = 2,517 TFLOP/s. Cloud-only. Verified against Epoch AI ml_hardware.csv.tpu_v7_ironwood Added Google TPU v7 Ironwood. BF16 = 2,307 TFLOP/s. Cloud-only. Verified against Epoch AI ml_hardware.csv.r100_rubin Added NVIDIA R100 (Rubin) placeholder. Announced GTC 2026, volume H2 2026. FP16 specs not yet disclosed; set to 0.maia_200 Added Microsoft Maia 200 placeholder. TSMC 3nm, 216GB HBM3E. Perf specs not yet disclosed; set to 0.b200 Corrected B200 fp16_tflops from 1,125 (was TF32) to 2,250 (dense FP16/BF16 Tensor). Verified against Epoch AI ml_hardware.csv and NVIDIA datasheet. 1,125 → 2,250 TFLOP/sScrutica. “Methodology.” v1.3.7, September 2026. https://scrutica.com/methodology
For a specific estimate, include the methodology version and the estimation path (for example “Hardware Path, v1.3.7”) so the calculation reproduces against the same parameter set. Published under CC BY-SA 4.0 except where source-attributed otherwise.
Cite the Compute Cost Index dataset (v2026.3, released 2026-07-17): scrutica.com/datasets/compute-cost-index (No DOI assigned; cite the version and checksum on the dataset page).
Cite the Sovereign AI Program Index dataset (v2026.3, released 2026-07-17): scrutica.com/datasets/sovereign-ai-index (No DOI assigned; cite the version and checksum on the dataset page).