Economics
Cloud rate cards and training-run costs; hyperscaler capital expenditure.
Cloud rate cards and training-run costs; hyperscaler capital expenditure.
Vintage 38 disclosure records across 14 labs · 9 models with records from two or more sourcescloud prices last verified 2026-03-29 – 2026-09-09 (oldest and newest rows; three providers refresh daily, three are manual snapshots)
Training compute · FLOP, logarithmic scale shared by every row
Dashed vertical line: 10²⁵ FLOP · EU AI Act systemic-risk trigger
The lab published the figure
Llama 3.1-405B
Meta Platforms, Inc. · Jul 2024
Lab published 3.8×10²⁵ FLOP. Third-party estimate: 3.8×10²⁵ FLOP.
Inflection-2
Inflection AI · Nov 2023
Lab published 1×10²⁵ FLOP. Third-party estimate: 1×10²⁵ FLOP.
The lab published the inputs only
Llama 4 Behemoth (preview)
Meta Platforms, Inc. · Apr 2025
No lab-published compute figure. Third-party estimate: 5.2×10²⁵ FLOP.
Inferred from disclosed inputs: 3.0×10²⁵–7.0×10²⁵ FLOP (tier 4).
Nemotron-4 340B
Nvidia · Jun 2024
No lab-published compute figure. Third-party estimate: 1.8×10²⁵ FLOP.
Inferred from disclosed inputs: 1.5×10²⁵–2.0×10²⁵ FLOP (tier 4).
Pangu Ultra
Huawei Technologies · Apr 2025
No lab-published compute figure. Third-party estimate: 1.1×10²⁵ FLOP.
Inferred from disclosed inputs: 1.0×10²⁵–1.3×10²⁵ FLOP (tier 4).
DeepSeek-V3
DeepSeek · Dec 2024
No lab-published compute figure.
The lab named only the hardware or the size
Gemini 1.0 Ultra
Google DeepMind · Dec 2023
No lab-published compute figure. Third-party estimate: 5×10²⁵ FLOP.
Scrutica-assigned lower bound: 1×10²⁵ FLOP (tier 4).
No quantitative lab disclosure recorded
Grok 4
xAI · Jul 2025
No lab-published compute figure. Third-party estimate: 5×10²⁶ FLOP.
GPT-4.5
OpenAI · Feb 2025
No lab-published compute figure. Third-party estimate: 3.8×10²⁶ FLOP.
Grok 3
xAI · Feb 2025
No lab-published compute figure. Third-party estimate: 3.5×10²⁶ FLOP.
Grok-2
xAI · Aug 2024
No lab-published compute figure. Third-party estimate: 3.0×10²⁵ FLOP.
Claude 3.5 Sonnet
Anthropic · Jun 2024
No lab-published compute figure. Third-party estimate: 2.7×10²⁵ FLOP.
Scrutica-assigned lower bound: 1×10²⁵ FLOP (tier 4).
GPT-4 (Jun 2023)
OpenAI · Jun 2023
No lab-published compute figure. Third-party estimate: 2.1×10²⁵ FLOP.
The lab documented that it would not say
GPT-4 (Mar 2023)
OpenAI · Mar 2023
No lab-published compute figure. Third-party estimate: 2.1×10²⁵ FLOP.
Scrutica-assigned lower bound: 1×10²⁵ FLOP (tier 4).
Each track represents one recorded frontier run. Runs are grouped by the lab’s disclosure: a compute figure, inputs from which compute can be estimated, partial information, no quantitative disclosure, or an explicit decision to withhold it.
The horizontal axis shows training compute in FLOP on a logarithmic scale. Filled circles show lab-published figures, open circles show third-party estimates, and brackets show ranges inferred from disclosed inputs. Scrutica grades a source on four tiers, tier 1 a primary filing or measurement and tier 4 an estimate; the brackets are tier 4.
An unmarked track has no recorded FLOP figure to plot. A disclosure of GPU-hours or spending alone can still determine its category without supplying a position on the compute axis.
The dashed line at 10²⁵ FLOP is the EU AI Act systemic-risk trigger. Some runs enter this dataset on a qualitative description of their capability. Scrutica assigns those records a 10²⁵ FLOP floor (tier 4). The assigned floor is omitted from the plotted compute figures.
The most recent run whose lab published a compute figure was released in July 2024; the 7 recorded runs released since fall in the “inputs only” and “no lab quantity recorded” rungs. Both counts are over the 14 runs on this ladder.
Compute estimates may use disclosed chip counts, hardware specifications or model size. A qualitative description such as “frontier-class” supplies no compute quantity. Of twelve training costs on record, exactly one came from the lab that trained the model.
Choose a compute budget, GPU model and procurement method to estimate the run’s cost. Cloud estimates use published rates from the Cost Index. The on-premises estimate includes three-year hardware amortisation, power, networking and facility costs. The default budget is Llama 3.1 405B’s disclosed 3.8×10²⁵ FLOP, and the EU AI Act’s 10²⁵ FLOP trigger is one of the preset budgets. How the estimate is derived
Of fourteen recorded frontier runs, two reach the 10²⁵ FLOP systemic-risk trigger on a figure the lab itself published, and eleven only on a third party’s estimate. One has no compute figure this axis can place, so it is counted in none of those classes. One of these runs has a recorded value of exactly 10²⁵ FLOP. Each run is counted once, against the best-authority figure it has.
| Run | What the lab published | Lab’s figure | Estimated | Cost on record |
|---|---|---|---|---|
| Llama 3.1-405B · Meta Platforms, Inc. | The lab published the figure (lab source) | 3.8×10²⁵ | 3.8×10²⁵ | $53M · estimated |
| Inflection-2 · Inflection AI | The lab published the figure (lab source) Source unavailable when checked 2026-09-02 (HTTP 404). Flagged citation | 1×10²⁵ | 1×10²⁵ | $13M · estimated |
| Llama 4 Behemoth (preview) · Meta Platforms, Inc. | The lab published the inputs only (lab source) | not published | 5.2×10²⁵ | $45M · estimated |
| Nemotron-4 340B · Nvidia | The lab published the inputs only (lab source) | not published | 1.8×10²⁵ | $21M · estimated |
| Pangu Ultra · Huawei Technologies | The lab published the inputs only (lab source) | not published | 1.1×10²⁵ | none |
| DeepSeek-V3 · DeepSeek | The lab published the inputs only (lab source) | not published | none | $5.6M · lab-stated |
| Gemini 1.0 Ultra · Google DeepMind | The lab named only the hardware or the size (lab source) | not published | 5×10²⁵ | $31M · estimated |
| Grok 4 · xAI | No quantitative lab disclosure recorded | not published | 5×10²⁶ | $388M · estimated |
| GPT-4.5 · OpenAI | No quantitative lab disclosure recorded | not published | 3.8×10²⁶ | $366M · estimated |
| Grok 3 · xAI | No quantitative lab disclosure recorded | not published | 3.5×10²⁶ | $218M · estimated |
| Grok-2 · xAI | No quantitative lab disclosure recorded | not published | 3.0×10²⁵ | $32M · estimated |
| Claude 3.5 Sonnet · Anthropic | No quantitative lab disclosure recorded (lab source) | not published | 2.7×10²⁵ | $26M · estimated |
| GPT-4 (Jun 2023) · OpenAI | No quantitative lab disclosure recorded | not published | 2.1×10²⁵ | none |
| GPT-4 (Mar 2023) · OpenAI | The lab documented that it would not say (lab source) | not published | 2.1×10²⁵ | $37M · estimated |
The table compares calculated costs with the published costs held for each run. It uses the recorded compute figure and a hardware and procurement configuration assigned to that model; rows with missing inputs explain why no comparison is available. When the published cost is itself an estimate, agreement can reflect shared assumptions. The Source column identifies the figure used for comparison.
Δ is the calculated cost minus the reported cost, as a percentage of the reported cost. The table counts differences within ±20% as agreement. When several sources supply a cost, the figure from the most authoritative source is used; the Source column identifies it.
5 of 11 reproductions fall inside ±20%; the 6 outside are Gemini 1.0 Ultra, Grok 3, Grok 4, Llama 3.1-405B, Llama 4 Behemoth (preview), Inflection-2, each with the reproduction configuration stated below. A further 3 of 14 runs are listed in the table without a reproduction; the Δ column states the reason.
| Model | FLOP | Hardware | Reported $ | Source | Scrutica calc | Δ |
|---|---|---|---|---|---|---|
DeepSeek-V3 DeepSeek · Dec 2024 | not published | H800 SXM5 × 2,048 | $5.6M | Lab paper (T1) | not run | no central FLOP on record; the lab published GPU-hours instead |
Grok 4 xAI · Jul 2025 | 5×10²⁶ T2 | H100 SXM5 × 200,000 (Colossus full) | $387.8M | Epoch AI (T2) | $527.2M $420.4M – $1.1B | +36% |
GPT-4.5 OpenAI · Feb 2025 | 3.8×10²⁶ T2 | undisclosed | $366.0M | Epoch AI (T2) | $400.7M $319.5M – $809.1M | +9% |
Grok 3 xAI · Feb 2025 | 3.5×10²⁶ T2 | H100 SXM5 × 80,000 (Colossus subset) | $217.8M | Epoch AI (T2) | $369.0M $294.3M – $745.3M | +69% |
Llama 4 Behemoth (preview) Meta Platforms, Inc. · Apr 2025 | 5.2×10²⁵ T2 | H100 SXM5 × 32,000 (FP8) | $44.6M | Epoch AI (T2) + 1 paired source | $54.7M $43.6M – $110.4M | +23% |
Gemini 1.0 Ultra Google DeepMind · Dec 2023 | 5×10²⁵ T2 | TPUv4 × ~57,000 (Epoch derivation) | $30.7M | Epoch AI (T2) + 1 paired source | $52.7M $42.0M – $106.5M | +72% |
Llama 3.1-405B Meta Platforms, Inc. · Jul 2024 | 3.8×10²⁵ T1 | H100 SXM5 80GB × 16,384 | $52.9M | Epoch AI (T2) + 1 paired source | $40.1M $31.9M – $80.9M | -24% |
Grok-2 xAI · Aug 2024 | 3.0×10²⁵ T2 | H100 SXM5 (count undisclosed) | $31.6M | Epoch AI (T2) | $31.2M $24.9M – $63.0M | -1% |
Claude 3.5 Sonnet Anthropic · Jun 2024 | 2.7×10²⁵ T2 | undisclosed | $25.9M | Epoch AI (T2) + 1 paired source | $28.5M $22.7M – $57.5M | +10% |
GPT-4 (Mar 2023) OpenAI · Mar 2023 | 2.1×10²⁵ T2 | A100 SXM4 × ~25,000 | $37.3M | Epoch AI (T2) + 1 paired source | $37.1M $29.6M – $74.9M | -1% |
Nemotron-4 340B Nvidia · Jun 2024 | 1.8×10²⁵ T2 | H100 80GB SXM5 × 6,144 | $21.3M | Epoch AI (T2) + 1 paired source | $19.0M $15.1M – $38.3M | -11% |
Inflection-2 Inflection AI · Nov 2023 | 1×10²⁵ T1 | H100 × 5,000 | $13.5M | Epoch AI (T2) + 1 paired source Lab disclosure Source unavailable when checked 2026-09-02 (HTTP 404). Flagged citation | $10.5M $8.4M – $21.3M | -22% |
GPT-4 (Jun 2023) OpenAI · Jun 2023 | 2.1×10²⁵ T2 | — | — | — | not run | no published cost to compare against; no reproduction configuration recorded for this run |
Pangu Ultra Huawei Technologies · Apr 2025 | 1.1×10²⁵ T2 | — | — | — + 1 paired source | not run | no published cost to compare against; no reproduction configuration recorded for this run |
Every reproduction below runs at MFU 0.40 and interconnect 0.85.
All panels use a 10²⁵ FLOP budget. The utilization and hardware panels use one-year reserved rates; the procurement panel compares rental terms with three-year on-premises amortisation. Hardware and procurement comparisons select the cheapest matching quote for each row, so the provider and region can change too.
The GPU and one-year reserved rate stay fixed while utilization changes. The baseline MFU is 0.40.
The 3 priced rows use quotes from 2 providers. Cost differences therefore reflect both hardware throughput and the available rates. H100 SXM5 and H200 SXM use the same recorded dense throughput of 989.5 TFLOP/s. At this fixed budget and utilization, the cheaper quote produces a 26% lower cost than the dearer one.
Two series. Cloud rows use the cheapest quote for each term, from 3 providers. The lighter bars use Azure throughout: among providers quoting all four cloud terms, it has the lowest on-demand rate.
The historical ledger retains twenty training-cost records. Nineteen also have an ECI score for the scatter and capability-threshold trajectories; Trinity Large remains in the ledger with its unsupported score withheld. The cost-only trajectory uses all twenty records. Each source retains its stated cost basis.
Cloud results use the cheapest recorded quote matching the selected GPU, procurement method and region. The on-premises option uses hardware and operating-cost assumptions. Neither calculation establishes what a particular lab paid; a comparison with a lab’s reported cost also depends on which expenses that report includes.
A total-development estimate would also need experimentation and staff costs. Cottier et al. estimate those costs for four models, with compute and staff modelled separately; their paper does not establish a universal multiplier for the runs here.
The cloud range varies model FLOP utilization and interconnect efficiency within Scrutica’s configured limits. The on-premises range varies utilization and power-usage effectiveness. These ranges show how the assumptions affect cost; they have no assigned probability of containing a run’s actual cost.
Coverage is the Epoch ≥10²⁵ frontier set plus verified lab-primary disclosures: 38 records over 29 models from 14 labs. Coverage depends on the source records collected here. The calculator uses the quoted price vintage; historical source figures retain their stated dollar basis, with no inflation or currency conversion applied.