Economics
Cloud rate cards and training-run costs; hyperscaler capital expenditure.
Cloud rate cards and training-run costs; hyperscaler capital expenditure.
Vintage38 disclosure records across 14 labs · 9 multi-source pairingsprice layer verified 2026-03-29 – 2026-09-03 (oldest and newest rows; three providers refresh daily, three are manual snapshots)
The disclosure ladder. One track per cross-validated frontier run, ordered by how much of the arithmetic its lab left for someone else to do: from a FLOP figure stated in the lab’s own paper down to a technical report that documents its refusal to state one.
Position encodes training compute on a log axis. The glyph encodes who is speaking, the only categorical claim the figure makes: a filled mark is a figure the lab published, an open ring is a third party’s estimate, a bracket is an inference from inputs the lab did publish and is authority tier 4.
A track with no mark is a run for which nobody has published a figure this axis can place: the axis is training compute in FLOP, so a lab that published GPU-hours or a dollar figure and no FLOP has an empty track and a rung that says so. An empty track is still drawn.
The 10²⁵ upright is reference apparatus. It marks the EU AI Act systemic-risk trigger because several runs land near it. A row whose only lower bound equals that figure shows the corpus’s admission floor for a run its lab calls frontier-class without publishing a number; the floor is Scrutica’s own, and the plate does not draw it as a lab mark.
Release dates make the ladder a chronology. The most recent run whose lab published a compute figure was released in July 2024; the 7 cross-validated runs released since fall on the inputs only and qualitative rungs, and none states a figure. The count is over the 14 runs on this ladder only, and it measures what labs published; what was trained is a different question.
Of fourteen cross-validated frontier runs, two reach the 1025 systemic-risk trigger on a figure the lab itself published, and eleven only on a third party’s estimate. No run crosses on this corpus’s admission floor alone; every row whose only lower bound is the floor also has a third-party estimate, and is counted there. One has no compute figure this axis can place, so it is counted in none of the three. One of the crossing runs sits exactly at the trigger, on a figure rounded to one significant digit, which is why the sentence above says "reach". Each run is counted once, against the best-authority figure it has; the threshold comes from the Act, and the classification comes from this corpus.
| Run | What the lab published | Lab’s figure | Estimated | Cost on record |
|---|---|---|---|---|
| Llama 3.1-405B · Meta Platforms, Inc. | The lab published the figure source | 3.8×10²⁵ | 3.8×10²⁵ | $53M · estimated |
| Llama 4 Behemoth (preview) · Meta Platforms, Inc. | The lab published the inputs only source | not published | 5.2×10²⁵ | $45M · estimated |
| Nemotron-4 340B · Nvidia | The lab published the inputs only source | not published | 1.8×10²⁵ | $21M · estimated |
| Pangu Ultra · Huawei Technologies | The lab published the inputs only source | not published | 1.1×10²⁵ | none |
| DeepSeek-V3 · DeepSeek | The lab published the inputs only source | not published | none | $5.6M · lab-stated |
| Gemini 1.0 Ultra · Google DeepMind | The lab named only the hardware or the size source | not published | 5×10²⁵ | $31M · estimated |
| Inflection-2 · Inflection AI | The lab named only the hardware or the size source | 1×10²⁵ | 1×10²⁵ | $13M · estimated |
| Grok 4 · xAI | The lab said only that the model is frontier-class | not published | 5×10²⁶ | $388M · estimated |
| GPT-4.5 · OpenAI | The lab said only that the model is frontier-class | not published | 3.8×10²⁶ | $366M · estimated |
| Grok 3 · xAI | The lab said only that the model is frontier-class | not published | 3.5×10²⁶ | $218M · estimated |
| Grok-2 · xAI | The lab said only that the model is frontier-class | not published | 3.0×10²⁵ | $32M · estimated |
| Claude 3.5 Sonnet · Anthropic | The lab said only that the model is frontier-class source | not published | 2.7×10²⁵ | $26M · estimated |
| GPT-4 (Jun 2023) · OpenAI | The lab said only that the model is frontier-class | not published | 2.1×10²⁵ | none |
| GPT-4 (Mar 2023) · OpenAI | The lab documented that it would not say source | not published | 2.1×10²⁵ | $37M · estimated |
Most public figures for the cost of a frontier training run are built on a training-compute number the lab never published. Of the fourteen runs above, two have a compute figure the lab itself published; for eleven the only number on record was produced by somebody else, working from whatever the lab did disclose: a GPU count, a hardware class, a parameter total, sometimes a sentence calling the model frontier-class. Of twelve training costs on record, exactly one came from the lab that trained the model. The estimates themselves are careful, and they are the reason any of this can be discussed. But a cost figure of this kind is an inference over an undisclosed quantity, priced at rates the buyer did not pay. The calculator below is one more such inference, built so every assumption is visible and adjustable.
The calculator prices three dimensions: how much compute the run needs, what hardware it runs on, and how the hardware is procured. It resolves them against the live price layer the Cost Index maintains, and against a three-year on-prem amortization in parallel; both are stated in the same unit. The default budget is Llama 3.1 405B’s disclosed 3.8×10²⁵ FLOP, and the EU AI Act’s 10²⁵ FLOP trigger is one of the preset budgets. How the estimate is derived
The same calculation, run against each disclosed model at its documented hardware, with the per-row gap to the figure on record shown. Where that figure is itself an estimate, which the ladder above shows is nearly everywhere, the comparison is between two inferences, and the table says which is which.
Each row reproduces a published cost figure inside the calculator above; Δ is the reproduction-versus-disclosed gap, with the ±20% band marking the coarse external-validity threshold. The table reads from frontier_model_runs (multi-source per cell, best-tier-available picked per column); where both Epoch and the lab disclose, the higher-authority figure leads with the source labelled. 5 of 11 reproductions fall inside ±20%; the 6 outside are Gemini 1.0 Ultra, Grok 3, Grok 4, Llama 3.1-405B, Llama 4 Behemoth (preview), Inflection-2, each with the assumption that moves it named below. Gaps of that size are the expected scale of an inference over quantities the labs did not disclose. A further 3 of 14 runs are listed in the table without a reproduction; the Δ column states the reason.
| Model | FLOP | Hardware | Reported $ | Source | Scrutica calc | Δ |
|---|---|---|---|---|---|---|
DeepSeek-V3 DeepSeek · Dec 2024 | not published | H800 SXM5 × 2,048 | $5.6M | Lab paper (T1) | not run | no central FLOP on record; the lab published GPU-hours instead |
Grok 4 xAI · Jul 2025 | 5×10²⁶ T2 | H100 SXM5 × 200,000 (Colossus full) | $387.8M | Epoch AI (T2) | $527.2M $420.4M–$1.1B | +36% |
GPT-4.5 OpenAI · Feb 2025 | 3.8×10²⁶ T2 | undisclosed | $366.0M | Epoch AI (T2) | $400.7M $319.5M–$809.1M | +9% |
Grok 3 xAI · Feb 2025 | 3.5×10²⁶ T2 | H100 SXM5 × 80,000 (Colossus subset) | $217.8M | Epoch AI (T2) | $369.0M $294.3M–$745.3M | +69% |
Llama 4 Behemoth (preview) Meta Platforms, Inc. · Apr 2025 | 5.2×10²⁵ T2 | H100 SXM5 × 32,000 (FP8) | $44.6M | Epoch AI (T2) + 1 paired source | $54.7M $43.6M–$110.4M | +23% |
Gemini 1.0 Ultra Google DeepMind · Dec 2023 | 5×10²⁵ T2 | TPUv4 × ~57,000 (Epoch derivation) | $30.7M | Epoch AI (T2) + 1 paired source | $52.7M $42.0M–$106.5M | +72% |
Llama 3.1-405B Meta Platforms, Inc. · Jul 2024 | 3.8×10²⁵ T1 | H100 SXM5 80GB × 16,384 | $52.9M | Epoch AI (T2) + 1 paired source | $40.1M $31.9M–$80.9M | -24% |
Grok-2 xAI · Aug 2024 | 3.0×10²⁵ T2 | H100 SXM5 (count undisclosed) | $31.6M | Epoch AI (T2) | $31.2M $24.9M–$63.0M | -1% |
Claude 3.5 Sonnet Anthropic · Jun 2024 | 2.7×10²⁵ T2 | undisclosed | $25.9M | Epoch AI (T2) + 1 paired source | $28.5M $22.7M–$57.5M | +10% |
GPT-4 (Mar 2023) OpenAI · Mar 2023 | 2.1×10²⁵ T2 | A100 SXM4 × ~25,000 | $37.3M | Epoch AI (T2) + 1 paired source | $37.1M $29.6M–$74.9M | -1% |
Nemotron-4 340B Nvidia · Jun 2024 | 1.8×10²⁵ T2 | H100 80GB SXM5 × 6,144 | $21.3M | Epoch AI (T2) + 1 paired source | $19.0M $15.1M–$38.3M | -11% |
Inflection-2 Inflection AI · Nov 2023 | 1×10²⁵ T1 | H100 × 5,000 | $13.5M | Epoch AI (T2) + 1 paired source | $10.5M $8.4M–$21.3M | -22% |
GPT-4 (Jun 2023) OpenAI · Jun 2023 | 2.1×10²⁵ T2 | — | — | — | not run | no published cost to compare against; no reproduction configuration recorded for this run |
Pangu Ultra Huawei Technologies · Apr 2025 | 1.1×10²⁵ T2 | — | — | — + 1 paired source | not run | no published cost to compare against; no reproduction configuration recorded for this run |
frontier_model_runs.A 10²⁵ FLOP run on H100 SXM5 (the EU AI Act §51 trigger), held at 1-year reserved cloud pricing, swept on three axes. Each panel moves one assumption, holds the rest, and prints the price operands that moved with it under every row. Each row is priced at whichever vendor is cheapest for that row, so a column reads as a set of separate quotes rather than one vendor’s rate card.
The only axis on this page whose sweep holds the vendor as well: every row resolves to the same 1-year reserved rate, so the spread here is the assumption alone. MFU 0.40 is the calculator's default.
The vendor moves under this sweep too: 2 providers supply these 3 rows, which is why a slower accelerator can price above a faster one. H100 SXM5 and H200 SXM share one compute die at 989.5 TFLOP/s dense, and the newer part's advantage is memory. At equal FLOP budget and equal MFU, the 26% between them is entirely price.
Two series. Each tier is priced at whichever vendor is cheapest for that tier, 3 different vendors across the column. The lighter bars are one vendor's own ladder: every tier at azure, the cheapest provider quoting all four cloud terms. Where the two series diverge, the reading is about who quotes what: the vendors with the cheapest spot and on-demand rates quote no reserved terms, so the cheapest reserved rate belongs to somebody else.
The calculator answers what a new run would cost. This is what the historical record looks like on the same axes: the frontier cost trajectory beside the falling cost of reproducing a previous generation’s capability, February 2023 to January 2026, in 2023 dollars.
The calculator prices a run at published rates. The largest labs do not buy at published rates, and several of them train on hardware they own or on capacity inside a parent company. For those runs the output is an upper-bound cloud-equivalent. It does not reproduce internal cost, and no public source closes that gap.
Total development cost (experimentation, salaries, the runs that failed) is reported at three to eight times the training-compute figure, and far higher for at least one lab whose disclosed number is an organisation-wide bound. The multiplier appears in the historical scatter and is excluded from the calculator’s headline, because folding it in would produce a figure with no source.
Model FLOP utilization is set from published figures for runs that reported it (PaLM at 0.46, Llama 3 at 0.38). The bounds the calculator shows reflect that documented range. An empirical distribution over the runs on this page is not available: most of them, as the ladder shows, reported no utilization figure.
Coverage is the Epoch ≥10²⁵ frontier set plus verified lab-primary disclosures: 38 records over 29 models from 14 labs. It is not a history of every training run. A model absent from it is absent from the corpus; absence does not place it below the threshold. All figures are in 2023 dollars; no inflation or FX conversion is applied.