<!--
Source: https://scrutica.com/economics/training-cost
Generated: 2026-09-04T08:30:09.562Z
Format: Markdown extraction of the rendered HTML at the source URL.
For the full agent guide see: https://scrutica.com/llms-full.txt
For the MCP server see: https://scrutica.com/api/mcp
-->

# Training Economics
## Two of fourteen frontier runs have a compute figure the lab itself published

**Vintage**38 disclosure records across 14 labs · 9 multi-source pairingsprice layer verified 2026-03-29 – 2026-09-04 (oldest and newest rows; three providers refresh daily, three are manual snapshots)

Training compute (FLOP, log)1×10²⁵1×10²⁶1×10²⁷10²⁵ · EU AI Act systemic-risk triggerstatedinputs onlyhardware or sizequalitativedeclined, on the recordLlama 3.1-405B · Meta Platforms, Inc.. The lab published 3.8×10²⁵. Third-party estimate: 3.8×10²⁵.Llama 3.1-405BMeta Platforms, Inc. · Jul 2024Llama 4 Behemoth (preview) · Meta Platforms, Inc.. The lab published no training-compute figure. Third-party estimate: 5.2×10²⁵. Inferred from disclosed inputs: 3.0×10²⁵–7.0×10²⁵ (tier 4).Llama 4 Behemoth (preview)Meta Platforms, Inc. · Apr 2025Nemotron-4 340B · Nvidia. The lab published no training-compute figure. Third-party estimate: 1.8×10²⁵. Inferred from disclosed inputs: 1.5×10²⁵–2.0×10²⁵ (tier 4).Nemotron-4 340BNvidia · Jun 2024Pangu Ultra · Huawei Technologies. The lab published no training-compute figure. Third-party estimate: 1.1×10²⁵. Inferred from disclosed inputs: 1.0×10²⁵–1.3×10²⁵ (tier 4).Pangu UltraHuawei Technologies · Apr 2025DeepSeek-V3 · DeepSeek. The lab published no training-compute figure.DeepSeek-V3DeepSeek · Dec 2024Gemini 1.0 Ultra · Google DeepMind. The lab published no training-compute figure. Third-party estimate: 5×10²⁵. No figure and no derivable range; recorded at this corpus's 1×10²⁵ admission floor, Scrutica's own bound (tier 4).Gemini 1.0 UltraGoogle DeepMind · Dec 2023Inflection-2 · Inflection AI. The lab published 1×10²⁵. Third-party estimate: 1×10²⁵.Inflection-2Inflection AI · Nov 2023Grok 4 · xAI. The lab published no training-compute figure. Third-party estimate: 5×10²⁶.Grok 4xAI · Jul 2025GPT-4.5 · OpenAI. The lab published no training-compute figure. Third-party estimate: 3.8×10²⁶.GPT-4.5OpenAI · Feb 2025Grok 3 · xAI. The lab published no training-compute figure. Third-party estimate: 3.5×10²⁶.Grok 3xAI · Feb 2025Grok-2 · xAI. The lab published no training-compute figure. Third-party estimate: 3.0×10²⁵.Grok-2xAI · Aug 2024Claude 3.5 Sonnet · Anthropic. The lab published no training-compute figure. Third-party estimate: 2.7×10²⁵. No figure and no derivable range; recorded at this corpus's 1×10²⁵ admission floor, Scrutica's own bound (tier 4).Claude 3.5 SonnetAnthropic · Jun 2024GPT-4 (Jun 2023) · OpenAI. The lab published no training-compute figure. Third-party estimate: 2.1×10²⁵.GPT-4 (Jun 2023)OpenAI · Jun 2023GPT-4 (Mar 2023) · OpenAI. The lab published no training-compute figure. Third-party estimate: 2.1×10²⁵. No figure and no derivable range; recorded at this corpus's 1×10²⁵ admission floor, Scrutica's own bound (tier 4).GPT-4 (Mar 2023)OpenAI · Mar 2023the lab published this figurea third party estimated itinferred from disclosed inputs (tier 4) · an empty track means nothing was published

_The disclosure ladder._ One track per cross-validated frontier run, ordered by how much of the arithmetic its lab left for someone else to do: from a FLOP figure stated in the lab’s own paper down to a technical report that documents its refusal to state one.

Position encodes training compute on a log axis. The glyph encodes who is speaking, the only categorical claim the figure makes: a filled mark is a figure the lab published, an open ring is a third party’s estimate, a bracket is an inference from inputs the lab did publish and is authority tier 4.

A track with no mark is a run for which nobody has published a figure _this axis can place_: the axis is training compute in FLOP, so a lab that published GPU-hours or a dollar figure and no FLOP has an empty track and a rung that says so. An empty track is still drawn.

**The 10²⁵ upright is reference apparatus.** It marks the EU AI Act systemic-risk trigger because several runs land near it. A row whose only lower bound equals that figure shows the corpus’s admission floor for a run its lab calls frontier-class without publishing a number; the floor is Scrutica’s own, and the plate does not draw it as a lab mark.

**Release dates make the ladder a chronology.** The most recent run whose lab published a compute figure was released in July 2024; the 7 cross-validated runs released since fall on the inputs only and qualitative rungs, and none states a figure. The count is over the 14 runs on this ladder only, and it measures what labs published; what was trained is a different question.

**Of fourteen cross-validated frontier runs, two reach the 1025 systemic-risk trigger on a figure the lab itself published, and eleven only on a third party’s estimate.** No run crosses on this corpus’s admission floor alone; every row whose only lower bound is the floor also has a third-party estimate, and is counted there. One has no compute figure this axis can place, so it is counted in none of the three. One of the crossing runs sits exactly at the trigger, on a figure rounded to one significant digit, which is why the sentence above says "reach". Each run is counted once, against the best-authority figure it has; the threshold comes from the Act, and the classification comes from this corpus.

Every run, what its lab said, and what had to be estimated

Run

What the lab published

Lab’s figure

Estimated

Cost on record

Llama 3.1-405B · Meta Platforms, Inc.

The lab published the figure [source](https://arxiv.org/abs/2407.21783)

3.8×10²⁵

3.8×10²⁵

$53M · estimated

Llama 4 Behemoth (preview) · Meta Platforms, Inc.

The lab published the inputs only [source](https://ai.meta.com/blog/llama-4-multimodal-intelligence/)

not published

5.2×10²⁵

$45M · estimated

Nemotron-4 340B · Nvidia

The lab published the inputs only [source](https://arxiv.org/abs/2406.11704)

not published

1.8×10²⁵

$21M · estimated

Pangu Ultra · Huawei Technologies

The lab published the inputs only [source](https://arxiv.org/abs/2504.07866)

not published

1.1×10²⁵

none

DeepSeek-V3 · DeepSeek

The lab published the inputs only [source](https://arxiv.org/abs/2412.19437)

not published

none

$5.6M · lab-stated

Gemini 1.0 Ultra · Google DeepMind

The lab named only the hardware or the size [source](https://arxiv.org/abs/2312.11805)

not published

5×10²⁵

$31M · estimated

Inflection-2 · Inflection AI

The lab named only the hardware or the size [source](https://inflection.ai/inflection-2)

1×10²⁵

1×10²⁵

$13M · estimated

Grok 4 · xAI

The lab said only that the model is frontier-class

not published

5×10²⁶

$388M · estimated

GPT-4.5 · OpenAI

The lab said only that the model is frontier-class

not published

3.8×10²⁶

$366M · estimated

Grok 3 · xAI

The lab said only that the model is frontier-class

not published

3.5×10²⁶

$218M · estimated

Grok-2 · xAI

The lab said only that the model is frontier-class

not published

3.0×10²⁵

$32M · estimated

Claude 3.5 Sonnet · Anthropic

The lab said only that the model is frontier-class [source](https://www.anthropic.com/news/claude-3-5-sonnet)

not published

2.7×10²⁵

$26M · estimated

GPT-4 (Jun 2023) · OpenAI

The lab said only that the model is frontier-class

not published

2.1×10²⁵

none

GPT-4 (Mar 2023) · OpenAI

The lab documented that it would not say [source](https://arxiv.org/abs/2303.08774)

not published

2.1×10²⁵

$37M · estimated

Most public figures for the cost of a frontier training run are built on a training-compute number the lab never published. Of the fourteen runs above, two have a compute figure the lab itself published; for eleven the only number on record was produced by somebody else, working from whatever the lab did disclose: a GPU count, a hardware class, a parameter total, sometimes a sentence calling the model frontier-class. Of twelve training costs on record, exactly one came from the lab that trained the model. The estimates themselves are careful, and they are the reason any of this can be discussed. But a cost figure of this kind is an inference over an undisclosed quantity, priced at rates the buyer did not pay. The calculator below is one more such inference, built so every assumption is visible and adjustable.

### What a run would cost

Cite

CSVJSONJSON+Prov

PrintScreenshot

The calculator prices three dimensions: how much compute the run needs, what hardware it runs on, and how the hardware is procured. It resolves them against the live price layer the [Cost Index](/economics/cost-index) maintains, and against a three-year on-prem amortization in parallel; both are stated in the same unit. The default budget is Llama 3.1 405B’s disclosed 3.8×10²⁵ FLOP, and the EU AI Act’s [10²⁵ FLOP trigger](/capacity/thresholds) is one of the preset budgets. [How the estimate is derived↗](/methodology#training-cost-decomposition "How this is derived: training-cost decomposition. Frontier vs replication cost curves in 2023 USD; total-development multiplier range applied to compute-only cost.")

### Against the runs that were disclosed

The same calculation, run against each disclosed model at its documented hardware, with the per-row gap to the figure on record shown. Where that figure is itself an estimate, which the ladder above shows is nearly everywhere, the comparison is between two inferences, and the table says which is which.

Cross-validation · disclosed runs

## Cross-validation against publicly-disclosed runs

Each row reproduces a published cost figure inside the calculator above; Δ is the reproduction-versus-disclosed gap, with the ±20% band marking the coarse external-validity threshold. The table reads from `frontier_model_runs` (multi-source per cell, best-tier-available picked per column); where both Epoch and the lab disclose, the higher-authority figure leads with the source labelled. **5 of 11 reproductions fall inside ±20%**; the 6 outside are Gemini 1.0 Ultra, Grok 3, Grok 4, Llama 3.1-405B, Llama 4 Behemoth (preview), Inflection-2, each with the assumption that moves it named below. Gaps of that size are the expected scale of an inference over quantities the labs did not disclose. A further 3 of 14 runs are listed in the table without a reproduction; the Δ column states the reason.

Model

FLOP

Hardware

Reported $

Source

Scrutica calc

Δ

DeepSeek-V3

DeepSeek · Dec 2024

not published

H800 SXM5 × 2,048

$5.6M

[Lab paper (T1)](https://arxiv.org/abs/2412.19437)

not run

no central FLOP on record; the lab published GPU-hours instead

Grok 4

xAI · Jul 2025

5×10²⁶

T2

H100 SXM5 × 200,000 (Colossus full)

$387.8M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

$527.2M

$420.4M–$1.1B

+36%

GPT-4.5

OpenAI · Feb 2025

3.8×10²⁶

T2

undisclosed

$366.0M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

$400.7M

$319.5M–$809.1M

+9%

Grok 3

xAI · Feb 2025

3.5×10²⁶

T2

H100 SXM5 × 80,000 (Colossus subset)

$217.8M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

$369.0M

$294.3M–$745.3M

+69%

Llama 4 Behemoth (preview)

Meta Platforms, Inc. · Apr 2025

5.2×10²⁵

T2

H100 SXM5 × 32,000 (FP8)

$44.6M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

\+ 1 paired source

$54.7M

$43.6M–$110.4M

+23%

Gemini 1.0 Ultra

Google DeepMind · Dec 2023

5×10²⁵

T2

TPUv4 × ~57,000 (Epoch derivation)

$30.7M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

\+ 1 paired source

$52.7M

$42.0M–$106.5M

+72%

Llama 3.1-405B

Meta Platforms, Inc. · Jul 2024

3.8×10²⁵

T1

H100 SXM5 80GB × 16,384

$52.9M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

\+ 1 paired source

$40.1M

$31.9M–$80.9M

\-24%

Grok-2

xAI · Aug 2024

3.0×10²⁵

T2

H100 SXM5 (count undisclosed)

$31.6M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

$31.2M

$24.9M–$63.0M

\-1%

Claude 3.5 Sonnet

Anthropic · Jun 2024

2.7×10²⁵

T2

undisclosed

$25.9M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

\+ 1 paired source

$28.5M

$22.7M–$57.5M

+10%

GPT-4 (Mar 2023)

OpenAI · Mar 2023

2.1×10²⁵

T2

A100 SXM4 × ~25,000

$37.3M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

\+ 1 paired source

$37.1M

$29.6M–$74.9M

\-1%

Nemotron-4 340B

Nvidia · Jun 2024

1.8×10²⁵

T2

H100 80GB SXM5 × 6,144

$21.3M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

\+ 1 paired source

$19.0M

$15.1M–$38.3M

\-11%

Inflection-2

Inflection AI · Nov 2023

1×10²⁵

T1

H100 × 5,000

$13.5M

[Epoch AI (T2)](https://epoch.ai/data/notable-ai-models)

\+ 1 paired source

$10.5M

$8.4M–$21.3M

\-22%

GPT-4 (Jun 2023)

OpenAI · Jun 2023

2.1×10²⁵

T2

—

—

—

not run

no published cost to compare against; no reproduction configuration recorded for this run

Pangu Ultra

Huawei Technologies · Apr 2025

1.1×10²⁵

T2

—

—

—

\+ 1 paired source

not run

no published cost to compare against; no reproduction configuration recorded for this run

DeepSeek-V3. Not reproduced: no central FLOP on record; the lab published GPU-hours instead. Paper §3.2 reports 2.788M H800 GPU-hours at internal $2/GPU-hr accounting → $5.576M final-run cost. On-prem 3-year amortization is the closest defensible reproduction baseline.

Grok 4. Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. xAI Grok 4 announcement attributes 200K-H100 cluster. Calculator runs against Epoch central 5.0e26.

GPT-4.5. Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. OpenAI does not disclose hardware. Calculator runs against Epoch central 3.8e26.

Grok 3. Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. xAI Memphis Colossus is on-prem; Grok 3 attributed to a Colossus subset. On-prem 3-year amortization is the appropriate reproduction baseline.

Llama 4 Behemoth (preview). Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. Meta blog discloses 32K-GPU cluster + FP8 training + 390 TFLOPs/GPU throughput; Epoch attributes H100 SXM5. On-prem 3-year amortization.

Gemini 1.0 Ultra. Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. Google paper discloses TPUv4 hardware class only; chip count not stated. Epoch derives ~57,000 TPUv4. Calculator runs on H100-equivalent throughput as proxy — figure indicative, not directly comparable.

Llama 3.1-405B. Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. Llama 3 paper §3.3.1 discloses 3.8×10²⁵ FLOPs and 16,384 H100 SXM5; Meta runs internal cluster (no retail rate published). On-prem 3-year amortization is the right baseline.

Grok-2. Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. xAI Memphis Colossus pre-Grok-3 scaling. Calculator runs against Epoch central 3.0e25.

Claude 3.5 Sonnet. Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. Anthropic does not disclose compute on system cards. Anthropic compute runs on AWS Trainium and GCP TPU under multi-year contracts at heavily-discounted rates; on-prem 3-year amortization is the closest defensible baseline.

GPT-4 (Mar 2023). Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. OpenAI Tech Report explicitly does not disclose hardware or compute. Microsoft built a dedicated supercomputer for OpenAI; the effective price is capex amortization, not retail cloud reserved. Cottier & Rahman (arXiv:2405.21015) total dev cost $90–100M.

Nemotron-4 340B. Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. NVIDIA paper discloses 768 DGX H100 nodes × 8 GPUs/node = 6,144 H100 80GB SXM5. FLOP not stated; calculator uses Epoch central estimate 1.8e25.

Inflection-2. Reproduction: On-prem · 3-year amortized at MFU 0.40 · interconnect 0.85. Inflection blog: "5,000 NVIDIA H100 GPUs in fp8 mixed precision for ~10²⁵ FLOPs". Lab disclosure of compute scale at the EU threshold.

Reading the Δ column

Δ is the calculator-vs-published gap under one procurement assumption per row. When the calculator runs lower than the disclosure, the cause is typically MFU 0.40 acting as a conservative default against an estimate that folded in R&D or assumed a higher implicit price; when it runs higher, the cause is typically a final-run-only published estimate against the cloud-pricing overhead the calculator includes. The column is a per-row audit point; it does not score the estimates.

[Source: Epoch AI](https://epoch.ai/data), CC BY 4.0 · lab-primary rows from each lab’s canonical paper / system card / announcement (URLs in the Source column above). Source table: `frontier_model_runs`.

Sensitivity

## Sensitivity: what would change the answer

A 10²⁵ FLOP run on H100 SXM5 (the EU AI Act §51 trigger), held at 1-year reserved cloud pricing, swept on three axes. Each panel moves one assumption, holds the rest, and prints the price operands that moved with it under every row. Each row is priced at whichever vendor is cheapest for that row, so a column reads as a set of separate quotes rather than one vendor’s rate card.

MFU

Effective utilization fraction · everything else held

MFU 0.25$103.9M

azure · East US · $62.92/GPU-hr · $190.78/peak-PFLOP-day

MFU 0.30$86.6M

azure · East US · $62.92/GPU-hr · $190.78/peak-PFLOP-day

Epoch 2022 LLM recommendation

MFU 0.40$64.9M

azure · East US · $62.92/GPU-hr · $190.78/peak-PFLOP-day

the panel default

MFU 0.50$52.0M

azure · East US · $62.92/GPU-hr · $190.78/peak-PFLOP-day

PaLM-class

MFU 0.60$43.3M

azure · East US · $62.92/GPU-hr · $190.78/peak-PFLOP-day

The only axis on this page whose sweep holds the vendor as well: every row resolves to the same 1-year reserved rate, so the spread here is the assumption alone. MFU 0.40 is the calculator's default.

Hardware

GPU model · 1-year reserved · MFU 0.40 · cheapest provider per model

A100 SXM$66.0M

aws · US West (Oregon) · $20.18/GPU-hr · $194.00/peak-PFLOP-day

H100 SXM5$64.9M

azure · East US · $62.92/GPU-hr · $190.78/peak-PFLOP-day

H200 SXM$48.0M

azure · West US 2 · $46.52/GPU-hr · $141.04/peak-PFLOP-day

B200—

No 1-year reserved pricing in substrate for B200. Try a different procurement tier or region.

The vendor moves under this sweep too: 2 providers supply these 3 rows, which is why a slower accelerator can price above a faster one. H100 SXM5 and H200 SXM share one compute die at 989.5 TFLOP/s dense, and the newer part's advantage is memory. At equal FLOP budget and equal MFU, the 26% between them is entirely price.

Procurement

H100 SXM5 · MFU 0.40 · cheapest provider per tier

Cloud · spot$18.8M

azure$18.8M

azure · East US 2 · $18.17/GPU-hr · $55.09/peak-PFLOP-day

Cloud · on-demand$35.4M

azure$106.0M

lambda · US (Texas) · $4.29/GPU-hr · $104.05/peak-PFLOP-day

Cloud · 1-year reserved$64.9M

azure$64.9M

azure · East US · $62.92/GPU-hr · $190.78/peak-PFLOP-day

Cloud · 3-year reserved$44.5M

azure$44.5M

aws · US East (N. Virginia) · $43.16/GPU-hr · $130.84/peak-PFLOP-day

On-prem · 3-year amortized$10.5M

three-year amortization — no vendor quote, so no vendor moves

**Two series.** Each tier is priced at whichever vendor is cheapest for that tier, 3 different vendors across the column. The lighter bars are one vendor's own ladder: every tier at azure, the cheapest provider quoting all four cloud terms. Where the two series diverge, the reading is about who quotes what: the vendors with the cheapest spot and on-demand rates quote no reserved terms, so the cheapest reserved rate belongs to somebody else.

### Cost against capability, over twenty runs

The calculator answers what a new run would cost. This is what the historical record looks like on the same axes: the frontier cost trajectory beside the falling cost of reproducing a previous generation’s capability, February 2023 to January 2026, in 2023 dollars.

### What this page cannot do

#### 1 · It cannot see what the labs actually paid

The calculator prices a run at published rates. The largest labs do not buy at published rates, and several of them train on hardware they own or on capacity inside a parent company. For those runs the output is an upper-bound cloud-equivalent. It does not reproduce internal cost, and no public source closes that gap.

#### 2 · It prices the final run only

Total development cost (experimentation, salaries, the runs that failed) is reported at three to eight times the training-compute figure, and far higher for at least one lab whose disclosed number is an organisation-wide bound. The multiplier appears in the historical scatter and is excluded from the calculator’s headline, because folding it in would produce a figure with no source.

#### 3 · Utilization is a calibrated parameter

Model FLOP utilization is set from published figures for runs that reported it (PaLM at 0.46, Llama 3 at 0.38). The bounds the calculator shows reflect that documented range. An empirical distribution over the runs on this page is not available: most of them, as the ladder shows, reported no utilization figure.

#### 4 · The corpus is the frontier set

Coverage is the Epoch ≥10²⁵ frontier set plus verified lab-primary disclosures: 38 records over 29 models from 14 labs. It is not a history of every training run. A model absent from it is absent from the corpus; absence does not place it below the threshold. All figures are in 2023 dollars; no inflation or FX conversion is applied.