Capabilities
Frontier-model capability read against an incomplete public compute record.
Reading the capability substrates…
Frontier-model capability read against an incomplete public compute record.
Reading the capability substrates…
METR measures how long a task can be before a model stops finishing it half the time. The measure is the closest thing the field has to a scalar for autonomous capability, and it is routinely read against training compute — the quantity every compute-threshold regime in force actually gates on.
That reading is available for 18 of the 49 models METR publishes. It is not available for any of the ten highest. Inside the range METR states its task suite measures reliably, the most capable model finishes tasks of 12.0 h; the most capable one with a public training-compute estimate finishes tasks of 3.4 h, a factor of 3.5 below it.
One measurement sits above that range. METR's note on the dataset this view draws reads “Measurements above 16 hrs are unreliable with our current task suite”, and METR excludes such points from its own trend fit. Claude Mythos (Preview) is drawn here rather than dropped, at 17.4 h — which would put the frontier a factor of 5.1 above the compute record instead of 3.5. Both figures are on this page because the gap between them is the same fact the page is about, one instrument further out: the frontier has run past not only the public compute record but past the measurement.
Regressing log horizon on log training compute over the 18 models that have both numbers gives an R² of 0.66: each decade of training compute buys a factor of 4.8 in task horizon, at a fitted slope of 0.685 ± 0.124 log-minutes per decade. That is a real relationship and a genuinely loose one — at a given compute level the residual spread is a factor of 3.8 either way, which is what you would expect when post-training, scaffolding and elicitation effort all move the measured value and none of them is in the fit.
Two things weaken the fit further and both are visible in its own operands. It spans 5.4 decades of compute, from 1.9×10²¹ to 5.0×10²⁶ FLOP, and the plate above shows that span is not evenly populated. And the 18 points sit at only 14 distinct compute values, because configuration variants of one training run share an estimate: the fit treats them as independent observations, and they are not.
The binding limitation is still the range. The fit is anchored on models topping out at 3.4 h, and the frontier — even inside METR's stated measurement range — is at 12.0 h. Every claim of the form “this much compute buys an agent that works unsupervised for N hours” at frontier N is extrapolation past the last observation, not interpolation within the data. That a tighter fit makes this worse rather than better is the point: a well-fitting model extrapolated past its last observation is the more seductive error.
METR's own headline is a claim about time rather than compute, measured across the whole set, and it is the better-founded of the two readings for exactly that reason. The dataset this view cites publishes it in its own header: the horizon doubles every 4.23 months over models released from 2023 on (3.43–5.19 months), and every 6.17 months over the whole published record. Read the two together with care: METR fits the doubling time with the above-16-hour points excluded, and the frontier horizon on this page includes them, so the two quantities here are computed under different inclusion rules.
Not drawnA regression line and prediction band over the covered subset appeared on the previous version of this view. They are stated here instead of drawn: a band rendered across a compute axis whose upper reaches contain no frontier observation invites exactly the extrapolation the paragraph above warns against.
An hour is METR's legible cut-point. 23 models clear it; 19 of them have no public training-compute estimate — the set any compute-to-capability argument at the frontier has to reach over. The other 4 are the whole of the compute record above an hour.
Compute on recordAll 19 rows: none published.
| Model | Developer | Task horizon | METR interval | METR release |
|---|---|---|---|---|
| Claude Mythos (Preview) · above METR's stated measurement range | Anthropic | 17.4 h | 8.5 h – 55.1 h | METR-Horizon-v1.1 |
| Claude Opus 4.6 | Anthropic | 12.0 h | 5.3 h – 60.6 h | METR-Horizon-v1.1 |
| Gemini 3.1 Pro | Google DeepMind | 6.4 h | 3.9 h – 11.6 h | METR-Horizon-v1.1 |
| GPT-5.2 (high) | OpenAI | 5.9 h | 3.3 h – 13.6 h | METR-Horizon-v1.1 |
| GPT-5.3 Codex | OpenAI | 5.8 h | 3.2 h – 13.6 h | METR-Horizon-v1.1 |
| GPT-5.4 | OpenAI | 5.7 h | 3.1 h – 12.8 h | METR-Horizon-v1.1 |
| Claude Opus 4.5 | Anthropic | 4.9 h | 2.7 h – 10.4 h | METR-Horizon-v1.1 |
| Claude Opus 4.5 (16K) | Anthropic | 4.8 h | 1.8 h – 20.4 h | METR-Horizon-v1.0 |
| Gemini 3 Pro | Google DeepMind | 3.7 h | 2.3 h – 6.3 h | METR-Horizon-v1.1 |
| GPT-5.1 Codex Max | OpenAI | 3.7 h | 2.2 h – 6.6 h | release not yet attributed |
| Claude Sonnet 4.5 (16K) | Anthropic | 2.0 h | 56 min – 4.2 h | METR-Horizon-v1.0 |
| o3 | OpenAI | 2.0 h | 1.2 h – 3.2 h | METR-Horizon-v1.1 |
| Claude Opus 4.1 (16K) | Anthropic | 1.9 h | 58 min – 3.5 h | METR-Horizon-v1.0 |
| Claude Opus 4.1 | Anthropic | 1.7 h | 59 min – 2.7 h | METR-Horizon-v1.1 |
| Claude Opus 4 | Anthropic | 1.7 h | 60 min – 2.7 h | METR-Horizon-v1.1 |
| o3 (medium) | OpenAI | 1.5 h | 46 min – 2.7 h | METR-Horizon-v1.0 |
| Claude Opus 4 (16K) | Anthropic | 1.4 h | 45 min – 2.4 h | METR-Horizon-v1.0 |
| o4-mini (medium) | OpenAI | 1.3 h | 35 min – 2.5 h | METR-Horizon-v1.0 |
| Claude Sonnet 4 (16K) | Anthropic | 1.2 h | 38 min – 2.2 h | METR-Horizon-v1.0 |
The compute axis above is an attribute of models. This is the same axis read as an attribute of places: how long a facility Scrutica tracks would need to run to accumulate a training budget the size of the largest run on the figure above that has a public estimate — GPT-5, 6.6×10²⁵ FLOP, Epoch AI. It is not the largest figure on that plate: Grok 4 sits at 5.0×10²⁶ FLOP. “Frontier scale” here means the frontier of the public compute record, which is the smaller thing and the page's whole subject.
Scrutica publishes 4,257 facility records. 512 of them have either a disclosed GPU count or a nameplate power figure, which is what the estimator needs — so this table is a 12.0% slice of the published substrate, and the other 3,745 are absent from it for want of a disclosure rather than for want of capacity. Of the 512, 151 clear the reference budget in 90 days of continuous training; the 24 fastest are listed. Rows are facility records, not campuses: where a site is recorded as more than one facility, each appears on its own line with its own figures.
| Facility | Country | Est. daily FLOP | Days to reference budget | Range across the estimator's assumptions |
|---|---|---|---|---|
| Stargate UAE (OpenAI/G42/Oracle) · estimated | AE | 8.8×10²⁵ power | <1 | <1–4 |
| Meta Louisiana Datacenter · estimated | US | 4.0×10²⁵ power | 2 | <1–8 |
| AWS Project Rainier (New Carlisle, IN) · estimated | US | 3.9×10²⁵ power | 2 | <1–8 |
| Meta Richland Parish (Hyperion) · estimated | US | 3.5×10²⁵ power | 2 | <1–9 |
| Crusoe Cheyenne Wyoming · estimated | US | 3.2×10²⁵ power | 2 | 1–10 |
| DataVolt Neom 1.5 GW Phase 2 · estimated | SA | 2.7×10²⁵ power | 2 | 1–12 |
| IREN Sweetwater 1 · estimated | US | 2.5×10²⁵ power | 3 | 1–13 |
| Oracle Vantage Data Centers Frontier · estimated | US | 2.5×10²⁵ power | 3 | 1–13 |
| Stargate Abilene (OpenAI/Oracle/SoftBank) · estimated | US | 2.1×10²⁵ power | 3 | 2–16 |
| OpenAI/Microsoft Mt Pleasant, Wisconsin Phase 2 · estimated | US | 2.0×10²⁵ hardware · H100 SXM5 assumed | 3 | 2–9 |
| Crusoe Goodnight in Claude, Texas · estimated | US | 1.8×10²⁵ power | 4 | 2–19 |
| 1 Gig Data Center East Fishkill, NY | US | 1.8×10²⁵ power | 4 | 2–19 |
| xAI Colossus 2 Memphis Phase 2 · estimated | US | 1.6×10²⁵ hardware · H100 SXM5 assumed | 4 | 3–12 |
| Colossus 2 · estimated | US | 1.5×10²⁵ hardware · H100 SXM5 assumed | 4 | 3–12 |
| Fluidstack France Gigawatt Campus · estimated | FR | 1.5×10²⁵ hardware · H100 SXM5 assumed | 5 | 3–13 |
| Meta Prometheus New Albany · estimated | US | 1.5×10²⁵ hardware · H100 SXM5 assumed | 5 | 3–13 |
| IREN Childress · estimated | US | 1.3×10²⁵ power | 5 | 2–25 |
| Reliance Industries Supercomputer · estimated | IN | 1.3×10²⁵ hardware · H100 SXM5 assumed | 5 | 4–14 |
| Project Rainier · estimated | US | 1.2×10²⁵ hardware · H100 SXM5 assumed | 6 | 4–16 |
| Microsoft Fairwater Atlanta · estimated | US | 1.1×10²⁵ power | 6 | 3–29 |
| Meta Prometheus · estimated | US | 1.1×10²⁵ power | 6 | 3–30 |
| IREN Sweetwater 2 · estimated | US | 1.1×10²⁵ power | 6 | 3–31 |
| Micron Fab 2 | US | 10²⁵ power | 6 | 3–32 |
| TeraWulf Lake Mariner Campus · estimated | US | 9.0×10²⁴ power | 7 | 4–37 |
Lower bound“Days to” assumes 24/7 single-model training at full facility capacity, which no real frontier run achieves once shared inference, failures, restarts and evaluation runs are accounted for. The range beside it is the estimator's parameter-sensitivity band — interconnect efficiency and model-FLOP utilisation across their documented ranges — and is not a confidence interval: there is no distributional model behind it. Derivations: Methodology.
A cumulative-FLOP threshold is a proxy. It is administrable — a provider knows its own training budget, and the figure is auditable after the fact — which is a real virtue and the reason the instrument exists in the form it does. What it is not is a measurement of the thing being regulated. The horizontal axis above is the proxy; the vertical axis is closer to the concern.
Two facts on this page bear on how tightly the proxy binds. The fit between them is real but loose (R² 0.66 over 18 observations at 14 distinct compute values), and it is unobserved exactly where the policy questions now are. A reader should take neither as an argument against compute thresholds — the alternative instruments have their own problems, and the capability side is floor-bound in a way that is arguably worse. It is an argument for knowing which of the two quantities a given claim actually rests on.