Capabilities
Frontier-model capability read against an incomplete public compute record.
Reading the capability substrates…
Frontier-model capability read against an incomplete public compute record.
Reading the capability substrates…
In print, and in their own published frameworks, 6 organizations have said that their capability numbers are lower bounds. OpenAI's Preparedness Framework calls one-time elicitation “a lower bound, rather than a ceiling.” METR attaches a figure — around two to three engineer-weeks of iterative development, spent on the two models it singles out — to the elicitation effort, not to the gap that effort left. The UK AI Security Institute names the specific choices that hold its numbers down. What none of them supplies is the multiplier that would convert a published score into a ceiling.
Two benchmark publishers happen to disclose more than one scaffold for the same model, which makes a piece of the gap measurable. It does not produce the missing multiplier. It produces something more useful: evidence that the multiplier cannot exist. Across the 8 paired models the fold change runs 1.56× to 4.68× — and 3 of the 24 model pairs that either column can separate are ordered one way by the first column and the other way by the last.
Cybench publishes each model's score three ways: the task alone, the task with its subtask decomposition available, and scoring on the decomposed subtasks. Nothing about the model changes between the columns. Everything a reader would infer from the number does.
| # | Model | Subtasks | Fold change | Against Unguided |
|---|---|---|---|---|
| 1 | o1-preview OpenAI | 46.8% | 4.68× | up 3 places |
| 2 | Claude 3.5 Sonnet (2024-06) Anthropic | 43.9% | 2.51× | down 1 place |
| 3 | Claude 3 Opus Anthropic | 36.8% | 3.68× | unchanged |
| 4 | GPT-4o (2024-11) OpenAI | 28.7% | 2.30× | down 2 places |
| 5 | Llama 3.1 405B Instruct Meta AI | 20.5% | 2.73× | unchanged |
| 6 | Mixtral 8x22B Instruct Mistral AI | 15.2% | 2.03× | up 1 place |
| 7 | Gemini 1.5 Pro Google DeepMind | 11.7% | 1.56× | down 1 place |
| 8 | Llama 3 70B Instruct Meta AI | 8.2% | 1.64× | unchanged |
A uniform bias is easy to reason about — if every score is low by the same factor, comparisons survive and only the absolute level moves. These are not that.
Both readings are correct and both are published by the same evaluator on the same page. A policy artefact that cites “the Cybench score” for either model has silently chosen one of them, and the choice is not visible in the citation. This is the part of the floor problem that a uniform-discount intuition does not cover: the elicitation gap is not a level shift, it is a reordering, and reorderings survive normalization.
OSWorld publishes each model at three step budgets — 15, 50 and 100 actions before the run is cut off. The budget is not a property of the model; it is a property of the evaluation.
| Model | 15 steps | 50 steps | 100 steps | 100 / 15 |
|---|---|---|---|---|
| Claude Sonnet 4.5 Anthropic | 42.9 | 58.1 | 62.9 | 1.47× |
| Claude Sonnet 4 Anthropic | 31.2 | 43.9 | 41.4 ↓ | 1.33× |
| Claude 3.7 Sonnet Anthropic | 27.1 | 35.8 | 35.6 ↓ | 1.31× |
| Computer Use Preview OpenAI | 26.0 | 31.3 | 30.5 ↓ | 1.17× |
| o3 (medium reasoning) OpenAI | 9.1 | 17.2 | 23.0 | 2.53× |
3 of the 5 models score lower at 100 steps than at 50 (marked ↓). More room is not monotonically better — an agent given a longer leash can spend it compounding an error, which is a finding about agent behaviour that the single published number for each model conceals entirely.
On the statistic this page is built on, though, the second axis is not the same thing — and that is worth more than the parallel would have been. 0 of the 10 pairs reverse between the 15 steps and 100 steps budgets, none of them tied. Budget moves levels; scaffold moves ranks. Every model gains from a longer leash and the leaderboard order survives it, where the scaffold choice reorders 3 of 24 separable pairs. The two are different kinds of evaluation choice, and only one of them threatens an argument that rests on a comparison.
And the budget can outweigh a model generation: Claude 3.7 Sonnet at 100 steps scores 35.6 against Claude Sonnet 4 at 15 steps scoring 31.2 — the older model ahead of its own direct successor, same lab and same family, on the same benchmark under the same scoring, because it was given more room. It is not an isolated case: 4 of the 10 ordered older-then-newer pairs in this table cross the same way.
| Older model, 100 steps | Score | Newer model, 15 steps | Score | Margin |
|---|---|---|---|---|
| Claude 3.7 Sonnet | 35.6 | Claude Sonnet 4 | 31.2 | 4.4 |
| Computer Use Preview | 30.5 | o3 (medium reasoning) | 9.1 | 21.4 |
| Claude 3.7 Sonnet | 35.6 | o3 (medium reasoning) | 9.1 | 26.5 |
| Claude 3.7 Sonnet | 35.6 | Computer Use Preview | 26.0 | 9.6 |
Verbatim from each linked primary source. The framing tags are descriptive — they say what kind of statement each is, not that the organization endorses this page's consolidation of them, which none has been asked to do.
OpenAI
Explicit floor language“we regard any one-time capability elicitation in a frontier model as a lower bound, rather than a ceiling, on capabilities that may emerge in real world use and misuse.”
OpenAI Preparedness Framework v2 (the "lower bound, rather than a ceiling" verbatim) · 2025-04-15 · §3.1 “Evaluation approach” (p. 9) — OpenAI Preparedness Framework v2. The floor is stated directly, in a single sentence.
METR
Qualified acknowledgement“We have put a limited amount of effort into eliciting models to get good performance on our tasks, so while our results are a reasonable lower bound, some models may have somewhat greater capabilities than we demonstrate. The most work was done to elicit o1 and the original Claude 3.5 Sonnet, each of which had around 2-3 engineer weeks of iterative development.”
Measuring AI Ability to Complete Long Software Tasks (METR) · 2025-03-18 · Appendix E.4 "Limitations and future work", subsection "More models with better elicitation" (v4) — METR Time Horizons paper, the most quantified disclosure on this page. The figure it quantifies is the elicitation effort spent. No figure here states the capability difference that effort produced.
UK AI Security Institute
Explicit floor language“The cap, alongside our use of a simple agent scaffold, artificially lowers success rates and understates what models can do with more tokens and stronger scaffolds.”
How fast is autonomous AI cyber capability advancing? · 2026-05-13 · Section “Cyber Time Horizons Results”, paragraph following Figure 1 — AISI ties the floor to two specific design choices it can name: the 2.5M-token-per-task budget and the simple agent scaffold. The cap (not "step cap") is a token cap.
Anthropic
Structural argument about evaluation“more general requirements to (a) match expected efforts of potential adversaries… Although still an aspirational goal, the science of evaluations is not currently mature enough to make confident predictions about the precise buffer we should require between current models and a Capability Threshold.”
Anthropic Responsible Scaling Policy v3.0 (evaluation-methodology requirements) · 2026-02-24 · Changelog (p. 17), note "Less prescriptive evaluation methodology" — Anthropic RSP v3.0, changelog note. The requirement is to match expected adversary effort, and the same note states that the buffer science is immature. This is a statement about measurement; it puts no multiplier on the gap.
Google DeepMind
Qualified acknowledgement“Risk assessment must take into account the fact that other actors may put significantly more effort into eliciting capabilities than we put into assessing risk, thus requiring conservatism in the form of evaluations.”
Google DeepMind Frontier Safety Framework v3.0 (out-elicitation acknowledgement, ML R&D CCLs) · 2025-09-22 · §1.3 Outline of Our Risk Assessment Process — “Note on Machine Learning R&D CCLs” — Frontier Safety Framework v3.0, in a note attached to the Machine Learning R&D capability levels. The note does not extend to the framework’s evaluations generally. DeepMind sets the limit itself in the next sentence: as a frontier AI company it does not expect other groups to put significantly more effort into that particular domain than it does. The scope is stated here because read loose the sentence sounds like a standard for every evaluation, and the document does not set one.
Apollo Research
Structural argument about evaluation“Since evals often aim to estimate an upper bound of capabilities, it is important to understand how to elicit maximal rather than average capabilities.”
We need a science of evals (Apollo Research) · 2024-01-22 · Introduction, before the first section heading — Apollo Research, "We need a science of evals." The structural argument is that elicitation quality determines what an upper-bound score can mean. The post supports it with a cited external 2023 study measuring up to 76 accuracy points of difference from prompt-format changes alone; that measurement is relayed from the cited study, and is not Apollo’s own.
The AISI cyber-range methodology paper (Folkerts, Payne, Inman et al., arXiv 2603.11214, 11 March 2026) defines the instrument: two ranges — a 32-step corporate-network attack and a 7-step ICS attack — under a 2.5M-token-per-task budget, across seven models from August 2024 to February 2026 at varying inference-time compute. The Mythos Preview evaluation (13 April 2026) is one run on it. On 13 May 2026 the institute named the two things holding its published numbers below the ceiling: the token cap, and a “simple agent scaffold” held constant across models so the cross-model comparison stays clean.
Inspect — the evaluation framework AISI itself maintains, at UKGovernmentBEIS/inspect_ai — added deepagent() in v0.3.213 on 27 April 2026: subagent delegation, persistent memory, structured planning. The published cyber methodology has not adopted it, and the reason is a real methodological commitment rather than an oversight. Running stronger scaffolds for later models would break the across-time comparison against every model already on the leaderboard. The institute's own reading is the cleanest one available: the held-constant call is correct for comparability, and the correct call understates capability.
AISI's public cyber evaluations cap each task at 2.5 million tokens of model output and use what the institute calls a “simple agent scaffold,” held constant across models so comparisons between them are like-for-like. AISI is explicit that both choices push scores down. They are the right choices for measuring relative progress and the wrong instrument for reading any single number as a ceiling.
Two weeks after the Mythos cyber evaluation shipped, the evaluation framework AISI maintains gained a multi-agent capability its own published methodology cannot use without forfeiting comparability against the models already measured. The distance between what the framework can run and what the public methodology reports is a known property of the instrument, documented by the institute.
Not that the Subtasks column or the 100-step budget is a capability ceiling. Each is the strongest mode its publisher chose to disclose, and the distance from there to a genuine maximum-elicitation result is unmeasured by construction — nobody ran it, and the figures above would not see it if they had.
Not that ranking instability generalises across evaluation choices. It is a property of the scaffold axis on this evidence and not of the budget axis: 0 of 10 OSWorld pairs reverse between the 15 steps and 100 steps budgets. Two datasets, one statistic, opposite results — and five models on one benchmark is a narrow base for either half, so neither is offered as a law about evaluation.
Not a multiplier. Two measured fold-change series exist in the public literature, and neither is a gap factor: the Cybench fold change, mean 2.64× and median 2.40× across 8 models; the OSWorld 100/15 lift, mean 1.56× and median 1.33× across 5 models. They are not commensurable — each captures a different cap-versus-scaffold delta on a different instrument — and the spread within even one of them defeats the idea of a single factor. The one further figure the sources put a number on is METR's elicitation budget, and that measures the effort spent, not the distance it closed. Anyone quoting 2×, 10× or 100× across the frontier is past what the sources support.
The structural reason the number is missing matters more than the number. Maximum-elicitation of cyber or CBRN capability is operationally close to building the misuse tool, and the legal, ethical and reputational considerations that keep public evaluators from publishing one apply identically across every domain Article 51 is drawn around. The evaluation regimes nearest the ceiling — internal pre-deployment work under RSP, Preparedness or FSF review, and any classified national-security evaluation the public need not be told exists — are exactly the ones whose results the public cannot read. The publication asymmetry is why the gap is open-ended, not evidence that a specific multiplier is known in some other room.
EU AI Act — Article 51(2) presumes systemic risk above 1025 FLOP, and the rebuttal cites capability evaluations.
The presumption is rebuttable: the provider demonstrates that the model does not pose systemic risk despite crossing the compute line, and the demonstration leans on capability evaluations. The July 2025 GPAI Guidelines codify a ±30% measurement tolerance on the cumulative-FLOP side. No comparable tolerance is documented on the capability side, because the capability side has not specified a measurement regime to put bounds on. A tolerance-bounded input feeding a floor-only output is the asymmetry this page renders.
And the reversals sharpen it. A rebuttal that argues from comparison — this model scores below that one, which was cleared — depends on an ordering that the evidence above shows is not stable across evaluation choices the citation does not record. The distinction is worth carrying into the reading: on the step-budget axis the ordering IS stable (0 of 10 pairs reverse), so it is the scaffold the rebuttal has to record, not the compute the agent was allowed to spend. A regime asking for one evaluation parameter should ask for that one.
United States — EO 14148 rescinded EO 14110, and the federal reporting regime sits on no cumulative-FLOP trigger.
EO 14110's 1026 Defense Production Act reporting trigger was rescinded in January 2025; there is currently no federal cumulative-FLOP reporting regime. State activity is the live front — California SB 1047, at 1026 plus $100M, was vetoed in September 2024. Whatever regime succeeds the rescinded order inherits the same capability-evidence substrate as the EU: floor-only, tolerance-undocumented.
Research community — the scaffold dimension is reported when a publisher varies it, and invisible when they do not.
Capability-frontier reporting treats the per-model published score as the unit of comparison. Where a publisher varied the scaffold and reported both columns, this page can measure what that variation does. Where a publisher reported one column — the overwhelming majority of published capability numbers — the same variation is present, unmeasured, and indistinguishable on the surface of the figure. Nothing in a citation says which case a reader has.
Scrutica's own records are strongest on the compute side. The capability side is Epoch, METR and AISI territory; the policy side is the EU AI Office and US OSTP. This page works where the three meet, and each surface below takes one slice of it.
Capacity Thresholds
Which facilities have the physical compute to put a model across the EU AI Act 1025 line. The atlas gives the short form of this argument as an addendum; the quantitative substrate is here.
The instrument anchor
The uncertainty an evaluator prints, beside the uncertainty this page recovers — the two at the same order of magnitude, on one axis.
Capacity Estimator
Per-run FLOP under three independent paths — hardware, power, cost. Multi-agent deployment needs roughly N× the budget of the single-agent evaluation that cleared it, and that cost equivalence lives here.
Derivation chains
Every Scrutica estimate, with its full derivation. The parent surface for the elicitation-gap addendum on the Thresholds view.
Per-model paired rows from Cybench (Unguided / Subtask-Guided / Subtasks) and OSWorld (15 / 50 / 100 step budgets). A row survives the paired-row filter on two conditions: every column the source publishes is populated for that entry, and the source attributes the entry to a specific model version. That gives 8 Cybench models and 5 OSWorld models. The second condition is what excludes a further seven OSWorld entries that do have all three budgets: they are open-weight agent systems the source records without a model version, and an agent system is itself a scaffold choice, so including them would vary two things at once. It does not select for the finding — two of those seven also score lower at 100 steps than at 50, the same direction as the rows shown. Scores render at the precision the source publishes, with no rounding and no interpolation.
Both leaderboards render their own rows in client-side JavaScript, so Scrutica reads them through Epoch AI's CSV ingest and stamps that ingest's vintage above. The filtered rows are printed here in full — the ladder plots all three Cybench columns for each model and states them in each model's tooltip, and the OSWorld table prints all three budgets — so every cell can be re-checked against the named leaderboard column without Scrutica in the loop.
Every unordered pair of models is classified against two columns: concordant when both order the pair the same way, inverted when they order it oppositely, tied when either column gives the two models the same score. Ties are reported separately rather than folded into either bucket — folding them into concordance would inflate agreement, and into the denominator would deflate disagreement. The rate a reader should use is 3 of 24.
The lower-bound use of a published number: the value is at least this. The disclosure panel deliberately preserves the differences between explicit floor language (OpenAI and UK AI Security Institute), qualified acknowledgement (METR and Google DeepMind) and a structural argument about evaluation generally (Anthropic and Apollo Research) rather than collapsing three framings into one concept.
Held-constant scaffolding is the right call for cross-model comparability, and the institute published the caveat itself, in plain English, on its own page. The problem sits one layer up: published cyber-capability scores get cited in policy chains as capability measurements when the publishers have labelled them floors. The calibration chain is the subject here, not the evaluation.
| Source | Publisher | Tier | Role in the argument |
|---|---|---|---|
| Our evaluation of Claude Mythos Preview’s cyber capabilities · 2026-04-13 · accessed 2026-05-20 | UK AI Security Institute | T1 | The Mythos cyber evaluation. The post reports outcomes (CTF challenges, the 32-step "Last Ones" corporate-network range); the scaffold mechanics live in Folkerts et al. 2026 below. |
| How fast is autonomous AI cyber capability advancing? · 2026-05-13 · accessed 2026-05-20 | UK AI Security Institute | T1 | AISI states in its own words that the published scores are floors. The reasons it gives are the 2.5M-token-per-task cap and the "simple agent scaffold" framing. |
| Cyber-range scaling: AISI capability evaluation methodology (Folkerts, Payne, Inman et al.) · 2026-03-11 · accessed 2026-05-20 | AISI / arXiv:2603.11214 | T1 | AISI cyber-range methodology paper: the 32-step corporate-network range, the 7-step ICS range, the 10M / 100M token budget series across seven models (Aug 2024 – Feb 2026). |
| Inspect AI framework — CHANGELOG (v0.3.213, deepagent + subagent delegation + persistent memory + structured planning) · 2026-04-27 · accessed 2026-05-20 | UK Government BEIS | T1 | Inspect’s deepagent() path was released two weeks after the Mythos report. It provides built-in research() / plan() / general() subagent factories, and the published cyber methodology has not yet adopted it. |
| Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models · 2024-08-15 · accessed 2026-05-20 | Zhang, Diffenderfer, Bhatt et al. (arXiv 2408.08926) | T2 | Cybench scaffold-mode definitions: Unguided / Subtask-Guided / Subtasks (per-step success rate). |
| Cybench HAL leaderboard (Unguided / Subtask-Guided / Subtasks columns) · rolling · accessed 2026-05-20 | Cybench / HAL evaluation harness | T2 | Per-model, per-scaffold-mode scores. The leaderboard renders its rows client-side; Scrutica reads it through Epoch AI’s CSV ingest, and the rows that ingest supplied are printed in full on this page. |
| OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments · 2024-04-11 · accessed 2026-05-20 | Xie, Zhang, Chen et al. (NeurIPS 2024; arXiv 2404.07972) | T2 | OSWorld step-budget protocol (15 / 50 / 100 steps). |
| OSWorld leaderboard (15- / 50- / 100-step budget rows) · rolling · accessed 2026-05-20 | OSWorld project | T2 | Per-model, per-step-budget scores. The leaderboard renders its rows client-side; Scrutica reads it through Epoch AI’s CSV ingest, and the rows that ingest supplied are printed in full on this page. Following the cited URL now lands on the project’s versioned leaderboard host (measured 2026-08-13). The canonical address is kept because it is what was read, and the redirect is disclosed so a reader who arrives at a version-suffixed page knows the citation anticipated it. Whether that version revises these rows is unverified; the scores here are the ones in the ingest at the accessed date. |
| Regulation (EU) 2024/1689 — Article 51(2) (presumption of systemic risk above 10²⁵ FLOP) · 2024-07-12 · accessed 2026-05-20 | European Parliament & Council | T1 | The cumulative-FLOP threshold the published capability evidence is meant to anchor. |
| Guidelines on GPAI obligations (±30% measurement tolerance) · 2025-07-18 · accessed 2026-05-20 | European Commission, AI Office | T1 | July 2025 GPAI Guidelines on cumulative-FLOP measurement. |
| Executive Order 14148 — Initial Rescissions of Harmful Executive Orders and Actions · 2025-01-20 · accessed 2026-05-20 | White House (45-W-EO-14148) | T1 | Rescinds EO 14110 (the 10²⁶ reporting trigger and the dual-use-AI reporting regime it sat on). |
| OpenAI Preparedness Framework v2 (the "lower bound, rather than a ceiling" verbatim) · 2025-04-15 · accessed 2026-05-20 | OpenAI | T1 | OpenAI: "we regard any one-time capability elicitation in a frontier model as a lower bound, rather than a ceiling, on capabilities that may emerge in real world use and misuse." The cleanest lab-side floor framing on public record. |
| Measuring AI Ability to Complete Long Software Tasks (METR) · 2025-03-18 · accessed 2026-08-14 | METR / arXiv:2503.14499v4 | T1 | METR’s elicitation disclosure. The paper treats its results as a lower bound because of the elicitation effort behind them, and puts a figure on that effort: around 2-3 engineer weeks of iterative development for the two models it singles out. The figure sizes the elicitation INPUT; no source here states what the resulting gap was. |
| Anthropic Responsible Scaling Policy v3.0 (evaluation-methodology requirements) · 2026-02-24 · accessed 2026-08-14 | Anthropic | T1 | Anthropic RSP v3.0: testing must satisfy general requirements including matching the expected efforts of potential adversaries, and its changelog states that the science of evaluations is not yet mature enough to set a confident buffer below a Capability Threshold. That is a statement about the maturity of measurement. It is not a claim that the published scores are floors. |
| Google DeepMind Frontier Safety Framework v3.0 (out-elicitation acknowledgement, ML R&D CCLs) · 2025-09-22 · accessed 2026-08-14 | Google DeepMind | T1 | Google DeepMind’s one passage on elicitation effort, in a note scoped to the Machine Learning R&D capability levels: risk assessment must allow for other actors eliciting harder than DeepMind assesses, which requires conservatism in the evaluations. The same note then says DeepMind does not expect to be out-elicited in that particular domain, so the acknowledgement has a stated limit. It is not a floor claim about DeepMind’s evaluations generally. |
| We need a science of evals (Apollo Research) · 2024-01-22 · accessed 2026-08-14 | Apollo Research | T1 | Apollo: evals "aim to estimate an upper bound of capabilities", so elicitation quality is load-bearing for what a score means. The structural argument for the floor-vs-ceiling gap; the prompt-format sensitivity figure the post gives (up to 76 accuracy points) is a cited external study’s measurement, relayed not produced. |
Citing a data pointFor a specific cell on the figure or a specific organizational disclosure, cite the underlying primary source above rather than this page. The “Cite this” control in the header offers BibTeX, APA, Chicago, MLA, CSL-JSON and DataCite forms for the analysis itself.