<!--
Source: https://scrutica.com/capabilities/elicitation-gap
Generated: 2026-09-04T01:58:10.373Z
Format: Markdown extraction of the rendered HTML at the source URL.
Registers: this page writes some passages twice, once for a technical reader and once
for a policy reader, and displays one of them. This document has the passages the
page displays; the 1 it does not display are omitted rather than
presented here as separate claims.
For the full agent guide see: https://scrutica.com/llms-full.txt
For the MCP server see: https://scrutica.com/api/mcp
-->

# Capabilities — Elicitation gap
## 3 of 24 separable model pairs change order between scaffolds

In print, and in their own published frameworks, 6 organizations have said that their capability numbers are lower bounds. OpenAI's Preparedness Framework calls one-time elicitation “a lower bound, rather than a ceiling.” METR attaches a figure — around two to three engineer-weeks of iterative development, spent on the two models it singles out — to the elicitation effort, not to the gap that effort left. The UK AI Security Institute names the specific choices that hold its numbers down. What none of them supplies is the multiplier that would convert a published score into a ceiling.

Two benchmark publishers happen to disclose more than one scaffold for the same model, which makes a piece of the gap measurable. It does not produce the missing multiplier. It produces something more useful: evidence that the multiplier cannot exist. Across the 8 paired models the fold change runs 1.56× to 4.68× — and 3 of the 24 model pairs that either column can separate are ordered one way by the first column and the other way by the last.

Version

v2.0.0 · 2026-05-20. A vintage-dated reading of primary-source disclosures, not a live feed; revised when a lab publishes a new floor statement. Benchmark ingest 2026-05-19.

Cite

### Same model, same benchmark, three scaffolds

Cybench publishes each model's score three ways: the task alone, the task with its subtask decomposition available, and scoring on the decomposed subtasks. Nothing about the model changes between the columns. Everything a reader would infer from the number does.

Parallel-coordinates plot: 8 models, each carried across the three scaffold modes Cybench publishes for the same model — Unguided, Subtask-guided, and Subtasks. Vertical position is the published score. Lines rise to the right by between 1.56 and 4.68 times, so there is no single gap factor. 3 of the 24 separable model pairs reverse their order between the outer columns, drawn as crossing lines; 4 further pairs are tied on one column and cannot be ordered. 3 models score lower under the middle scaffold than unguided.0%10%20%30%40%50%UnguidedSubtask-guidedSubtasksLlama 3 70B Instruct (Meta AI) — Unguided 5.0%, Subtask-guided 7.5%, Subtasks 8.2%. Fold change 1.64×Llama 3 70B InstructGemini 1.5 Pro (Google DeepMind) — Unguided 7.5%, Subtask-guided 5.0%, Subtasks 11.7%. Fold change 1.56×. Scores lower under the middle scaffold than unguided.Gemini 1.5 ProMixtral 8x22B Instruct (Mistral AI) — Unguided 7.5%, Subtask-guided 5.0%, Subtasks 15.2%. Fold change 2.03×. Scores lower under the middle scaffold than unguided.Mixtral 8x22B InstructLlama 3.1 405B Instruct (Meta AI) — Unguided 7.5%, Subtask-guided 15.0%, Subtasks 20.5%. Fold change 2.73×Llama 3.1 405B InstructGPT-4o (2024-11) (OpenAI) — Unguided 12.5%, Subtask-guided 17.5%, Subtasks 28.7%. Fold change 2.30×. Its order against at least one other model reverses between the outer columns.GPT-4o (2024-11)Claude 3 Opus (Anthropic) — Unguided 10.0%, Subtask-guided 12.5%, Subtasks 36.8%. Fold change 3.68×. Its order against at least one other model reverses between the outer columns.Claude 3 OpusClaude 3.5 Sonnet (2024-06) (Anthropic) — Unguided 17.5%, Subtask-guided 15.0%, Subtasks 43.9%. Fold change 2.51×. Scores lower under the middle scaffold than unguided.. Its order against at least one other model reverses between the outer columns.Claude 3.5 Sonnet (2024-06)o1-preview (OpenAI) — Unguided 10.0%, Subtask-guided 10.0%, Subtasks 46.8%. Fold change 4.68×. Its order against at least one other model reverses between the outer columns.o1-previeworder holds across the columnsorder reverses against at least one other model

Cybench per-model scores under the three scaffold modes the benchmark publishes, 8 models with all three columns populated. Vertical position is the published score on the source's own scale; the axis runs to 50% because the highest published cell is 46.8%. Slope is the fold change, crossings are order reversals, and a dip at the middle column is a model that scored lower under the nominally stronger scaffold. _Three models publish an identical 7.5% Unguided score and are drawn on one coincident mark — the source does not separate them, so neither does the figure._ Nothing here is a ceiling: the Subtasks column is the strongest scaffold the publisher chose to disclose, and the distance from it to a maximum-elicitation result is unmeasured. Source: Cybench leaderboard via the Epoch CSV ingest; scores at source-published precision, no rounding or interpolation.

1.56–4.68×The fold change from Unguided to Subtasks, across 8 models. Widest: o1-preview. Narrowest: Gemini 1.5 Pro. A single “gap factor” is wrong for almost every model in the set.

3 of 24Model pairs whose order reverses between the outer columns, of the pairs both columns separate. A further 4 of the 28 total pairs tie on one column and cannot be ordered either way.

3 of 8Models scoring lower under the middle scaffold than unguided. The scaffold ladder is not a total order.

Read the leaderboard fromUnguidedSubtask-guidedSubtasksscored on the decomposed subtasks

The same 8 models, ranked by the column you chose. The last column states each model's change in position against the Unguided ranking — the ordering a reader would get from the column closest to “the model, unassisted.” The reading you have selected is in this page’s address, so it can be cited.

#

Model

Subtasks

Fold change

Against Unguided

1

o1-preview OpenAI

46.8%

4.68×

up 3 places

2

Claude 3.5 Sonnet (2024-06) Anthropic

43.9%

2.51×

down 1 place

3

Claude 3 Opus Anthropic

36.8%

3.68×

unchanged

4

GPT-4o (2024-11) OpenAI

28.7%

2.30×

down 2 places

5

Llama 3.1 405B Instruct Meta AI

20.5%

2.73×

unchanged

6

Mixtral 8x22B Instruct Mistral AI

15.2%

2.03×

up 1 place

7

Gemini 1.5 Pro Google DeepMind

11.7%

1.56×

down 1 place

8

Llama 3 70B Instruct Meta AI

8.2%

1.64×

unchanged

### Rank reversals between scaffolds

A uniform bias is easy to reason about — if every score is low by the same factor, comparisons survive and only the absolute level moves. These are not that.

-   **Claude 3.5 Sonnet (2024-06)** outscores **o1-preview** on Unguided — 17.5% against 10.0% — and loses to it on Subtasks, 43.9% against 46.8%.
-   **GPT-4o (2024-11)** outscores **Claude 3 Opus** on Unguided — 12.5% against 10.0% — and loses to it on Subtasks, 28.7% against 36.8%.
-   **GPT-4o (2024-11)** outscores **o1-preview** on Unguided — 12.5% against 10.0% — and loses to it on Subtasks, 28.7% against 46.8%.

Both readings are correct and both are published by the same evaluator on the same page. A policy artefact that cites “the Cybench score” for either model has silently chosen one of them, and the choice is not visible in the citation. This is the part of the floor problem that a uniform-discount intuition does not cover: the elicitation gap is not a level shift, it is a reordering, and reorderings survive normalization.

Denominator

28 model pairs exist across the 8 models. 4 of them tie on one column and cannot be ordered by it, so the honest denominator for a disagreement rate is the 24 pairs both columns separate. 21 of those hold their order; 3 reverse.

The window closed

Every model above was released in 2024, and that is a fact about disclosure rather than about this page’s freshness. Of the 9 Cybench entries released up to November 2024, 8 published more than one scaffold column. Of the 13 released since — through February 2026 — none did. The measurable window did not narrow; it closed, and it closed as the models became the ones governance is about. Nothing here says the gap went away — only that the public series that let anyone measure it stopped being published.

### The same thing on a second axis: how much room the agent is given

OSWorld publishes each model at three step budgets — 15, 50 and 100 actions before the run is cut off. The budget is not a property of the model; it is a property of the evaluation.

Model

15 steps

50 steps

100 steps

100 / 15

Claude Sonnet 4.5 Anthropic

42.9

58.1

62.9

1.47×

Claude Sonnet 4 Anthropic

31.2

43.9

41.4 ↓

1.33×

Claude 3.7 Sonnet Anthropic

27.1

35.8

35.6 ↓

1.31×

Computer Use Preview OpenAI

26.0

31.3

30.5 ↓

1.17×

o3 (medium reasoning) OpenAI

9.1

17.2

23.0

2.53×

**3** of the 5 models score _lower_ at 100 steps than at 50 (marked ↓). More room is not monotonically better — an agent given a longer leash can spend it compounding an error, which is a finding about agent behaviour that the single published number for each model conceals entirely.

On the statistic this page is built on, though, the second axis is _not_ the same thing — and that is worth more than the parallel would have been. **0 of the 10 pairs reverse** between the 15 steps and 100 steps budgets, none of them tied. Budget moves levels; scaffold moves ranks. Every model gains from a longer leash and the leaderboard order survives it, where the scaffold choice reorders 3 of 24 separable pairs. The two are different kinds of evaluation choice, and only one of them threatens an argument that rests on a comparison.

And the budget can outweigh a model generation: **Claude 3.7 Sonnet** at 100 steps scores 35.6 against **Claude Sonnet 4** at 15 steps scoring 31.2 — the older model ahead of its own direct successor, same lab and same family, on the same benchmark under the same scoring, because it was given more room. It is not an isolated case: **4 of the 10** ordered older-then-newer pairs in this table cross the same way.

Every crossing, not the widest one. Ordered by lineage first — a model against its own successor is the case where nothing but the budget and the generation differs; the widest margin pairs a computer-use model with a reasoning model, which is a wider gap and a weaker comparison.

Older model, 100 steps

Score

Newer model, 15 steps

Score

Margin

Claude 3.7 Sonnet

35.6

Claude Sonnet 4

31.2

4.4

Computer Use Preview

30.5

o3 (medium reasoning)

9.1

21.4

Claude 3.7 Sonnet

35.6

o3 (medium reasoning)

9.1

26.5

Claude 3.7 Sonnet

35.6

Computer Use Preview

26.0

9.6

### The same admission, 6 times, in their own words

Verbatim from each linked primary source. The framing tags are descriptive — they say what kind of statement each is, not that the organization endorses this page's consolidation of them, which none has been asked to do.

OpenAI

Explicit floor language

> “we regard any one-time capability elicitation in a frontier model as a lower bound, rather than a ceiling, on capabilities that may emerge in real world use and misuse.”

[OpenAI Preparedness Framework v2 (the "lower bound, rather than a ceiling" verbatim)](https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf) · 2025-04-15 · §3.1 “Evaluation approach” (p. 9) — OpenAI Preparedness Framework v2. The floor is stated directly, in a single sentence.

METR

Qualified acknowledgement

> “We have put a limited amount of effort into eliciting models to get good performance on our tasks, so while our results are a reasonable lower bound, some models may have somewhat greater capabilities than we demonstrate. The most work was done to elicit o1 and the original Claude 3.5 Sonnet, each of which had around 2-3 engineer weeks of iterative development.”

[Measuring AI Ability to Complete Long Software Tasks (METR)](https://arxiv.org/abs/2503.14499v4) · 2025-03-18 · Appendix E.4 "Limitations and future work", subsection "More models with better elicitation" (v4) — METR Time Horizons paper, the most quantified disclosure on this page. The figure it quantifies is the elicitation effort spent. No figure here states the capability difference that effort produced.

UK AI Security Institute

Explicit floor language

> “The cap, alongside our use of a simple agent scaffold, artificially lowers success rates and understates what models can do with more tokens and stronger scaffolds.”

[How fast is autonomous AI cyber capability advancing?](https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing) · 2026-05-13 · Section “Cyber Time Horizons Results”, paragraph following Figure 1 — AISI ties the floor to two specific design choices it can name: the 2.5M-token-per-task budget and the simple agent scaffold. The cap (not "step cap") is a token cap.

Anthropic

Structural argument about evaluation

> “more general requirements to (a) match expected efforts of potential adversaries… Although still an aspirational goal, the science of evaluations is not currently mature enough to make confident predictions about the precise buffer we should require between current models and a Capability Threshold.”

[Anthropic Responsible Scaling Policy v3.0 (evaluation-methodology requirements)](https://www-cdn.anthropic.com/e670587677525f28df69b59e5fb4c22cc5461a17.pdf) · 2026-02-24 · Changelog (p. 17), note "Less prescriptive evaluation methodology" — Anthropic RSP v3.0, changelog note. The requirement is to match expected adversary effort, and the same note states that the buffer science is immature. This is a statement about measurement; it puts no multiplier on the gap.

Google DeepMind

Qualified acknowledgement

> “Risk assessment must take into account the fact that other actors may put significantly more effort into eliciting capabilities than we put into assessing risk, thus requiring conservatism in the form of evaluations.”

[Google DeepMind Frontier Safety Framework v3.0 (out-elicitation acknowledgement, ML R&D CCLs)](https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3.pdf) · 2025-09-22 · §1.3 Outline of Our Risk Assessment Process — “Note on Machine Learning R&D CCLs” — Frontier Safety Framework v3.0, in a note attached to the Machine Learning R&D capability levels. The note does not extend to the framework’s evaluations generally. DeepMind sets the limit itself in the next sentence: as a frontier AI company it does not expect other groups to put significantly more effort into that particular domain than it does. The scope is stated here because read loose the sentence sounds like a standard for every evaluation, and the document does not set one.

Apollo Research

Structural argument about evaluation

> “Since evals often aim to estimate an upper bound of capabilities, it is important to understand how to elicit maximal rather than average capabilities.”

[We need a science of evals (Apollo Research)](https://www.apolloresearch.ai/science/we-need-a-science-of-evals) · 2024-01-22 · Introduction, before the first section heading — Apollo Research, "We need a science of evals." The structural argument is that elicitation quality determines what an upper-bound score can mean. The post supports it with a cited external 2023 study measuring up to 76 accuracy points of difference from prompt-format changes alone; that measurement is relayed from the cited study, and is not Apollo’s own.

### The AISI case, in detail

The AISI cyber-range methodology paper (Folkerts, Payne, Inman et al., arXiv 2603.11214, 11 March 2026) defines the instrument: two ranges — a 32-step corporate-network attack and a 7-step ICS attack — under a 2.5M-token-per-task budget, across seven models from August 2024 to February 2026 at varying inference-time compute. The Mythos Preview evaluation (13 April 2026) is one run on it. On 13 May 2026 the institute named the two things holding its published numbers below the ceiling: the token cap, and a “simple agent scaffold” held constant across models so the cross-model comparison stays clean.

Inspect — the evaluation framework AISI itself maintains, at `UKGovernmentBEIS/inspect_ai` — added `deepagent()` in v0.3.213 on 27 April 2026: subagent delegation, persistent memory, structured planning. The published cyber methodology has not adopted it, and the reason is a real methodological commitment rather than an oversight. Running stronger scaffolds for later models would break the across-time comparison against every model already on the leaderboard. The institute's own reading is the cleanest one available: the held-constant call is correct for comparability, and the correct call understates capability.

### What this page does not claim

Not that the Subtasks column or the 100-step budget is a capability ceiling. Each is the strongest mode its publisher chose to disclose, and the distance from there to a genuine maximum-elicitation result is unmeasured by construction — nobody ran it, and the figures above would not see it if they had.

Not that ranking instability generalises across evaluation choices. It is a property of the scaffold axis on this evidence and not of the budget axis: 0 of 10 OSWorld pairs reverse between the 15 steps and 100 steps budgets. Two datasets, one statistic, opposite results — and five models on one benchmark is a narrow base for either half, so neither is offered as a law about evaluation.

Not a multiplier. Two measured fold-change series exist in the public literature, and neither is a gap factor: the Cybench fold change, mean 2.64× and median 2.40× across 8 models; the OSWorld 100/15 lift, mean 1.56× and median 1.33× across 5 models. They are not commensurable — each captures a different cap-versus-scaffold delta on a different instrument — and the spread within even one of them defeats the idea of a single factor. The one further figure the sources put a number on is METR's elicitation budget, and that measures the effort spent, not the distance it closed. Anyone quoting 2×, 10× or 100× across the frontier is past what the sources support.

The structural reason the number is missing matters more than the number. Maximum-elicitation of cyber or CBRN capability is operationally close to building the misuse tool, and the legal, ethical and reputational considerations that keep public evaluators from publishing one apply identically across every domain Article 51 is drawn around. The evaluation regimes nearest the ceiling — internal pre-deployment work under RSP, Preparedness or FSF review, and any classified national-security evaluation the public need not be told exists — are exactly the ones whose results the public cannot read. The publication asymmetry is why the gap is open-ended, not evidence that a specific multiplier is known in some other room.

### A floor-bound capability layer under a cumulative-FLOP ceiling

**EU AI Act — Article 51(2) presumes systemic risk above 1025 FLOP, and the rebuttal cites capability evaluations.**

The presumption is [rebuttable](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689): the provider demonstrates that the model does not pose systemic risk despite crossing the compute line, and the demonstration leans on capability evaluations. The July 2025 GPAI Guidelines codify a ±30% measurement tolerance on the cumulative-FLOP side. No comparable tolerance is documented on the capability side, because the capability side has not specified a measurement regime to put bounds on. A tolerance-bounded input feeding a floor-only output is the asymmetry this page renders.

And the reversals sharpen it. A rebuttal that argues from comparison — this model scores below that one, which was cleared — depends on an ordering that the evidence above shows is not stable across evaluation choices the citation does not record. The distinction is worth carrying into the reading: on the step-budget axis the ordering IS stable (0 of 10 pairs reverse), so it is the scaffold the rebuttal has to record, not the compute the agent was allowed to spend. A regime asking for one evaluation parameter should ask for that one.

**United States — EO 14148 rescinded EO 14110, and the federal reporting regime sits on no cumulative-FLOP trigger.**

EO 14110's 1026 Defense Production Act reporting trigger was [rescinded](https://www.federalregister.gov/documents/2025/01/28/2025-01901/initial-rescissions-of-harmful-executive-orders-and-actions) in January 2025; there is currently no federal cumulative-FLOP reporting regime. State activity is the live front — California SB 1047, at 1026 plus $100M, was vetoed in September 2024. Whatever regime succeeds the rescinded order inherits the same capability-evidence substrate as the EU: floor-only, tolerance-undocumented.

**Research community — the scaffold dimension is reported when a publisher varies it, and invisible when they do not.**

Capability-frontier reporting treats the per-model published score as the unit of comparison. Where a publisher varied the scaffold and reported both columns, this page can measure what that variation does. Where a publisher reported one column — the overwhelming majority of published capability numbers — the same variation is present, unmeasured, and indistinguishable on the surface of the figure. Nothing in a citation says which case a reader has.

### Where this page fits on Scrutica

Scrutica's own records are strongest on the compute side. The capability side is Epoch, METR and AISI territory; the policy side is the EU AI Office and US OSTP. This page works where the three meet, and each surface below takes one slice of it.

[

Capacity · Thresholds

Capacity Thresholds

Which facilities have the physical compute to put a model across the EU AI Act 1025 line. The atlas gives the short form of this argument as an addendum; the quantitative substrate is here.

](/capacity/thresholds)[

Capabilities

The instrument anchor

The uncertainty an evaluator prints, beside the uncertainty this page recovers — the two at the same order of magnitude, on one axis.

](/capabilities)[

Capacity · Engine

Capacity Estimator

Per-run FLOP under three independent paths — hardware, power, cost. Multi-agent deployment needs roughly N× the budget of the single-agent evaluation that cleared it, and that cost equivalence lives here.

](/capacity)[

Methodology

Derivation chains

Every Scrutica estimate, with its full derivation. The parent surface for the elicitation-gap addendum on the Thresholds view.

](/methodology)

### Method, and every source

#### What the figure shows

Per-model paired rows from Cybench (Unguided / Subtask-Guided / Subtasks) and OSWorld (15 / 50 / 100 step budgets). A row survives the paired-row filter on two conditions: every column the source publishes is populated for that entry, _and_ the source attributes the entry to a specific model version. That gives **8** Cybench models and **5** OSWorld models. The second condition is what excludes a further seven OSWorld entries that do have all three budgets: they are open-weight agent _systems_ the source records without a model version, and an agent system is itself a scaffold choice, so including them would vary two things at once. It does not select for the finding — two of those seven also score lower at 100 steps than at 50, the same direction as the rows shown. Scores render at the precision the source publishes, with no rounding and no interpolation.

Both leaderboards render their own rows in client-side JavaScript, so Scrutica reads them through Epoch AI's CSV ingest and stamps that ingest's vintage above. The filtered rows are printed here in full — the ladder plots all three Cybench columns for each model and states them in each model's tooltip, and the OSWorld table prints all three budgets — so every cell can be re-checked against the named leaderboard column without Scrutica in the loop.

#### How the pair counts are taken

Every unordered pair of models is classified against two columns: concordant when both order the pair the same way, inverted when they order it oppositely, tied when either column gives the two models the same score. Ties are reported separately rather than folded into either bucket — folding them into concordance would inflate agreement, and into the denominator would deflate disagreement. The rate a reader should use is **3** of **24**.

#### What “floor” means here

The lower-bound use of a published number: the value is at least this. The disclosure panel deliberately preserves the differences between explicit floor language (OpenAI and UK AI Security Institute), qualified acknowledgement (METR and Google DeepMind) and a structural argument about evaluation generally (Anthropic and Apollo Research) rather than collapsing three framings into one concept.

#### Why this is not a critique of AISI

Held-constant scaffolding is the right call for cross-model comparability, and the institute published the caveat itself, in plain English, on its own page. The problem sits one layer up: published cyber-capability scores get cited in policy chains as capability measurements when the publishers have labelled them floors. The calibration chain is the subject here, not the evaluation.

#### Primary sources

Source

Publisher

Tier

Role in the argument

[Our evaluation of Claude Mythos Preview’s cyber capabilities](https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities) · 2026-04-13 · accessed 2026-05-20

UK AI Security Institute

T1

The Mythos cyber evaluation. The post reports outcomes (CTF challenges, the 32-step "Last Ones" corporate-network range); the scaffold mechanics live in Folkerts et al. 2026 below.

[How fast is autonomous AI cyber capability advancing?](https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing) · 2026-05-13 · accessed 2026-05-20

UK AI Security Institute

T1

AISI states in its own words that the published scores are floors. The reasons it gives are the 2.5M-token-per-task cap and the "simple agent scaffold" framing.

[Cyber-range scaling: AISI capability evaluation methodology (Folkerts, Payne, Inman et al.)](https://arxiv.org/abs/2603.11214) · 2026-03-11 · accessed 2026-05-20

AISI / arXiv:2603.11214

T1

AISI cyber-range methodology paper: the 32-step corporate-network range, the 7-step ICS range, the 10M / 100M token budget series across seven models (Aug 2024 – Feb 2026).

[Inspect AI framework — CHANGELOG (v0.3.213, deepagent + subagent delegation + persistent memory + structured planning)](https://github.com/UKGovernmentBEIS/inspect_ai/blob/main/CHANGELOG.md) · 2026-04-27 · accessed 2026-05-20

UK Government BEIS

T1

Inspect’s deepagent() path was released two weeks after the Mythos report. It provides built-in research() / plan() / general() subagent factories, and the published cyber methodology has not yet adopted it.

[Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models](https://arxiv.org/abs/2408.08926) · 2024-08-15 · accessed 2026-05-20

Zhang, Diffenderfer, Bhatt et al. (arXiv 2408.08926)

T2

Cybench scaffold-mode definitions: Unguided / Subtask-Guided / Subtasks (per-step success rate).

[Cybench HAL leaderboard (Unguided / Subtask-Guided / Subtasks columns)](https://cybench.github.io/) · rolling · accessed 2026-05-20

Cybench / HAL evaluation harness

T2

Per-model, per-scaffold-mode scores. The leaderboard renders its rows client-side; Scrutica reads it through Epoch AI’s CSV ingest, and the rows that ingest supplied are printed in full on this page.

[OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments](https://arxiv.org/abs/2404.07972) · 2024-04-11 · accessed 2026-05-20

Xie, Zhang, Chen et al. (NeurIPS 2024; arXiv 2404.07972)

T2

OSWorld step-budget protocol (15 / 50 / 100 steps).

[OSWorld leaderboard (15- / 50- / 100-step budget rows)](https://os-world.github.io/) · rolling · accessed 2026-05-20

OSWorld project

T2

Per-model, per-step-budget scores. The leaderboard renders its rows client-side; Scrutica reads it through Epoch AI’s CSV ingest, and the rows that ingest supplied are printed in full on this page. Following the cited URL now lands on the project’s versioned leaderboard host (measured 2026-08-13). The canonical address is kept because it is what was read, and the redirect is disclosed so a reader who arrives at a version-suffixed page knows the citation anticipated it. Whether that version revises these rows is unverified; the scores here are the ones in the ingest at the accessed date.

[Regulation (EU) 2024/1689 — Article 51(2) (presumption of systemic risk above 10²⁵ FLOP)](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689) · 2024-07-12 · accessed 2026-05-20

European Parliament & Council

T1

The cumulative-FLOP threshold the published capability evidence is meant to anchor.

[Guidelines on GPAI obligations (±30% measurement tolerance)](https://digital-strategy.ec.europa.eu/en/policies/guidelines-gpai-providers) · 2025-07-18 · accessed 2026-05-20

European Commission, AI Office

T1

July 2025 GPAI Guidelines on cumulative-FLOP measurement.

[Executive Order 14148 — Initial Rescissions of Harmful Executive Orders and Actions](https://www.federalregister.gov/documents/2025/01/28/2025-01901/initial-rescissions-of-harmful-executive-orders-and-actions) · 2025-01-20 · accessed 2026-05-20

White House (45-W-EO-14148)

T1

Rescinds EO 14110 (the 10²⁶ reporting trigger and the dual-use-AI reporting regime it sat on).

[OpenAI Preparedness Framework v2 (the "lower bound, rather than a ceiling" verbatim)](https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf) · 2025-04-15 · accessed 2026-05-20

OpenAI

T1

OpenAI: "we regard any one-time capability elicitation in a frontier model as a lower bound, rather than a ceiling, on capabilities that may emerge in real world use and misuse." The cleanest lab-side floor framing on public record.

[Measuring AI Ability to Complete Long Software Tasks (METR)](https://arxiv.org/abs/2503.14499v4) · 2025-03-18 · accessed 2026-08-14

METR / arXiv:2503.14499v4

T1

METR’s elicitation disclosure. The paper treats its results as a lower bound because of the elicitation effort behind them, and puts a figure on that effort: around 2-3 engineer weeks of iterative development for the two models it singles out. The figure sizes the elicitation INPUT; no source here states what the resulting gap was.

[Anthropic Responsible Scaling Policy v3.0 (evaluation-methodology requirements)](https://www-cdn.anthropic.com/e670587677525f28df69b59e5fb4c22cc5461a17.pdf) · 2026-02-24 · accessed 2026-08-14

Anthropic

T1

Anthropic RSP v3.0: testing must satisfy general requirements including matching the expected efforts of potential adversaries, and its changelog states that the science of evaluations is not yet mature enough to set a confident buffer below a Capability Threshold. That is a statement about the maturity of measurement. It is not a claim that the published scores are floors.

[Google DeepMind Frontier Safety Framework v3.0 (out-elicitation acknowledgement, ML R&D CCLs)](https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3.pdf) · 2025-09-22 · accessed 2026-08-14

Google DeepMind

T1

Google DeepMind’s one passage on elicitation effort, in a note scoped to the Machine Learning R&D capability levels: risk assessment must allow for other actors eliciting harder than DeepMind assesses, which requires conservatism in the evaluations. The same note then says DeepMind does not expect to be out-elicited in that particular domain, so the acknowledgement has a stated limit. It is not a floor claim about DeepMind’s evaluations generally.

[We need a science of evals (Apollo Research)](https://www.apolloresearch.ai/science/we-need-a-science-of-evals) · 2024-01-22 · accessed 2026-08-14

Apollo Research

T1

Apollo: evals "aim to estimate an upper bound of capabilities", so elicitation quality is load-bearing for what a score means. The structural argument for the floor-vs-ceiling gap; the prompt-format sensitivity figure the post gives (up to 76 accuracy points) is a cited external study’s measurement, relayed not produced.

Citing a data pointFor a specific cell on the figure or a specific organizational disclosure, cite the underlying primary source above rather than this page. The “Cite this” control in the header offers BibTeX, APA, Chicago, MLA, CSL-JSON and DataCite forms for the analysis itself.