<!--
Source: https://scrutica.com/capabilities/training-evidence
Generated: 2026-07-27T06:19:33.981Z
Format: Markdown extraction of the rendered HTML at the source URL.
For the full agent guide see: https://scrutica.com/llms-full.txt
For the MCP server see: https://scrutica.com/api/mcp
-->

# Training Claim Evidence Dashboard
Scrutica

# Capabilities

What frontier models can do, on what compute, per what evidence — and what the public scores hide. Benchmark performance plotted against training compute; autonomous-task time horizons against the statutory lines they approach; announced training runs cross-checked against the physical evidence; and the elicitation gap that makes every published score a floor rather than a ceiling. Start with the scaling curve, then follow a model into the compute that could train it, or down into the floor its score really represents.

Data vintagesScaling6dElicitation68dstaticIn Context148dstatic