Capabilities
What frontier models can do, on what compute, per what evidence — and what the public scores hide. Benchmark performance plotted against training compute; autonomous-task time horizons against the statutory lines they approach; announced training runs cross-checked against the physical evidence; and the elicitation gap that makes every published score a floor rather than a ceiling. Start with the scaling curve, then follow a model into the compute that could train it, or down into the floor its score really represents.
Data vintages