<!--
Source: https://scrutica.com/datasets
Generated: 2026-07-28T10:25:11.434Z
Format: Markdown extraction of the rendered HTML at the source URL.
For the full agent guide see: https://scrutica.com/llms-full.txt
For the MCP server see: https://scrutica.com/api/mcp
-->

# Datasets
Scrutica

# Data & Citation

The packaged Scrutica datasets and how to cite them. Each bundle ships as a CSV with a paired JSON Schema, methodology Markdown, SHA-256 checksum, and Croissant JSON-LD descriptor; the Citation view carries the reuse forms for the datasets and for any number rendered on the site.

## Packaged datasets

Two Scrutica datasets, each shipped as a CSV with a paired JSON Schema, methodology Markdown, SHA-256 checksum, and Croissant JSON-LD descriptor (MLCommons 1.0). Licensed CC-BY-4.0; mirrored on Zenodo for DOI minting and on Hugging Face once the depositions land.

The supply-chain interdependence graph is not among them: the full graph behind [/supply-chain](/supply-chain) and [/cascade](/cascade) is licensed-substrate-majority and stays on-site, and its public-provenance subset draws chiefly on CSET ETO chip-explorer data (CC BY-NonCommercial 4.0), which the site-wide CC-BY on these bundles cannot carry — so those edge-level analyses are published on-site rather than packaged for download.

[

## Sovereign AI Program Index v2026.3

Sovereign-AI investment programs with announced / committed / disbursed spend stages and qualitative interdependence ratings.

34 rows·CC-BY-4.0·Released 2026-07-17·SHA-256 19a19e3f91…·DOI pending Zenodo deposition



](/datasets/sovereign-ai-index)[

## Compute Cost Index v2026.3

Cloud GPU pricing normalized to $/petaFLOP-day across on-demand, spot, and reserved tiers, using dense BF16 throughput.

508 rows·CC-BY-4.0·Released 2026-07-17·SHA-256 9d4ad13062…·DOI pending Zenodo deposition



](/datasets/compute-cost-index)

What is in each bundle

Each bundle is five files: a CSV (UTF-8, no BOM, ISO 8601 dates); a JSON Schema typing the columns; methodology Markdown that names sources and known limitations; a single-line SHA-256 over the CSV; and the Croissant JSON-LD descriptor (MLCommons 1.0) so ML frameworks can load the dataset without a manual schema mapping. Zenodo depositions mint DataCite DOIs that propagate onward to OpenAlex and Semantic Scholar, and the Hugging Face mirror ships in parallel — auto-generating its own Croissant on that side. CC-BY-4.0 throughout.

Refresh cadence

Bundles are regenerated by the platform’s packaging pipeline from the same canonical substrate every page on this site reads. Versioning is calendar-based (year.release-within-year), and fields are additive across versions; each refresh mints a fresh Zenodo version DOI under the same concept DOI, so a citation made against a prior release still resolves to that release.

Notify me about

New dataMajor featuresRegulatoryResearch

Email addressSubscribe

[CC BY-SA 4.0·© 2026 Scrutica](https://creativecommons.org/licenses/by-sa/4.0/ "Creative Commons Attribution-ShareAlike 4.0")

Built for the AI governance research community