Skip to main content

Data & Citation

The packaged Scrutica datasets and how to cite them. Each bundle ships as a CSV with a paired JSON Schema, methodology Markdown, SHA-256 checksum, and Croissant JSON-LD descriptor; the Citation view carries the reuse forms for the datasets and for any number rendered on the site.

Packaged datasets

Two Scrutica datasets, each shipped as a CSV with a paired JSON Schema, methodology Markdown, SHA-256 checksum, and Croissant JSON-LD descriptor (MLCommons 1.0). Licensed CC-BY-4.0; mirrored on Zenodo for DOI minting and on Hugging Face once the depositions land.

The supply-chain interdependence graph is not among them: the full graph behind /supply-chain and /cascade is licensed-substrate-majority and stays on-site, and its public-provenance subset draws chiefly on CSET ETO chip-explorer data (CC BY-NonCommercial 4.0), which the site-wide CC-BY on these bundles cannot carry — so those edge-level analyses are published on-site rather than packaged for download.

What is in each bundle

Each bundle is five files: a CSV (UTF-8, no BOM, ISO 8601 dates); a JSON Schema typing the columns; methodology Markdown that names sources and known limitations; a single-line SHA-256 over the CSV; and the Croissant JSON-LD descriptor (MLCommons 1.0) so ML frameworks can load the dataset without a manual schema mapping. Zenodo depositions mint DataCite DOIs that propagate onward to OpenAlex and Semantic Scholar, and the Hugging Face mirror ships in parallel — auto-generating its own Croissant on that side. CC-BY-4.0 throughout.

Refresh cadence

Bundles are regenerated by the platform’s packaging pipeline from the same canonical substrate every page on this site reads. Versioning is calendar-based (year.release-within-year), and fields are additive across versions; each refresh mints a fresh Zenodo version DOI under the same concept DOI, so a citation made against a prior release still resolves to that release.