# Scrutica Sovereign AI Program Index — Data Dictionary

Column-level reference for `sovereign-ai-index_v2026.3.csv` — version 2026.3, released 2026-07-17, 34 rows, CC-BY-4.0.

Sovereign-AI investment programs with announced / committed / disbursed spend stages and qualitative interdependence ratings (NVIDIA / TSMC / U.S. / BIS reach).

Generated from the dataset's JSON Schema (`sovereign-ai-index_v2026.3.schema.json`); this Markdown is the human-readable view of that machine-readable contract.

## Conventions

- **Nullability** — every column is optional unless its **Required** cell says `Yes`. A null is written as an empty CSV cell.
- **Lists** — multi-value columns are semicolon-joined (`a; b; c`).
- **Currency** — monetary columns are USD unless a `currency_of_record` column on the same row states otherwise.
- **Dates** — ISO 8601 (`YYYY-MM-DD`).
- **Estimates** — columns suffixed `_is_estimated` flag analyst-derived values (`true`) versus primary-source values (`false` / null), per Scrutica's transparency convention: no derived number is presented as if measured.
- **Provenance** — every row carries `data_source` and, where the source is public, `source_url`, so any value traces back to a primary source.

## Columns (28)

| # | Column | Type | Required | Description |
|--:|--------|------|:--------:|-------------|
| 1 | `country_code` | string | Yes | ISO 3166-1 alpha-2 lowercase country code |
| 2 | `program_name` | string | Yes | Free-text program name as Scrutica indexes it (combines stacked sub-programs where attribution is shared) |
| 3 | `program_status` | string |  | Free-text status (announced, committed, operational, legislation_pending, etc.) |
| 4 | `timeline_start_year` | number |  | Year the program began (per primary source) |
| 5 | `timeline_target_year` | number |  | Year by which the announced capacity / spend is targeted to be in place |
| 6 | `announced_usd` | number |  | Total announced spend in USD, summed across announced sub-programs (inflation note: large pledges are commonly stacked across overlapping initiatives) |
| 7 | `announced_usd_is_estimated` | boolean |  | True when announced_usd combines analyst-assigned conversions or aggregations |
| 8 | `announced_govt_only_usd` | number |  | Subset of announced_usd attributable to government commitments only (excludes private capex pledges) |
| 9 | `committed_usd` | number |  | USD subset of announced spend that has cleared legislative / board commitment, where verifiable |
| 10 | `committed_usd_is_estimated` | boolean |  | True when committed_usd is analyst-assigned |
| 11 | `disbursed_usd` | number |  | USD subset of committed spend that has been actually disbursed / contracted, where verifiable |
| 12 | `disbursed_usd_is_estimated` | boolean |  | True when disbursed_usd is analyst-assigned |
| 13 | `nvidia_dependency` | string |  | Qualitative dependency on NVIDIA accelerators (low / medium / high) |
| 14 | `tsmc_dependency` | string |  | Qualitative dependency on TSMC foundry (low / medium / high) |
| 15 | `us_dependency` | string |  | Qualitative dependency on U.S.-jurisdiction inputs (low / medium / high) |
| 16 | `bis_jurisdiction_reach` | string |  | Whether the program sits inside BIS Entity-List / FDPR jurisdictional reach (low / medium / high) |
| 17 | `primary_chips` | string |  | Semicolon-joined list of primary accelerators identified on the program (model strings as published) |
| 18 | `key_partners` | string |  | Semicolon-joined list of partner organisations |
| 19 | `structural_non_disclosure` | boolean |  | True when material capacity / spend is not disclosed by structure (e.g., classified, military-aligned) |
| 20 | `structural_non_disclosure_reason` | string |  | Free-text rationale for structural_non_disclosure |
| 21 | `currency_of_record` | string |  | Native reporting currency (e.g., USD, EUR, CNY) of the announcement |
| 22 | `fx_rate_used` | number |  | FX rate applied to convert currency_of_record into USD |
| 23 | `fx_rate_source` | string |  | Source of fx_rate_used |
| 24 | `fx_rate_date` | string |  | ISO 8601 date of fx_rate_used |
| 25 | `source_count` | number |  | Number of independent primary sources backing the row |
| 26 | `data_source` | string |  | Primary source attribution |
| 27 | `source_url` | string |  | Public URL of the primary source |
| 28 | `key_notes` | string |  | Free-text analyst note |

## Integrity & citation

- **SHA-256** of `sovereign-ai-index_v2026.3.csv`: `19a19e3f91fad40a2978b5b733f338eaa13f774f40906be003bf8133868dec56`
- **Canonical page** (BibTeX · APA · DOI once minted): https://scrutica.com/datasets/sovereign-ai-index
- **Methodology**: https://scrutica.com/datasets/sovereign-ai-index → Methodology, and https://scrutica.com/methodology

_Scrutica · sovereign-ai-index · v2026.3 · 2026-07-17 · CC-BY-4.0_
