Reproducible protocol for measuring country-level attention to artificial intelligence in Policy Commons documents.
Historical release v7.6: 269,853 master records; 182,018 analysis document-country rows; 135 sovereign-country codes. These are historical counts, not the current corpus size. The 24 September 2026 amendment below supersedes earlier blanket media exclusions and the prohibition on removing confirmed false positives from active corpora.
Videos, audio recordings and blogs are excluded by default. A video or recording may be retained as a labelled media exception when its abstract substantively describes AI policy and source metadata explicitly states a country. A blog may be retained when its title, abstract or keywords explicitly establishes AI-policy relevance and source metadata states a country. A country mention, an AI-adjacent sector or a generic event topic list alone does not satisfy this exception. Keep the format and exception basis available for sensitivity analysis.
Match AI as a complete token, including punctuation-delimited compounds such as AI-powered. Letters inside airport, airplane, availability, obtaining or explaining are not AI evidence. Recognised full phrases and multilingual terminology remain valid. Metadata-only screening identifies candidates for review; absence of AI terminology in metadata is not proof of irrelevance because indexed full text may contain substantive AI content. Record whether a decision is based on metadata, original text or human adjudication.
Source-reported publication country, inferred policy target country and organisation headquarters country are distinct attributes. Where publication-country evidence is missing, explicit national policy scope may support up to three target-country assignments, labelled as inference. International organisations and coordination bodies retain their organisational category; headquarters location does not automatically assign their documents to a national corpus. CGMS is an international coordination body, with its secretariat location in Germany (DEU) stored separately.
Retain substantive sectoral AI policy, including rural healthcare, AI-powered clinical tools, cybersecurity, arts and culture. Record these supported topics separately. Precision medicine alone is not AI-policy evidence. Creative Australia's report is eligible because it explicitly discusses AI and Australian arts participation. The RHTP case is evaluated as rural-healthcare implementation material, including AI clinical tools and cybersecurity, rather than relabelled solely as cybersecurity.
User-confirmed exact false positives are removed by stable identifier from active collected, deduplicated and analytical corpora. Maintain a deletion ledger with the identifier, decision, evidence, rule version and affected files. Immutable archival evidence may remain outside the active processing pipeline; a tombstone prevents later re-ingestion. Format exclusions and unresolved cases retain explicit disposition fields. No new run counts or corpus-wide full-text validation are asserted by this methodological amendment.
A substantive official transcript or recap may establish the media exception when the collected abstract is truncated; preserve the original abstract and store the enrichment and its source separately. A genre word occurring in a keyword field alone is insufficient to establish that a record is a video or blog. Confirm format through source metadata, source URL or the original page. A metadata census is not full-text verification.
Specialist review outcomes: all 238 initial exclusion candidates were reviewed, producing 179 provisional exclusions and 59 holds. Of 22 exception candidates reviewed, 13 were approved and 9 were held. These are candidate-review outcomes, not a statement that updates to the active corpus masters are complete. Holds require additional evidence.
Applied metadata-screen results: 730,058 deduplicated documents were screened. This round excluded 181 records from analysis (179 provisional media/blog exclusions and two human decisions), retained 13 specialist-reviewed media exceptions plus the human-reviewed RHTP exception, and held 68 media candidates pending additional evidence. Two human-confirmed false positives were deleted from each active master; a previously deleted third identifier was also purged from collection inputs. The current masters contain 730,175 source-stage rows, 730,056 deduplicated documents and 730,068 country-expanded rows. These are storage counts, not counts of eligible national-policy documents. The broader metadata review queue contains 28,825 records, including 27,747 substring-only candidates; these are not automatically classified as irrelevant.
Detailed examples and evidence links are in the dated error-register amendment.
(record_id, analysis_country) rows with unit weight 1.Include government departments, public authorities, local government, publicly funded research institutes, government-funded research centres and private research units carrying out commissioned national research. A company headquarters or branch location alone is insufficient.
Exclude ordinary business, product and promotional documents from clearly multinational private firms. Preserve them in the source master and retain government-research exceptions with evidence. EU, international and regional organisations are non-national unless a national office, commission, policy jurisdiction or country-specific mandate is demonstrated.
Store published_in_country, issuer_country, country_scope_class, government_research_evidence and country_decision_source separately. `Published in` is provider evidence, not an automatic issuer country.
Human-review rule: retain official proceedings, workshops and agendas when they document a government inquiry, regulation, standards work, public administration, testimony, recommendation or implementation. Exclude only logistical event records and purely scholarly proceedings.
Full cited examples are in the Policy Commons contamination and country-attribution criteria.
| Release | Rows | Change |
|---|---|---|
| v7.3 | 182,385 | HKG aggregated to CHN. |
| v7.4 | 182,347 | Additional agendas, proceedings and non-national networks excluded. |
| v7.5 | 182,018 | Exact multinational/private registry, remaining proceedings and product-manual misattributions applied. |
| v7.6 | 182,018 | National policy-interest scope fields added; 17 ambiguous non-national cases retained for manual review. |
Seed 20260922: 100/100 pages opened, country values matched after the Kosovo issuer exception. Seed 20260923: 100/100 pages opened after one retry; the sample was non-overlapping with prior samples. AI-term absence was never treated as exclusion evidence.
For every release, preserve the input, rules, seeds, manifests, evidence, exclusion queue, checksums and executable code.
Records whose normalised title is exactly Unidentified Document Analysis are retained as source evidence but excluded from content analysis, even if a provider country field is populated. Eight such records were found in the 730,056-document deduplicated master, seven with a country code. A country field cannot rescue unintelligible content.
A missing or punctuation-only abstract, blank publication country and absent substantive keywords place a record in original-source recovery. There were 63 blank-country records with unusable abstracts; 62 are held for source recovery, while one human-reviewed blog was excluded and retained. A generic WEB bibliographic type alone is not evidence that every such item is a blog. Metadata sparsity does not prove a document is non-AI.
Among 4,429 rows with blank country in the deduplicated master, exact publisher-name linkage yielded 616 unique Policy Commons organisation IDs. Another 873 rows had multiple IDs with an identical organisation type: store only the type consensus. Thirty-six rows had conflicting types and remain ambiguous. None of these links assigns publication country automatically. A separate 449-record queue has a long abstract and explicit AI-policy terms but no national country; it requires original-source checks.
The Digital Trust Council report is retained in the private-nonprofit auxiliary corpus. The original 19-page report establishes substantive AI certification content. The Policy Commons organisation entry classifies its issuer as Nonprofit; its official website reports 501(c)(3) registration. Record US legal registration separately and leave publication country unresolved. The current country-review HTML contains 4,353 rows after the two newly excluded blank-country cases were removed from that worklist.