Reviewed workflow · updated 23 September 2026

Policy Commons collection and Korean policy-report integration

Two sequential integration steps with real directory names, numbered tasks, source evidence, release checks and current status. Step 1 integration and strict quarantine are complete. Step 2 covers Korean issuers only.

Current status and units

Step 1 has been completed as an evidence-preserving integration of the frozen v8.3 candidate document master and the newly collected English and multilingual Policy Commons RIS/CSV master. Exact stable-pair resolution produced 738,737 unique records. A subsequent strict contamination review quarantined 298,880 records and issued a primary strict-analysis candidate set of 439,857 records. The protected v7.10 document-country file remains unchanged as a comparison baseline.

270,974Step 1 Source A: frozen v8.3 candidate document records
559,181Step 1 Source B: reviewed multilingual RIS/CSV document records
1,417,517unique source observations supporting Source B
182,010protected v7.10 document-country rows for comparison
738,737Step 1 integrated evidence-preserving document records
439,857strict primary-analysis candidate records after quarantine

Count rule: 270,974 + 559,181 − 50 internal duplicate rows − 91,368 exact cross-dataset overlaps = 738,737 integrated records. The 1,417,517 count is a Source B observation count and the 182,010 count is the protected document-country baseline, so neither is added to the document total. The strict quarantine partitions 738,737 records into Q1 210,153, Q2 82,380, Q3 6,347 and 439,857 analysis candidates.

Protected v7.10 baseline→Step 1 Policy Commons integration→Step 2 Korean integration→Reviewed comparison and promotion

Verified upper-level directories

iCloud root: /Users/deep1003/Library/Mobile Documents/com~apple~CloudDocs/data3/

The five collector modules
  1. 01_world_governments__POCO_POGO/ · World Governments
  2. 02_world_cities__POCO_PCWL/ · World Cities
  3. 03_public_health_and_social_care__POCO_POHE/ · Public Health and Social Care
  4. 04_north_american_state_provincial_territorial_governments__POCO_PSPL/ · North American State, Provincial and Territorial Governments
  5. 05_global_think_tanks__POCO_ALL/ · Global Think Tanks

Step 1: integrate the current Policy Commons collection

Computational integration complete Strict quarantine verified Analytical promotion remains review-gated

Merge inputs: Source A is the frozen v8.3 candidate master copied from Downloads. Source B is the reviewed English and multilingual RIS/CSV integration produced from the collector. Target workspace: policycommons_merge_workspace_20260921/03_stage1_current_master_plus_collector/.

RoleDataset and exact pathVerified scaleBoundary
Source A
local baseline input
policy_commons_ai_full_collection_20260910/imports_from_downloads_20260921/04_legacy_releases_and_checkpoints/releases/policycommons_local_enrichment_20260920_r1/final_augmented_release_v8_3_20260921/policycommons_augmented_master_frozen_v8_3.parquet270,974 candidate document records; 269,853 base-master records before the reviewed incrementCandidate document master, not the protected v7.10 analytical release
Source B
multilingual target input
policy_commons_ai_full_collection_20260910/output/policycommons_ai_ris_csv_integrated_20260922_v3/integrated_master.parquet559,181 unique documents: 545,717 paired RIS/CSV, 11,866 RIS-only and 1,598 CSV-only; 1,417,517 unique source observationsEnglish and multilingual collection after specialist metadata review; not yet merged with Source A
Comparison baseline
not a merge input
policycommons_golden_datasets_20260921/02_authoritative_analysis_v7_10/document_country_pairs_final_v7_10.csv182,010 document-country rowsProtected analytical baseline used only for comparison and reviewed promotion
Step 1 integrated outputstep1_v83_plus_collector_20260923_v6_reviewed/738,737 unique documents; 830,155 crosswalk rows; 1,473,871 linked source observationsEvidence-preserving master; no source record overwritten
Strict quarantine outputstep1_v83_plus_collector_20260923_v7_strict_quarantine/210,153 Q1; 82,380 Q2; 6,347 Q3; 439,857 strict analysis candidatesOnly strict_analysis_use_allowed=true records enter primary analysis

Validated count: the 830,155 pre-deduplication rows resolved to 738,737 documents after removing 50 internal duplicate rows and 91,368 exact cross-dataset overlaps. Two same-COI/different-artifact groups, DOI/title collision candidates and 16 provenance orphans remain in explicit review or recovery ledgers rather than being silently merged.

Task 1. Secure and reconcile files

Real directories: /Users/deep1003/Downloads/; policy_commons_ai_full_collection_20260910/imports_from_downloads_20260921/{00_inventory,05_transfer_audit}/; policycommons_merge_workspace_20260921/03_stage1_current_master_plus_collector/00_frozen_inputs/.

  1. 1.1 Reconcile each frozen Downloads entry with its recorded iCloud destination by byte size and SHA-256. Record missing, changed or unreadable files individually.
  2. 1.2 Verify remote iCloud upload and restore a test batch. Local copy hashes alone do not prove remote durability.
  3. 1.3 Freeze a dated collector snapshot without reading partially written files. Record completeness by module and year, including incomplete slices.
  4. 1.4 Preserve manifests and checksums; check symlinks, scripts and active checkpoint paths. Prepare a source deletion inventory only after the transfer checks pass.

Transfer limitation The reconciliation ledger contains 6,072 inventory entries: 6,057 local source-to-destination SHA-256 matches, 14 file-access cases requiring review and one non-regular item. This proves local copy agreement for the matched files, not independent remote iCloud durability.

Task 2. Normalise RIS and CSV observations

Real directories: policy_commons_ai_full_collection_20260910/artificial_intelligence_1950_2026_by_module/; .../imports_from_downloads_20260921/02_upload_batches/; policycommons_merge_workspace_20260921/03_stage1_current_master_plus_collector/01_normalised_observations/.

  1. 2.1 Parse RIS and CSV separately, retaining raw fields and empty values.
  2. 2.2 Record verified file pairs separately from record-level matches. Keep RIS-only, CSV-only and conflicting cases.
  3. 2.3 Produce a source-observation table with identifier, query, module, year, page, export batch, original file path and checksum.

Task 3. Link and deduplicate document versions

Real directories: policycommons_merge_workspace_20260921/02_shared_crosswalks/identity_keys/ and .../03_stage1_current_master_plus_collector/02_identity_and_conflicts/.

  1. 3.1 Link namespaced Policy Commons IDs or COIs first, then canonical document URLs.
  2. 3.2 Review DOI, title, year and issuer collisions before merging; preserve different editions and translations.
  3. 3.3 Build a source-to-document crosswalk, conflict queue and reconciliation of input occurrences to canonical document versions.

Task 4. Validate and issue Step 1 outputs

Real directories: policycommons_merge_workspace_20260921/03_stage1_current_master_plus_collector/{03_quality_assurance,04_versioned_release}/; .../05_final_analysis_releases/; .../06_comparative_evaluation/baseline_vs_stage1/.

  1. 4.1 Build a versioned document master and document-country table with distinct declared units.
  2. 4.2 Identify national analytical rows and separately retain international organisations and multinational private issuers.
  3. 4.3 Check schema, uniqueness, counts, provenance, conflicts, checksums and source-review samples.
  4. 4.4 Compare the issued Step 1 release with the protected v7.10 baseline.

Verified release issued The v6 integration and v7 strict-quarantine releases contain manifests, SHA-256 checksums, identity and field-reconciliation ledgers, country partitions, quarantine indexes and independent verification reports. The v7 partition check passed: all 738,737 records are unique within and mutually exclusive across Q1, Q2, Q3 and the strict analysis set.

Strict partitionRecordsPrimary-analysis treatment
Q1 high-confidence quarantine210,153Excluded. Includes explicit non-AI/non-document records, simple agendas and minutes, calls and solicitations, event material, newsletters, simple press releases, recruitment notices and explicit non-document media.
Q2 precautionary review82,380Excluded pending adjudication. Includes ambiguous event material, identity candidates, repeated abstracts, title-summary risk and substantive cross-source conflicts.
Q3 metadata recovery6,347Excluded pending recovery and reclassification.
Strict analysis candidates439,857Permitted for primary analysis under the strict release rule.

Step 2: add earlier Korean policy-report datasets

Korean candidate files exist Integration pending

Source: kim_jongrip_ai_stpi_full_datasets_share_20260630/. Prepared inputs: policy_commons_ai_full_collection_20260910/stage2_stpi_korean_policy_reports_20260921/. Workspace: policycommons_merge_workspace_20260921/04_stage2_stpi_kor_chn_jpn/. The final directory has a legacy name; this step covers Korea only.

Korean fileRowsCurrent meaning
Approved input7,749Protected manuscript baseline, 144 original columns
Recent increment21Official-source additions, 1 August to 21 September 2026; not a full census
Combined candidate7,770Baseline plus increment; no Policy Commons merge
Cleaned candidate7,522After collision quarantine; further issuer and AI review required
Collision review ledger248Quarantined, not irreversibly deleted
Expanded pool16,954Overlapping source review pool; not an appendable increment

Task 5. Validate the existing Korean inputs

Real directories: policy_commons_ai_full_collection_20260910/stage2_stpi_korean_policy_reports_20260921/; kim_jongrip_ai_stpi_full_datasets_share_20260630/; policycommons_merge_workspace_20260921/04_stage2_stpi_kor_chn_jpn/00_frozen_inputs/.

  1. 5.1 Check source manifests, hashes, 144-column layout and row counts for the 7,749 baseline and 21 additions.
  2. 5.2 Trace the 7,770 combined rows through the 248 collision cases and 7,522 retained candidates. Review quarantined identities before final disposition.
  3. 5.3 Check the 16,954-row expanded pool for overlap, source form, issuer and full-text AI evidence before promoting individual records.

Task 6. Standardise Korean issuer and document evidence

Real directories: policycommons_merge_workspace_20260921/02_shared_crosswalks/{field_schema,issuer_country}/ and .../04_stage2_stpi_kor_chn_jpn/01_schema_and_country_review/.

  1. 6.1 Map actual source fields into a shared schema while retaining raw columns and source keys.
  2. 6.2 Confirm Korean issuer country from the issuing institution; keep it separate from topical country, author affiliation and publication place.
  3. 6.3 Check policy-report form, AI relevance and suspected title/abstract linkage errors against original sources. Full-text AI evidence remains valid even when metadata omits an AI term.

Task 7. Compare Korean records with Step 1

Real directories: policycommons_merge_workspace_20260921/03_stage1_current_master_plus_collector/04_versioned_release/ and .../04_stage2_stpi_kor_chn_jpn/02_identity_and_conflicts/.

  1. 7.1 Match typed provider identifiers, canonical URLs and DOIs with namespace and version checks.
  2. 7.2 Use normalised title, year and institution for human-review candidates; title alone is never an automatic merge key.
  3. 7.3 Record confirmed overlaps, new documents, different versions and unresolved collisions in a crosswalk.

Task 8. Validate and issue Step 2 outputs

Real directories: policycommons_merge_workspace_20260921/04_stage2_stpi_kor_chn_jpn/{03_quality_assurance,04_versioned_release}/; .../05_final_analysis_releases/; .../06_comparative_evaluation/baseline_vs_stage2/.

  1. 8.1 Add confirmed new Korean documents to a versioned master, retaining source provenance and a Korea-specific reconciliation.
  2. 8.2 Draw a reproducible stratified sample across source, year, institution and review category; inspect original documents.
  3. 8.3 Compare v7.10, Step 1 and Step 2 results before any reviewed promotion for analysis.

No release yet The Step 2 versioned-release directory was empty at review.

Release checks and deliverables

DeliverableUnit and checkDestination
Source observation tableOne source occurrence; reconcile files, RIS/CSV rows and repeated searches03_stage1_current_master_plus_collector/01_normalised_observations/
Document identity crosswalkSource occurrence to canonical version; retain conflicts and versions03_stage1_current_master_plus_collector/02_identity_and_conflicts/ and 04_stage2_stpi_kor_chn_jpn/02_identity_and_conflicts/
Versioned document mastersStep 1 v6: 738,737 unique document records; protected source-specific values and checksums03_stage1_current_master_plus_collector/04_versioned_release/step1_v83_plus_collector_20260923_v6_reviewed/
Strict quarantine releaseQ1, Q2, Q3 and strict-analysis partitions; mutually exclusive and count-conserving03_stage1_current_master_plus_collector/04_versioned_release/step1_v83_plus_collector_20260923_v7_strict_quarantine/
Final analysis categoriesDocument-country rows and separate international/private/unresolved sets05_final_analysis_releases/
Comparative auditBaseline versus each step, plus review samples and metrics06_comparative_evaluation/

Step 1 v6 and strict-quarantine v7 are issued and independently verified. Step 2 remains pending. Original-source validation is not yet complete, so heuristic title-summary cases remain quarantined rather than being labelled confirmed corruption.

Specialist review incorporated

Specialist reviews checked identity and deduplication, country and provenance, format contamination, administrative and event contamination, and metadata integrity. Their corrections are reflected here:

Source documents: migration and integration README, Korean staging README, Step 1 specialist resolution, strict-quarantine report, and workspace README.

Classification amendment, 24 September 2026

The counts above describe historical Step 1 releases. The applied-screen paragraph below reports the subsequent update to the active collected, deduplicated and analytical corpora.

Videos, recordings and blogs are excluded by default, with documented exceptions for substantive AI-policy content and an explicit country in source metadata. Media exceptions require substantive abstract evidence; blog exceptions require explicit AI-policy evidence in the title, abstract or keywords. Generic event announcements, AI-adjacent sectors and country mentions alone do not establish eligibility.

AI acronym checks use token boundaries and accept compounds such as AI-powered. Substrings in airport, airplane, availability, obtaining and explaining do not establish AI relevance. Metadata screening is distinct from original-text adjudication: missing metadata keywords alone do not justify deleting a full-text search result.

Preserve publication country, inferred target country and organisation headquarters as separate evidence roles. International coordination bodies, including CGMS, remain international; the German secretariat location is dictionary information. Supported health and cultural policy records remain eligible, including rural-health AI clinical tools, cybersecurity and Australian arts participation.

Remove human-confirmed exact false positives from all three active corpus stages by stable identifier and log the removal. Archival evidence may be retained outside the active pipeline. Retain format labels, exception grounds, evidence and unresolved review status for reproducibility and sensitivity analysis.

An official transcript or substantive recap may support a media exception where the collected abstract is truncated. Preserve the collected text and separately cite any enrichment. A genre keyword alone does not establish document format. Metadata-wide screening does not constitute full-text verification.

Specialist candidate review: all 238 initial exclusion candidates were reviewed: 179 provisional exclusions and 59 holds. Review of 22 exception candidates produced 13 approvals and 9 holds. These counts describe the review stage only; they do not certify completion of master-file updates.

Applied screen: 730,058 deduplicated records screened; 181 analysis exclusions, 13 specialist-reviewed media exceptions plus the RHTP exception, and 68 media holds. Active source-stage, deduplicated and country-expanded masters now contain 730,175, 730,056 and 730,068 rows respectively. Storage counts include excluded and pending records. Exact source removals affected 21 files and 31 occurrences for three distinct identifiers. The larger metadata-review queue contains 28,825 documents and is not a confirmed contamination count. See the updated curation protocol and review criteria and case evidence.